An Image Segmentation Method Based on Dilated Heterogeneous Convolution
By introducing hollow heterogeneous convolution into the image segmentation network, using channel grouping and multi-scale cavity rate, the problem of restricted applications of hollow convolution is solved, and a more efficient image segmentation effect is achieved.
Patent Information
- Application Number
- CN202211185277.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-27
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-09-27
AI Technical Summary
When used in multi-scale applications, existing hollow convolutions are limited by a single receptive field and cannot be used freely across the network like the basic convolution module, resulting in inconvenient multi-scale applications.
An image segmentation method based on hollow heterogeneous convolution is proposed. By grouping the convolution kernel channels and assigning different void rates to each channel packet, each channel packet adopts a different void rate, thereby achieving multi-scale effect in a single filter.
Without increasing the number of parameters, this method improves the effective receptive field of the convolution kernel, improves the accuracy of image segmentation, and effectively solves the grid problem, which has stronger versatility.
Smart Images

Figure CN115631137B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to an image segmentation method based on dilated heterogeneous convolution. Background Art
[0002] Traditional segmentation networks use downsampling or pooling layers in the encoder to compress feature maps to obtain a larger receptive field, which enables the network to effectively extract global features and semantic information in images. Subsequently, the resolution of the feature maps is restored by upsampling to output pixel-level prediction values end-to-end. However, the loss of detailed information in the feature maps caused by the reduction in resolution is inevitable. To overcome the above difficulties, methods such as FCN and Unet use skip connections to supplement uncompressed high-resolution feature maps to the decoder, and there are also methods like HRNet that take a different approach to fuse low-resolution features of parallel sub-networks while maintaining high resolution in the backbone network. As a basic module, dilated convolution is widely used to replace the pooling layer in the traditional encoder because it can expand the receptive field while retaining the high resolution of the feature maps. Its emergence effectively solves the bottleneck caused by feature compression. At the same time, multi-scale schemes based on dilated convolution have gradually become one of the research directions. For example, ASPP and ESPNet use parallel dilated convolution modules with multiple dilation rates to achieve multi-scale fusion, or use dilated convolution with different dilation rates in segmentation or object detection networks like HDC, DRN, and YOLOF to extract multi-scale context information.
[0003] The advantages of dilated convolution are due to the structural gain brought by the grid structure, but its disadvantages are still quite obvious: limited by the fact that dilated convolution only allows a single receptive field, existing methods can only set the dilation rate on the filters at different levels of the network to obtain multi-scale context information. However, these methods cannot be freely used everywhere in the network like the basic convolution module, making the multi-scale application of dilated convolution inconvenient. Summary of the Invention
[0004] The purpose of the present invention is to propose an image segmentation method based on dilated heterogeneous convolution for the above problems, which can add multiple scales to a single filter without increasing the number of parameters, thereby enhancing the effective receptive field of the convolution kernel and greatly improving the accuracy of image segmentation.
[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0006] An image segmentation method based on dilated heterogeneous convolution proposed by the present invention includes the following steps:
[0007] S1. Construct a convolutional neural network model, where one or more filters in the convolutional neural network model are replaced with dilated heterogeneous convolutional filters, and the settings of the dilated heterogeneous convolutional filters are as follows:
[0008] S11. Divide each dilated heterogeneous convolutional filter into n groups according to its number of channels C, and each group contains convolutional kernels;
[0009] S12. Preset a combination of dilation rates as a set of non-repeating hyperparameters r = [r1, r2, …, r i , …, r n , where r i is the i-th dilation rate, i = 1, 2, …, n, and corresponding to the convolutional kernels assigned to each group, so that the convolutional kernels in each group have the same dilation rate, and the convolutional kernels between different groups have different dilation rates, specifically as follows:
[0010] First, obtain a sequence S1 = <r1 - r2 - … - r i - … - r n > according to the original order in the hyperparameters r, and assign the values in the sequence S1 as dilation rates to the n groups of convolutional kernels of the first dilated heterogeneous convolutional filter in order; when it comes to the second dilated heterogeneous convolutional filter, shift the values in the sequence S1 as a whole to the right, so that the value at the end of the sequence S1 is moved to the first place to obtain a sequence S2 = <r n - r1 - r2 - … - r i - … - r n-1 >, and assign the values in the sequence S2 as dilation rates to the n groups of convolutional kernels of the second dilated heterogeneous convolutional filter in order; similarly, for the j-th dilated heterogeneous convolutional filter, and so on, j = 1, 2, …, N, where N is the total number of dilated heterogeneous convolutional filters, and when N is greater than the total number of permutations and combinations n of the sequence, perform cyclic operations, and finally N dilated heterogeneous convolutional filters will be formed in order;
[0011] S2. Use the dataset to train and validate the constructed convolutional neural network model to obtain the final convolutional neural network model;
[0012] S3. Input the image to be segmented into the final convolutional neural network model to obtain an output feature map;
[0013] S4. Use the image segmentation network model to process the output feature map to obtain an image segmentation result.
[0014] Preferably, the convolutional neural network model is a Resnet50 network.
[0015] Preferably, the Resnet50 network includes a convolutional layer, a pooling layer, a first residual block, a second residual block, a third residual block, and a fourth residual block connected in sequence. The first residual block includes three bottleneck layers connected in series, the second residual block includes four bottleneck layers connected in series, the third residual block includes six bottleneck layers connected in series, and the fourth residual block includes three bottleneck layers connected in series. Each bottleneck layer is composed of a filter with a convolution kernel size of 1*1, a filter with a convolution kernel size of 3*3, a filter with a convolution kernel size of 1*1, a BN layer, and a RELU activation function connected in series. Moreover, the filters with a convolution kernel size of 3*3 in the three bottleneck layers of the fourth residual block are replaced with atrous heterogeneous convolution filters with a convolution kernel size of 3*3.
[0016] Preferably, the convolution kernel size of the convolutional layer is 7*7.
[0017] Preferably, the dataset is the ade20k dataset.
[0018] Preferably, the constructed convolutional neural network model is trained and verified using the dataset. Specifically, images with a size of 512*512 in the dataset are input into the convolutional neural network model with a batchsize of 12 for training for 120 epochs.
[0019] Preferably, the image segmentation network model is one of the deeplabv3 network model, the pspnet network model, and the upernet network model.
[0020] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0021] The convolutional neural network model of this method uses atrous heterogeneous convolution filters. Different from using traditional filters, the atrous heterogeneous convolution filters group the channels of the convolution kernel and assign different dilation rates to each channel group, so that each channel group uses different dilation rates to reduce the overlap of the atrous regions during sampling, achieving an improvement in the effective coverage area, thereby enhancing the receptive field of the atrous convolution. And through the misalignment of the effective sampling regions of the convolution kernel, complementary effective pixels after projection are realized, which can effectively solve the grid problem. Moreover, the atrous heterogeneous convolution filters proposed by this method do not bring any additional number of parameters, and their number of parameters is exactly the same as that of traditional filters. Therefore, the improvement in performance does not come at any additional cost, and the accuracy of image segmentation can be greatly improved without increasing the number of parameters, and it has stronger generality. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is a flowchart of the image segmentation method based on atrous heterogeneous convolution of the present invention;
[0023] Figure 2 is a schematic structural diagram of the convolutional neural network model of the present invention;
[0024] Figure 3 Schematic diagrams of the dilated convolution filter (a) of the prior art and the dilated heterogeneous convolution filter (b) of the present invention;
[0025] Figure 4 Schematic diagram of the setting process of the dilated heterogeneous convolution filter of the present invention;
[0026] Figure 5 Visualization diagrams of the effective spatial coverage areas of the dilated convolution filter (a) of the prior art and the dilated heterogeneous convolution filter (b) of the present invention. Detailed implementation manners
[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0028] It should be noted that, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used in the specification of this application herein are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0029] As Figures 1-5 shown, an image segmentation method based on dilated heterogeneous convolution includes the following steps:
[0030] S1. Construct a convolutional neural network model, and one or more filters in the convolutional neural network model are replaced with dilated heterogeneous convolution filters. The settings of the dilated heterogeneous convolution filters are as follows:
[0031] S11. Divide each dilated heterogeneous convolution filter equally into n groups according to its number of channels C, and each group contains convolution kernels;
[0032] S12. Preset a combination of dilation rates as a set of non-repeating hyperparameters r = [r1, r2,..., r i ,..., r n , where r i is the i-th dilation rate, i = 1, 2,..., n, and correspondingly assign it to the convolution kernels on each group, so that the convolution kernels of each group have the same dilation rate, and the convolution kernels between different groups have different dilation rates. Specifically as follows:
[0033] First, obtain the sequence S1 = <r1 - r2 -... - r i -... - rn >, and the values in sequence S1 are sequentially and correspondingly assigned as the dilation rates to the n groups of convolutional kernels of the first dilated heterogeneous convolutional filter; when it comes to the second dilated heterogeneous convolutional filter, the values in sequence S1 are shifted to the right as a whole, and the value at the end of sequence S1 is moved to the first place to obtain sequence S2 = <r n -r1-r2-…-r i -…-r n-1 >, and the values in sequence S2 are sequentially and correspondingly assigned as the dilation rates to the n groups of convolutional kernels of the second dilated heterogeneous convolutional filter; similarly, for the j-th dilated heterogeneous convolutional filter, and so on, j = 1, 2, …, N, where N is the total number of dilated heterogeneous convolutional filters, and when N is greater than the total number of permutations and combinations n of the sequence, the loop operation is performed, and finally N dilated heterogeneous convolutional filters will be formed in order.
[0034] In one embodiment, the convolutional neural network model is a Resnet50 network.
[0035] In one embodiment, the Resnet50 network includes a convolutional layer, a pooling layer, a first residual block, a second residual block, a third residual block, and a fourth residual block connected in sequence. The first residual block includes three serially connected bottleneck layers, the second residual block includes four serially connected bottleneck layers, the third residual block includes six serially connected bottleneck layers, and the fourth residual block includes three serially connected bottleneck layers. Each bottleneck layer is composed of a filter with a convolution kernel size of 1*1, a filter with a convolution kernel size of 3*3, a filter with a convolution kernel size of 1*1, a BN layer, and a RELU activation function connected in series, and the filters with a convolution kernel size of 3*3 in the three bottleneck layers of the fourth residual block are replaced by dilated heterogeneous convolutional filters with a convolution kernel size of 3*3.
[0036] In one embodiment, the convolution kernel size of the convolutional layer is 7*7.
[0037] Specifically, as Figure 2 shown, in this embodiment, the ResNet50 network is used as the backbone for image segmentation. For the ade20k dataset, the input of the Resnet50 network is a natural landscape image with a size of 512×512 and 3 channels. The first layer of the ResNet50 network is a convolutional layer with a convolution kernel size of 7*7. After the convolutional layer, a pooling layer is connected. After the pooling layer, four residual blocks are connected. The four residual blocks sequentially contain 3, 4, 6, and 3 bottleneck layers. Each bottleneck layer is composed of a group of convolutional kernels of 1*1, 3*3, and 1*1, as well as a BN layer and a RELU activation function connected in series.
[0038] For the fourth residual block in the backbone, replace the 3×3 convolutional kernels in the three bottleneck layers it contains with dilated heterogeneous convolution filters, which are sequentially denoted as the first dilated heterogeneous convolution filter, the second dilated heterogeneous convolution filter, and the third dilated heterogeneous convolution filter. Taking the number of channels C = 16, n = 4, and r = [1, 2, 3, 4] as an example, to ensure that each channel of the input feature map can be reasonably received by convolutional kernels of each scale, the dilated heterogeneous convolution filters will obtain the sequence S1 = <1 - 2 - 3 - 4> in the original order in the hyperparameter r. The values in this sequence are used as dilation rates and are sequentially (e.g., from top to bottom) assigned to the 4 groups of convolutional kernels of the first dilated heterogeneous convolution filter; when it comes to the second dilated heterogeneous convolution filter, the values in the sequence are shifted to the right as a whole, and the last value is moved to the first place to obtain the new sequence S2 = <4 - 1 - 2 - 3>. The values in the sequence S2 are used as dilation rates and are sequentially assigned to the n groups of convolutional kernels of the second dilated heterogeneous convolution filter; when it comes to the third dilated heterogeneous convolution filter, the values in the sequence are shifted to the right again as a whole, and the last value of the sequence S2 is moved to the first place to obtain the new sequence S3 = <3 - 4 - 1 - 2>. The values in the sequence S3 are used as dilation rates and are sequentially assigned to the n groups of convolutional kernels of the third dilated heterogeneous convolution filter. Finally, 3 dilated heterogeneous convolution filters with different arrangements of dilation rates will be generated in order. It should be noted that when N is greater than the total number of permutations and combinations n of the sequence, a loop operation is performed, that is, starting from the sequence S1, the assignment is looped again, and N is the total number of dilated heterogeneous convolution filters.
[0039] As shown in the figure, Figure 3 Figure (a) shows the structural schematic diagram of the traditional dilated convolution filter and the unfolded diagrams of its channels when the dilation rate is 2 and the number of channels is 4. Figure (b) shows the structural schematic diagram of the dilated heterogeneous convolution filter of the present application and the unfolded diagrams of its channels when the dilation rate combination is [1, 2, 3, 4] and the number of channels is 4.
[0040] Figure 4 Figure (a) shows the differences in dilation rates of the standard convolution filter, the traditional dilated convolution filter, and the dilated heterogeneous convolution filter at the channel level. For example, the standard convolution filter uses a single dilation rate r = 1, the traditional dilated convolution filter uses a single dilation rate r = 3, and the dilated heterogeneous convolution filter uses the dilation rate combination [1, 2, 3, 4]. Figure (b) shows the specific setting process of N dilated heterogeneous convolution filters.
[0041] Figure 5Figure (a) shows a traditional dilated convolution filter, such as the effective spatial coverage area when the dilation rate r = 4. The gray block area is the effective spatial coverage area (skeleton). Figure (b) shows the effective spatial coverage areas of the dilated heterogeneous convolution filter of the present invention with different dilation rates, such as when the dilation rates r = 1, 2, 3, 4. The gray block area is the effective spatial coverage area.
[0042] S2. Use the data set to train and validate the constructed convolutional neural network model to obtain the final convolutional neural network model.
[0043] In one embodiment, the data set is the ade20k data set.
[0044] In one embodiment, using the data set to train and validate the constructed convolutional neural network model specifically means inputting the 512*512 images in the data set into the convolutional neural network model with a batch size of 12 for training for 120 epochs.
[0045] S3. Input the image to be segmented into the final convolutional neural network model to obtain an output feature map.
[0046] S4. Use the deeplabv3 network model to process the output feature map to obtain the image segmentation result.
[0047] In one embodiment, the image segmentation network model is one of the deeplabv3 network model, the pspnet network model, and the upernet network model. The decode_head of the image segmentation in this embodiment uses the classic deeplabv3 network model.
[0048] The convolutional neural network model of this method uses a dilated heterogeneous convolution filter. Different from using a traditional filter, the dilated heterogeneous convolution filter groups the channels of the convolution kernel and assigns different dilation rates to each channel group, so that each channel group uses different dilation rates to reduce the overlap of the dilated areas during sampling, realizes the improvement of the effective coverage area, thereby enhancing the receptive field of the dilated convolution, and realizes the complementarity of effective pixels after projection through the dislocation of the effective sampling areas of the convolution kernel. It can effectively solve the grid problem, and the dilated heterogeneous convolution filter proposed by this method does not bring any additional number of parameters, and its number of parameters is exactly the same as that of the traditional filter. Therefore, the improvement of performance does not pay any additional cost, and it can greatly improve the accuracy of image segmentation without increasing the number of parameters and has stronger versatility.
[0049] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0050] The above-described embodiments only express the embodiments of the present application that are relatively specific and detailed in description, but should not be construed as a limitation on the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. An image segmentation method based on dilated heterogeneous convolution, characterized in that: The image segmentation method based on dilated heterogeneous convolution includes the following steps: S1. Construct a convolutional neural network model, where one or more filters in the convolutional neural network model are replaced by dilated heterogeneous convolution filters. The settings of the dilated heterogeneous convolution filters are as follows: S11. Divide each hole heterogeneous convolution filter into n groups according to its number of channels C, and each group contains convolution kernels; S12. The preset void ratio combination is a set of non-repeating hyperparameters r = [r1, r2, …, r i , …, r n , where r i is the i-th void ratio, i = 1, 2, …, n, corresponding to the convolutional kernels assigned to each group, such that the convolutional kernels of each group have the same void ratio, and the convolutional kernels between different groups have different void ratios, specifically as follows: First, obtain the sequence S1 = <r1 - r2 - … - r i -…- r n > in the original order of the hyperparameter r, and sequentially assign the values in the sequence S1 as the dilation rates to the n groups of convolutional kernels of the first dilated heterogeneous convolutional filter; when it comes to the second dilated heterogeneous convolutional filter, shift the values in the sequence S1 as a whole to the right, and the value at the end of the sequence S1 is shifted to the first place to obtain the sequence S2 = <r n - r1 - r2 - … - r i -…- r n-1 >, and sequentially assign the values in the sequence S2 as the dilation rates to the n groups of convolutional kernels of the second dilated heterogeneous convolutional filter; similarly, for the j-th dilated heterogeneous convolutional filter, and so on, j = 1, 2, …, N, where N is the total number of dilated heterogeneous convolutional filters, and when N is greater than the total number of permutations and combinations n of the sequence, perform cyclic operations, and finally N dilated heterogeneous convolutional filters will be formed in order; S2. Use a dataset to train and validate the constructed convolutional neural network model to obtain a final convolutional neural network model; S3. Input the image to be segmented into the final convolutional neural network model to obtain an output feature map; S4. Use an image segmentation network model to process the output feature map to obtain an image segmentation result; The convolutional neural network model is a Resnet50 network. The Resnet50 network includes a convolutional layer, a pooling layer, a first residual block, a second residual block, a third residual block, and a fourth residual block connected in sequence. The first residual block includes three bottleneck layers connected in series. The second residual block includes four bottleneck layers connected in series. The third residual block includes six bottleneck layers connected in series. The fourth residual block includes three bottleneck layers connected in series. Each bottleneck layer is composed of a filter with a kernel size of 1*1, a filter with a kernel size of 3*3, a filter with a kernel size of 1*1, a BN layer, and a RELU activation function connected in series. Moreover, the filters with a kernel size of 3*3 in the three bottleneck layers of the fourth residual block are replaced by dilated heterogeneous convolution filters with a kernel size of 3*3.
2. The image segmentation method based on dilated heterogeneous convolution according to claim 1, characterized in that: The convolutional kernel size of the convolutional layer is 7*7.
3. The image segmentation method based on dilated heterogeneous convolution according to claim 1, characterized in that: The dataset is the ade20k dataset.
4. The image segmentation method based on dilated heterogeneous convolution according to claim 1, characterized in that: The step of using the dataset to train and validate the constructed convolutional neural network model is specifically to input images with a size of 512*512 in the dataset into the convolutional neural network model for training for 120 epochs with a batchsize of 12.
5. The image segmentation method based on dilated heterogeneous convolution according to claim 1, characterized in that: The image segmentation network model is one of the deeplabv3 network model, the pspnet network model, and the upernet network model.
Citation Information
Patent Citations
Improved semantic segmentation method based on DeepLabv3+
CN113139551A
Omni-scale convolution for convolutional neural networks
WO2022133814A1