A method for image segmentation based on FEFNet network structure
By reasonably setting the hollow rate and performing feature fusion in the FEFNet network structure, the problem of extracting feature information of objects of different sizes in the prior art is solved, and efficient and accurate feature extraction and classification accuracy are achieved.
Patent Information
- Application Number
- CN202310327553.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-30
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-03-30
AI Technical Summary
The prior art is difficult to efficiently and accurately extract the characteristic information of objects of different sizes, and the convolution of the cavity will lead to a grid effect when the cavity rate is too large, reducing the effectiveness of feature extraction of small objects.
The FEFNet network structure is adopted, and the hollow rate is reasonably set to {3,5} and {5,7}, and feature fusion is carried out multiple times in the network to achieve information complementarity to improve network performance.
Effectively extract feature information of objects of different sizes, improve classification accuracy, avoid grid effects, and significantly improve the effectiveness and detection speed of the algorithm.
Smart Images

Figure CN116342957B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of FEFNet network structure, and specifically relates to a method for image segmentation based on the FEFNet network structure. Background Art
[0002] Convolutional Neural Network (CNN) is a network model commonly used in the field of images. CNN is particularly suitable for the field of images because of its characteristics of scaling, translation, and rotation invariance. At the same time, CNN also has the characteristics of local connection. Convolution is a local connection. The convolution operation is continuously superimposed, and information is transmitted layer by layer. The field of vision gradually expands from the local to the whole, which is in line with the characteristics of people looking at things from the local to the whole. The size of the receptive field of the convolution operation depends on the size of the convolution kernel. CNN includes a variety of layers and operations with different functions, which will be introduced one by one below.
[0003] 1. Convolutional Layer
[0004] The convolution layer uses convolution operations to extract image features, which is similar to filtering operations in digital image processing. The convolution kernel slides on the image in the form of a sliding window, and the sliding trajectory of the convolution kernel starts from the upper left corner of the image to the lower right corner of the image. Specifically, the convolution kernel uses pixels as the basic unit, and each time the sliding window moves stride pixels to the right until it reaches the rightmost side, then returns to the leftmost side and moves stride pixels downward, and so on, until it reaches the lower right corner of the image.
[0005] The size, number of channels, number, stride, and padding of the convolution kernel can all be set in advance as parameters. Different parameters can produce different convolution results. The number of convolution kernels is the depth of the output feature map.
[0006] Assume that the size of the input image is C1×H1×W1, the size of the convolution kernel is C1×K×K, there are N convolution kernels, the padding value is P, and the stride value is S. The calculation formulas for the height, width, and number of channels of the output feature map are as follows:
[0007]
[0008] C 2 =N
[0009] When performing convolution operations, if you do not want the feature map to be downsampled too much, you can add all-0 pixels to the edge of the feature map, that is, padding. For example, when padding is 2, two rows of all-0 pixels are added to the top, bottom, left, and right of the image. The purpose of padding is mainly to adjust the output size, because in the process of repeated convolution operations, if there is no padding operation, the feature map will become smaller and smaller, making it impossible to continue convolution. For example, using a convolution of size 3×3, stride 1, and padding 1, the size of the feature map after convolution is the same as the original Figure 1 To.
[0010] 2. Pooling layer
[0011] Pooling is a commonly used downsampling method in convolutional neural networks. One of the main functions of the pooling layer is to operate on the feature map without learning additional parameters to reduce the size of the feature map, while reducing overfitting. Pooling methods mainly include maximum pooling and average pooling, both of which can achieve nonlinear transformation of feature maps. The specific method of pooling is to select a fixed pooling area and traverse the feature map in the form of a sliding window. If the selected pooling method is maximum pooling, the maximum value in the sliding window is taken as the corresponding value in the result. If it is average pooling, the average of the values in the sliding window is taken as the corresponding value in the result. The sliding trajectory of the sliding window is similar to the convolution operation, and the step size is pixels each time. For example Figure 1 As shown in the figure, two 4×4 matrices are subjected to maximum pooling and average pooling respectively using a 2×2 pooling window with a step size of 2, resulting in two 2×2 matrices. The pooling operation reduces the size of the feature map to 1 / 2 of the original size.
[0012] Average pooling takes the average of the values in the pooled area, so it can retain more of the image background. Max pooling can retain as many texture features of the image as possible, and is more commonly used in the image field. The pooling layer has several features. First, unlike the convolutional layer, it does not need to learn weight parameters. It only performs specific calculations on the values in the target window. Secondly, the calculation of the pooling layer is performed channel by channel, and the pooling operation will not cause changes in the number of output channels. Finally, the pooling layer is insensitive to small changes in the input data. For example, for maximum pooling, as long as the maximum value in the window remains unchanged, changes in the remaining values will have no effect on the output result.
[0013] Average pooling and maximum pooling both refer to common pooling. There is also a type of pooling called global pooling, where the pooling area is the entire feature map. There are two types of pooling: global average pooling and global maximum pooling. The global pooling layer can give the network a global view. Global pooling can replace the fully connected layer without introducing additional parameters, and is simple to use.
[0014] 3. ReLU activation function
[0015] ReLU is the simplest activation function and works very well. It can make the network converge quickly and is generally used in the middle layer of the neural network. However, it also has a disadvantage. For values less than 0, the gradient will always be 0, and the weight cannot be updated in this propagation. This situation is called the Dead ReLU problem. There are generally two reasons for this situation. One is that the parameters are not initialized well, which is relatively rare. The other is that the learning rate is too high, resulting in excessive parameter updates. The ReLU calculation method is as follows: Figure 2 As shown:
[0016] f(x)=max(x,0)
[0017] 4. Upsampling
[0018] Most image semantic segmentation networks can be divided into two stages: encoding and decoding. The encoding stage is to extract image features by gradually downsampling through methods such as convolution and pooling. Therefore, the image is downsampled in the encoding stage to reduce its size, while the receptive field is increased, thereby capturing the high-level features of the input feature map. The decoder needs to restore the feature map to the size of the original image. Upsampling is a process of increasing the size of the feature map, but upsampling usually also causes a loss of accuracy and the feature map cannot be fully restored. Common upsampling methods include transposed convolution, nearest neighbor interpolation, and bilinear interpolation. Transposed convolution is also called deconvolution.
[0019] 1. Deconvolution
[0020] Ordinary convolution maps multiple values to one value, while deconvolution maps one input value to multiple output values, which is a one-to-many mapping relationship. Deconvolution can actually be regarded as a special forward convolution. Before convolution, the size of the feature map is enlarged by padding it with 0, and then convolution is performed to upsample the feature map.
[0021] 2. Nearest neighbor interpolation method
[0022] The idea of the nearest neighbor interpolation method is easy to understand. As the name suggests, the nearest neighbor interpolation is to make the value of the new pixel inserted after upsampling equal to the value of the nearest pixel around it. Due to the simple operation, the upsampling result is generally not very good, and there will be a more obvious jagged feeling.
[0023] 3. Bilinear interpolation
[0024] The upsampling method used in this article is bilinear interpolation. Bilinear interpolation does not have the jagged feeling of the nearest neighbor interpolation method, and can guarantee the accuracy of the image to a certain extent. It is simple and efficient. The core idea is to perform linear interpolation operations in the X direction and the Y direction respectively.
[0025] like Figure 3 The coordinate diagram of each point is shown in the figure. The horizontal and vertical coordinates represent the position of the pixel point, and f(·) represents the gray value of the pixel point. If the gray values of the pixels M11, M12, M21, and M22 closest to point P are known, the gray value of P can be calculated using the surrounding pixel values. First, perform two linear interpolation operations in the x-axis direction. The calculation method is as follows: Figure 4 Then, a linear interpolation is performed in the y-axis direction to obtain the gray value of the pixel where point P is located, that is, f(P).
[0026]
[0027] Prior art 1
[0028] Atrous convolution is also called dilated convolution or expanded convolution. Simply put, it is to add some spaces (zeros) between the convolution kernel elements to expand the size of the convolution kernel. Expanding the size of the convolution kernel means expanding the receptive field. Atrous convolution can expand the receptive field without losing resolution and keep the relative spatial position of pixels unchanged. This can enhance the extraction of target features without increasing model parameters to improve the effect of the algorithm. Using a variable a to measure the atrous rate of atrous convolution, the relationship between the actual convolution kernel size after adding the atrous and the original convolution kernel size is: K = K + (k-1)(a-1).
[0029] Figure 4 The convolution kernel size and receptive field are shown when the dilation rate is 1, 2, and 4. The rectangular box surrounded by red dots represents the convolution kernel size with different dilation rates. In this rectangular box, the values of the pixels between the red dots are all 0. The receptive field corresponding to different dilation rates is also different. When the dilation rate is 1, the corresponding receptive field is 3x3, when the dilation rate is 2, the corresponding receptive field is 7x7, and when the dilation rate is 3, the corresponding receptive field is 15x15. Figure 4 The blue rectangular frame is in the middle.
[0030] Disadvantages of the prior art 1
[0031] Although the atrous convolution changes the receptive field of the convolution kernel by controlling the size of the atrous ratio, if the atrous ratio is too large, it will cause a gridding effect. The reason for this problem is that when the atrous ratio is too large, not all input pixels are calculated, that is, the convolution kernel is discontinuous. The use of atrous convolution is effective for the segmentation of large objects, but it may not be beneficial for the segmentation of small objects. On the contrary, when the atrous ratio of the atrous convolution in the network is set too large, it will reduce the effectiveness of the network in extracting features of small objects. Figure 5As shown in the figure, this visualization is the effect of DABNet on the Cityscapes test set. In the DABNet network, the maximum void ratio reaches 16, which may be one of the reasons why DABNet performs poorly on the test set. How to simultaneously process feature extraction of objects of different sizes is the key to designing a good void convolutional network.
[0032] Prior art 2
[0033] Most of the early methods2,3,4 have achieved good results. PSPNet5 uses pyramid pooling module (PPM) and DeepLabv26 applies Atrous spatial pyramid pooling (ASPP) to explore contextual information. They have achieved remarkable results, but are limited by inference speed and computing power. PSPNet contains 65.7 million parameters and DeepLabV2 contains 54.6 million parameters. Such a scale of parameters indicates that they do not perform well in inference speed (less than one frame per second), so they do not meet the standards of lightweight semantic segmentation networks. In addition, ENet7 designed a small encoder-decoder network model with an inference speed of 76.9, but the classification accuracy is only 58.3%. ESPNet8 effectively uses spatial pyramid modules and convolution decomposition methods to collect multi-scale contextual information and reduce parameters. The fewer parameters, the faster the inference speed. Although ESPNet performs well in inference speed, it is not satisfactory in classification accuracy. Recently, some new networks have emerged. MSCFNet9 created an effective asymmetric residual (EAR) module and attention module for the multi-scale context fusion network, which improved the segmentation accuracy and inference speed. DABNet10 built a bottleneck structure to extract context information using deep asymmetric convolution and dilated convolution. However, the DAB module explores the intrinsic correlation between feature maps through simple linear operations, which cannot fully utilize the intrinsic relationship between feature maps. In addition, the use of dilated convolutions with large dilation rates will lead to grid problems, which DABNet does not solve well. These networks often only focus on the segmentation accuracy or detection speed of the network, and have never designed a network that can take into account both segmentation accuracy and detection speed.
[0034] Lightweight semantic segmentation networks need to find a balance between the amount of computation (GFLOPS), the number of parameters (Parameters), the detection speed (FPS) and the accuracy (mIoU). They hope to achieve high detection speed and model accuracy while using as little computation and parameters as possible, so as to meet the needs of autonomous driving, robot navigation segmentation, medical image analysis and other fields.
[0035] The present invention aims to solve the problem of how to use dilated convolution to efficiently and accurately extract feature information of objects of different sizes. In addition, in order to improve the classification accuracy, this patent reasonably integrates the local and boundary feature information of the object to further improve the effectiveness of the algorithm. Figure 6 This is the effect diagram of this patent on the cityscapes11 test set. Figure 2 In comparison, the present invention performs better in segmenting small objects. Summary of the invention
[0036] In order to solve the above technical problems, the present invention provides a FEFNet network structure to solve how to use dilated convolution to efficiently and accurately extract feature information of objects of different sizes while improving classification accuracy.
[0037] A FEFNet network structure includes the following steps:
[0038] Step S1: Perform network extraction on the image to obtain underlying feature information;
[0039] Step S2: Perform jump links on the network input image to enrich feature information;
[0040] Step S3: down-sample the feature information to obtain down-sampling module information;
[0041] Step S4: Process the downsampling module information with the FEF module to obtain the required image.
[0042] Preferably, step S1 includes the following sub-steps:
[0043] Sub-step S11: The first 3x3 convolution has 3 input channels, 32 output channels, 3 kernels, 2 strides, and 1 padding.
[0044] Sub-step S12: After the convolution, a BN operation is performed on the result;
[0045] Sub-step S13: Then the activation function is PReLu.
[0046] Preferably, step S2 specifically implements the skip link part by average pooling, the size of the average pooling is 3, the step length is 2, and the padding is 1.
[0047] Preferably, step S3 includes the following sub-steps:
[0048] Sub-step S31: performing a convolution operation on the input feature map, with a convolution kernel size of 3, a step size of 2, a padding of 1, and the number of input and output channels being given parameters;
[0049] Sub-step S32: Determine the input dimension and output dimension of the sampling module; if the input dimension is smaller than the output dimension, proceed to sub-steps S33, S34, and S35; if the input dimension is not smaller than the output dimension, directly perform sub-step S35 on the result of sub-step S31;
[0050] Sub-step S33: performing a maximum pooling operation with a size of 2 and a step size of 2 on the input feature map;
[0051] Sub-step S34: Dimensionally concatenate the results of sub-step S31 and sub-step S32;
[0052] Sub-step S35: Perform BN and PReLu operations on the feature map, and the result obtained is the result of the downsampling module and is output.
[0053] Preferably, step S4 includes the following sub-steps:
[0054] Sub-step S41: extracting local information of the target and reducing the dimension of the feature map through a convolution operation with a convolution kernel size of 3, an output channel number that is half the number of input channels, a step size of 1, and a padding of 1;
[0055] Sub-step S42: Use two boundary feature extractors to extract boundary feature information of the target, refer to ERFNet 13 In this method, each boundary extractor uses depth-wise separable convolution and dilated convolution techniques;
[0056] Sub-step S43: In sub-step S42, after the operation of the first boundary extractor, its result is added to the boundary feature map and the result of sub-step S41; after the second boundary extractor, its result is dimensionally spliced with the FEF module input feature map, the boundary feature map and the result of sub-step S41;
[0057] Sub-step S44: The FEF module converts the dimension into the input dimension of the FEF module through a 1x1 convolution.
[0058] Preferably, sub-step S44 is specifically:
[0059] The number of input channels of sub-step S44 is twice the number of channels of the input feature map of the FEF module, the number of output channels is the number of channels of the input feature map of the FEF module, the convolution kernel size is 1, the step size is 1, and the padding is 0.
[0060] The beneficial effects of the FEFNet network structure of the present invention are as follows:
[0061] 1. The algorithm of the present invention sets the void ratio reasonably. The void ratios set in the algorithm of the present invention are {3,5} and {5,7} respectively. Such a void ratio combination can better take into account objects of different sizes at the same time.
[0062] 2. After reasonably solving the problems of dilated convolution, feature fusion is performed multiple times in the network to achieve information complementarity, thereby improving network performance. This is more efficient than other algorithms, and the experimental results also verify this.
[0063] 3. The accuracy of the present invention is good. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show a certain embodiment of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0065] Figure 1 Figure 2 is a diagram of two pooling calculation methods of the present invention;
[0066] Figure 2 It is the ReLU activation function diagram of the present invention;
[0067] Figure 3 is a bilinear interpolation graph of the present invention;
[0068] Figure 4 The convolution kernel size and receptive field diagram of different void ratios of the present invention;
[0069] Figure 5 A DABNet visualization diagram of the present invention;
[0070] Figure 6 It is a visualization diagram of the algorithm of the present invention;
[0071] Figure 7 It is the FEFNet network structure diagram of the present invention;
[0072] Figure 8 It is a DownSampling module diagram of the present invention; DETAILED DESCRIPTION
[0073] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0074] The network structure of the present invention is as follows Figure 8 As shown, the input of the network is a three-channel RGB image of 512x1024. The detailed steps of the present invention will be introduced one by one below.
[0075] Backbone network: In order to reduce the amount of network parameters while ensuring the ability to extract rich feature information,
[0076] The present invention does not use the conventional backbone feature extraction network (such as Resnet12, etc.). Although the backbone extraction network can extract more accurate and rich target feature information and improve network performance, due to the characteristics of the conventional backbone feature extraction network with deep network layers and large number of parameters, it often cannot meet the requirements of lightweight semantic segmentation research in terms of detection speed. Therefore, in order to better balance the detection speed and detection accuracy of the network,
[0077] The backbone extraction part of the present invention uses three conventional 3x3 convolutions to achieve the underlying feature extraction of the target (image).
[0078] The first 3x3 convolution has 3 input channels, 32 output channels, 3 kernels, 2 strides, and 1 padding. After the convolution, a BN operation is performed on the result to prevent overfitting, and then the activation function is PReLu.
[0079] The second and third 3x3 convolutions are similar to the first 3x3 convolution operations, except that the second and third convolutions have 32 input channels, 32 output channels, and a stride of 1. Like a convolution, both convolutions are followed by BN and PReLu operations.
[0080] Through these three convolution operations, rich underlying feature information can also be extracted.
[0081] Although doing so will slightly reduce the performance of the network, it can help the network perform well in detection speed, and a slight decrease in network performance is acceptable in lightweight semantic segmentation research.
[0082] Jump Links ( Figure 7 The skip link downsamples the input image (RGB three-channel image) of the network and combines it with the underlying feature information obtained in step S1 of the network to obtain rich feature information.
[0083] The skip link part of the present invention is implemented by average pooling, and the average pooling size is 3, the step size is 2, and the padding is 1.
[0084] Downsample once (corresponding to Figure 7 The dashed line 1 / 2) performs this operation once, downsampling twice (corresponding to Figure 7 The dotted line 1 / 4) performs this operation twice, downsampling three times (corresponding to Figure 7 The dotted line 1 / 8) performs this operation three times.
[0085] DownSampling: The DownSampling operation is used to downsample the input feature map (for example, the network input image (size is 512*1024), after the maximum pooling operation with a size of 2 and a step size of 2, the size becomes 256*512), so as to expand the receptive field and enable the algorithm to extract accurate and rich feature information in the subsequent process.
[0086] like Figure 8 As shown, the DownSampling module of the present invention first performs a convolution operation on the input feature map (the convolution kernel size is 3, the step size is 2, the padding is 1, and the number of input and output channels is a given parameter).
[0087] Then, a maximum pooling operation with a size of 2 and a stride of 2 is performed on the input feature map.
[0088] Finally, the two results are dimensionally concatenated and output as the result of the DownSampling module.
[0089] Since the present invention does not use a conventional feature extraction network in the backbone feature extraction part, some feature information will be lost, resulting in the extracted feature information being inaccurate. In order to make up for the feature information of the backbone part, this network fuses the feature maps multiple times to achieve the purpose of enriching the feature information.
[0090] This network is down-sampled twice and up-sampled ( Figure 7 Before the UpSampling operation, the feature map is fused with the corresponding skip link feature map and then downsampled.
[0091] Such operations can enrich the feature information of the target, make up for the lost underlying feature information, and provide better support for the subsequent extraction of high-level and semantic information.
[0092] However, too many DownSampling operations will lead to the loss of target features and position information, which is not conducive to segmentation. Therefore, the present invention only performs three downsampling operations in total (the first is the first 3x3 convolution of the backbone feature extraction part, the second and third are Figure 8 The two DownSampling operations in ).
[0093] FEF block: To reduce the number of parameters and computational complexity, and to find a balance between FPS and mIoU.
[0094] FEF module of the present invention
[0095] First, a convolution operation with a convolution kernel size of 3, half the number of output channels as the input channels, a step size of 1, and a padding of 1 is used to extract the local information of the target and reduce the dimension of the feature map.
[0096] Although 1x1 convolution can halve the dimension and add much fewer parameters than 3x3 convolution, the use of 1x1 convolution will cause the loss of local information. If 1x1 convolution is used, more network layers are required to make up for the lost information, which will increase more parameters and network complexity. Therefore, in order to achieve the effect of extracting local information and halving the dimension at the same time, the present invention uses 3x3 convolution to extract local feature information.
[0097] In addition, the present invention uses two boundary feature extractors to extract boundary feature information of the target. Referring to the method of ERFNet13, each boundary extractor adopts depth-separable convolution and dilated convolution technology.
[0098] Depthwise separable convolution refers to using an nx1 and a 1xn depthwise separable convolution instead of a regular nxn convolution. Using depthwise separable convolution technology can reduce computational costs and speed up training, but the disadvantage is that it will reduce network performance. Using dilated convolution instead of regular 3x3 convolution is to take into account targets of different sizes.
[0099] In FEF block 1, the hole rate of the first boundary feature extractor is 3 (the number of input and output channels of the 3x1 convolution are both the number of channels after dimensionality reduction, the convolution kernel size is (3, 1), the step size is 1, the hole rate is (current hole rate, 1), and the padding is (current hole rate, 0); the number of input and output channels of the 1x3 convolution are both the number of channels after dimensionality reduction, the convolution kernel size is (1, 3), the step size is 1, the hole rate is (1, current hole rate), and the padding is (0, current hole rate)). After the first boundary extractor, the present invention adds it to the local feature information to enrich the feature information of the target, and also plays a role in accurately locating the target and accurately classifying pixels, and uses it as the input of the second boundary extractor. The hole rate of the second boundary feature extractor is 5, and the hole rate is set slightly larger than that of the first boundary feature extractor. The reason is that if the size of the target is large, then when using a convolution with a smaller hole rate to extract its feature information, it is still possible that the extracted information is inaccurate and the semantic classification results of the same target are discontinuous (such as Figure 6 ) and other issues. As shown in Table 1.
[0100] Table 1
[0101]
[0102] After passing through the second boundary extractor, the present invention concatenates the result with the local feature map and the FEF module input feature map to maximize the use of these feature information, making the local and boundary feature information of the target more accurate and rich. Finally, the FEF module converts the dimension into the input dimension of the FEF module through a 1x1 convolution (the number of input channels is twice the number of channels of the FEF module input feature map, the number of output channels is the number of channels of the FEF module input feature map, the convolution kernel size is 1, the step size is 1, and the padding is 0).
[0103] There are three such FEF modules in FEF block 1. After FEF block 1, its result is dimensionally concatenated with the downsampled feature map of the network input image, and then the DownSampling operation is performed as a whole, which plays a role in downsampling the feature map and expanding the receptive field. In FEF block 2, the hole rate of the first boundary feature extractor is 5, and the hole rate of the second boundary feature extractor is 7. FEF block 2 contains 7 such FEF modules.
[0104] After executing FEF block 2, the output of FEF block 2 is compared with the downsampled result of the network input ( Figure 8 The feature information is enriched by combining the 1 / 8 dotted line in the middle and upper part of the image, and then a 1x1 convolution is used to convert the dimension to the number of categories in the dataset (19 for Cityscapes and 11 for CamVid14) (the convolution kernel size is 1, the step size is 1, and the padding is 0). Finally, the commonly used bilinear interpolation method is used for decoding and the segmented image is obtained.
[0105] In order to ensure that the algorithm can take into account targets of different sizes at the same time, as shown in Table 2, the present invention conducts experiments on multiple sets of data.
[0106] Table 2
[0107]
[0108] It can be seen from Table 2 that when the void ratio of the FEF-2 module becomes larger, it is not conducive to the results. This is because when the void ratio is larger, the network is prone to grid effect, which will reduce the performance of the network. Therefore, {3,5} and {5,7} are selected as the void ratios of the two parts.
[0109] Table 3 The present invention also experiments on the number of FEF modules in FEF block. It can be found that when the number of FEF modules in FEF block 1 increases, mIoU shows an upward trend, but the increase is slow, while when the number of FEF modules in FEF block 2 increases, mIoU also shows an upward trend, and the increase is large. This is because FEF block 2 contains a large amount of high-level feature information, and reasonable processing of high-level feature information is the key.
[0110] Table 3
[0111]
[0112] As shown in Table 3, the mIoU of this patent on the cityscapes test set reached 72.6%, which is better than other algorithms. In addition, although this algorithm is not the best in detection speed, it is still in the upper middle position. Overall, this algorithm has achieved very good results.
[0113] The reason why such a result can be achieved is that the patented algorithm reasonably sets the void ratio. In this algorithm, the void ratios set are {3,5} and {5,7} respectively. Such a void ratio combination can better take into account objects of different sizes at the same time.
[0114] Secondly, after reasonably solving the problems of dilated convolution, feature fusion is performed multiple times in the network to achieve information complementarity, thereby improving network performance. This is more efficient than other algorithms (DABNet10, etc.), and the experimental results also verify this.
[0115] Table 2 shows the performance of the present invention on the CamVid dataset.
[0116] Compared with the Cityscapes dataset, the CamVid dataset contains more large objects (such as cars, buildings, etc.). With a larger dilation rate, DABNet9 can extract richer feature information. In contrast, we try to supplement the feature information by exploring the intrinsic correlation between feature maps. Although our results are not as good as those of DABNet, we still achieve good results. We also compare our method with 5 other lightweight networks. Compared with SegNet2, ENet7, and ESPNet8, our method improves by 23.2%, 13.7%, and 13.7% in accuracy, respectively. Although we have more parameters than ENet and ESPNet, our method performs well in accuracy. It is worth noting that we only use about 3.5% of the parameters of SegNet and get better results, which is a huge improvement. For ICNet15, although it produces good results, we get good results with fewer parameters.
Claims
1. A method for image segmentation based on FEFNet network structure, It is characterized in that The following steps are involved: Step S1: extract features from the image to obtain underlying feature information; Three conventional 3×3 convolutions are used to extract the underlying features of the image; Step S2: Performing skip links on the network input image to enrich feature information; skip links are to downsample the network input image and combine it with the underlying feature information obtained in step S1 to obtain enriched feature information; Step S3: down-sampling the rich feature information to obtain down-sampling module information; Step S4: Process the downsampling module information through the FEF module, perform convolution and upsampling to obtain the required image; Sub-step S41: extracting local information of the target and reducing the dimension of the feature map through a convolution operation with a convolution kernel size of 3, an output channel number that is half the number of input channels, a step size of 1, and a padding of 1; Sub-step S42: using two boundary feature extractors to extract boundary feature information of the target. Referring to the method of ERFNet, each boundary extractor adopts depthwise separable convolution and dilated convolution technology. Sub-step S43: In sub-step S42, after the operation of the first boundary extractor, its result is added to the boundary feature map and the result of sub-step S41; after the second boundary extractor, its result is dimensionally spliced with the FEF module input feature map, the boundary feature map and the result of sub-step S41; Sub-step S44: The FEF module converts the dimension into the input dimension of the FEF module through a 1x1 convolution.
2. The method for image segmentation based on the FEFNet network structure according to claim 1, It is characterized in that The step S1 comprises the following sub-steps: Sub-step S11: The first 3x3 convolution has 3 input channels, 32 output channels, 3 kernels, 2 strides, and 1 padding. Sub-step S12: After the convolution, a BN operation is performed on the result; Sub-step S13: Then the activation function is PReLu.
3. The method for image segmentation based on the FEFNet network structure according to claim 1, It is characterized in that Specifically, the skip link part of step S2 is implemented by average pooling, and the size of the average pooling is 3, the step length is 2, and the padding is 1.
4. The method for image segmentation based on the FEFNet network structure according to claim 1, It is characterized in that The step S3 comprises the following sub-steps: Sub-step S31: performing a convolution operation on the input feature map, with a convolution kernel size of 3, a step size of 2, a padding of 1, and the number of input and output channels being given parameters; Sub-step S32: Determine the input dimension and output dimension of the sampling module; if the input dimension is smaller than the output dimension, proceed to sub-steps S33, S34, and S35; if the input dimension is not smaller than the output dimension, directly perform sub-step S35 on the result of sub-step S31; Sub-step S33: performing a maximum pooling operation with a size of 2 and a step size of 2 on the input feature map; Sub-step S34: Dimensionally concatenate the results of sub-step S31 and sub-step S32; Sub-step S35: Perform BN and PReLu operations on the feature map, and the result obtained is the result of the downsampling module and is output.
5. The method for image segmentation based on the FEFNet network structure according to claim 1, It is characterized in that The step S44 is specifically as follows: The number of input channels of sub-step S44 is twice the number of channels of the input feature map of the FEF module, the number of output channels is the number of channels of the input feature map of the FEF module, the convolution kernel size is 1, the step length is 1, and the padding is 0.