Semantic Segmentation Method and Device Based on Context Cascade and Multi-Scale Feature Refinement

Through the semantic segmentation method of context cascade and multi-scale feature refinement, combined with lightweight backbone network and attention mechanism, the existing network has solved the problems of large computing overhead and slow inference speed, and achieved efficient real-time semantic segmentation effect.

CN116543155BActive Publication Date: 2025-07-11HAINAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310508273.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-08
Publication Date
2025-07-11
Estimated Expiration
2043-05-08

AI Technical Summary

Technical Problem

The existing semantic segmentation network has high computational overhead when pursuing accuracy and is difficult to meet real-time requirements, resulting in slow inference speed in practical application scenarios with limited resources.

Method used

The semantic segmentation method based on context cascading and multi-scale feature refinement is adopted. Through the backbone network, context cascading module, multi-scale feature refinement module and upsampling module, combined with the lightweight backbone network and attention mechanism, multi-scale context information is captured and spatial details are refined to improve the segmentation effect.

Benefits of technology

On the basis of maintaining the low calculation amount and parameter amount, the accuracy and inference speed of semantic segmentation are improved, and the balance of real-time and accuracy is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116543155B_ABST
    Figure CN116543155B_ABST
Patent Text Reader

Abstract

The present invention discloses a semantic segmentation method and device based on context cascading and multi-scale feature refinement. The method is applied to a convolutional neural network, and the convolutional neural network includes a semantic segmentation network based on context cascading and multi-scale feature refinement. The method includes: inputting an image into a backbone network for feature encoding; then inputting it into a context cascading module, performing a cascading operation on feature maps with different receptive fields at each level to obtain a multi-scale context information feature map with global features; inputting the feature map in the low-dimensional stage into a multi-scale feature refinement module, obtaining multi-scale spatial information in the low-dimensional stage through channel splitting and convolution, obtaining a low-dimensional multi-scale spatial feature map after attention guidance, and deeply fusing it with the multi-scale context information feature map with global features, and realizing the prediction of the feature map through upsampling. The present invention realizes a better balance between segmentation accuracy and inference speed when mining multi-scale context information and refining spatial details on resource-constrained platforms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular, to a semantic segmentation method and device based on context cascading and multi-scale feature refinement. Background Art

[0002] Currently, many popular semantic segmentation networks focus on accuracy, and these networks require a large amount of computational overhead, resulting in a very slow inference speed and making it difficult to be deployed in actual application scenarios. On the other hand, in order to pursue real-time inference speed, many works sacrifice the performance of the segmentation network. Therefore, in the field of semantic segmentation, balancing accuracy and real-time performance has become a difficult challenge.

[0003] Since the advent of the fully convolutional network FCN, it has brought semantic segmentation to a new direction. Different from previous methods, the biggest change it made is to replace the last fully connected layer of the original CNN with a convolutional layer to achieve pixel-level dense prediction, which greatly improves the segmentation accuracy. Since then, many semantic segmentation models have emerged, all adopting the FCN architecture, such as U-Net, SegNet, DeepLab series, RefineNet, PSPNet, etc. There are also some models that focus on accuracy. These above-mentioned segmentation models have achieved high accuracy on the CityScapes dataset. These semantic segmentation models all adopt large and complex backbone networks and have many computationally expensive operations. Although they can fully extract features in the image, a large number of complex computational operations also result in a very slow inference speed of the network, which cannot meet some application scenarios with real-time requirements.

[0004] In view of the above problems, in the case of poor resources, it is particularly important to meet the real-time requirements of real-time semantic segmentation and comprehensively consider aspects such as the number of parameters, computational complexity, accuracy, and inference speed in the segmentation network to achieve a high prediction accuracy with a fast inference speed. Summary of the Invention

[0005] To solve the existing technical problems, embodiments of the present invention provide a semantic segmentation method and device based on context cascading and multi-scale feature refinement. The technical solution is as follows:

[0006] In a first aspect, a processing method for image segmentation is provided, characterized in that the method is applied to a convolutional neural network, and the convolutional neural network includes a semantic segmentation network based on context cascading and multi-scale feature refinement, wherein the semantic segmentation network based on context cascading and multi-scale feature refinement further includes: a backbone network, a context cascading module, a multi-scale feature refinement module, and an upsampling module; the method includes:

[0007] After inputting the image into the backbone network, the semantic information in the image is encoded;

[0008] The feature map processed by the backbone network is input into the context cascading module, and the feature maps with different receptive fields at each level are cascaded to obtain a feature map with multi-scale context information with global features;

[0009] The feature map processed by the backbone network is input into the multi-scale feature refinement module. Through channel splitting and convolution, multi-scale spatial information in the low-dimensional stage is obtained, and after being guided by attention, a low-dimensional multi-scale spatial feature map is obtained;

[0010] The low-dimensional multi-scale spatial feature map and the feature map with multi-scale context information with global features are deeply fused in the multi-scale feature refinement module;

[0011] The deeply fused feature map is input into the upsampling module, and after upsampling, a feature map with the same size as the original image is obtained.

[0012] Further, the context cascading module includes: a dense cascading dilated convolution module with multiple different dilation rate combinations. The feature map processed by the backbone network is input into the context cascading module, and the feature maps with different receptive fields at each level are cascaded to obtain a feature map with multi-scale context information with global features, including:

[0013] After the feature map processed by the backbone network enters the context cascading module and undergoes channel compression, it is sequentially input into the dense cascading dilated convolution modules with different dilation rate combinations. In each dense cascading dilated convolution module, depthwise separable convolutions with different dilation rates and channel reduction are performed in sequence to extract multi-scale context information of different-sized targets in the feature map. After channel concatenation, it is fused with the original input feature map, and the dense cascading dilated convolution modules with different dilation rate combinations obtain their respective multi-scale feature maps with different receptive fields;

[0014] The context cascading module cascades the multi-scale feature maps with different receptive fields at each level to obtain a multi-scale context information feature map with local features;

[0015] The feature map processed by the backbone network undergoes global pooling operation after channel compression in the context cascading module, and a feature map with global feature representation is obtained through upsampling;

[0016] The context cascading module performs a cascading operation on the feature map with global feature representation and the multi-scale context information feature map with local features in a short-term dense cascading manner to obtain a feature map with multi-scale context information with global features.

[0017] Furthermore, there are three dense cascaded dilated convolution modules, and the different dilation rates of the three dense cascaded dilated convolution modules are set from small to large.

[0018] Furthermore, inputting the feature map processed by the backbone network into the multi-scale feature refinement module to obtain a low-dimensional multi-scale spatial feature map through channel splitting and convolution, including:

[0019] The feature map in the low-dimensional stage is set with four branches through channel splitting;

[0020] The four branches respectively pass through depthwise separable convolutions with different dilation rates in parallel to enrich the multi-scale spatial information in the low-dimensional stage.

[0021] Furthermore, the multi-scale spatial information in the low-dimensional stage is guided by an attention mechanism to add constraints to obtain a low-dimensional multi-scale spatial feature map guided by attention.

[0022] Furthermore, the attention mechanism adopts channel attention.

[0023] Furthermore, the multi-scale feature refinement module integrates the low-dimensional multi-scale spatial feature map guided by attention and the feature map with multi-scale context information having global features, integrates the high-dimensional features and low-dimensional features, and refines the spatial detail information.

[0024] Furthermore, the upsampling module restores the result feature map to the size of the original image.

[0025] In a second aspect, a processing device for image segmentation is provided, characterized in that the device can be applied to a convolutional neural network, the convolutional neural network includes a semantic segmentation network based on context cascading and multi-scale feature refinement, and the device includes:

[0026] A backbone network module for feature encoding of the input image to extract semantic information;

[0027] A context cascading module for channel fusion of the output feature maps with different receptive fields at all levels of the feature map processed by the backbone network module through a cascading operation to obtain a feature map with multi-scale context information having global features;

[0028] A multi-scale feature refinement module for guiding the captured low-dimensional multi-scale spatial features by high-dimensional features according to an attention mechanism, integrating the high-dimensional features and low-dimensional features of the network, and refining the spatial detail information;

[0029] An upsampling module for restoring the size of the result feature map to the size of the original input image.

[0030] Further, it is characterized in that the context cascade module further includes:

[0031] A dense cascade dilated convolution module for extracting target features of different sizes in an image according to multi-view receptive fields.

[0032] The beneficial effects brought by the technical solution provided by the embodiment of the present invention are as follows: Through a semantic segmentation method and device based on context cascade and multi-scale feature refinement, a new efficient real-time semantic segmentation network (CCMFRNet) is proposed based on a context cascade module (Context Cascade Module, CCM) and a multi-scale feature refinement module (Multi-scale Feature Refinement Module, MFRM). The CCM fuses channels in a short-term dense cascade manner using three dense cascade dilated convolution modules (Dense Cascade Dilated Convolution Module, DCDM) with different dilation rate combinations to capture rich multi-scale context information and improve the segmentation effect. The MFRM uses SE channel attention to enable high-dimensional features to guide the captured low-dimensional multi-scale spatial features, enrich the feature space in the low-dimensional stage, promote deep feature fusion between low-dimensional features and high-dimensional features, and effectively and efficiently refine spatial detail information. Both the CCM and the MFRM can effectively improve the network learning ability, making the CCMFRNet have better convergence and higher accuracy, and achieving a better balance between accuracy and efficiency. Description of the Drawings

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0034] Figure 1 is the overall structural block diagram of an image segmentation processing method provided by an embodiment of the present invention;

[0035] Figure 2 is the flowchart of a dense cascade dilated convolution module (DCDM) in an image segmentation processing method provided by an embodiment of the present invention;

[0036] Figure 3 is the flowchart of a context cascade module (CCM) in an image segmentation processing method provided by an embodiment of the present invention;

[0037] Figure 4 is the device structure diagram of an image segmentation processing method provided by an embodiment of the present invention. Detailed implementation manners

[0038] To make the objectives, technical solutions and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0039] As Figure 1 shown, the embodiment of the present invention provides a processing method for image segmentation. The method is applied to a convolutional neural network, and the convolutional neural network includes a semantic segmentation network (CCMFRNet) based on context cascade and multi-scale feature refinement. Among them, CCMFRNet is used for real-time semantic segmentation. It mainly consists of a backbone network, a context cascade module (Context Cascade Module, CCM), a multi-scale feature refinement module (Multi-scale Feature Refinement Module, MFRM), and an upsampling module. CCMFRNet aims to capture multi-scale context information and refine spatial detail information in an efficient manner to achieve an overall balance between accuracy and real-time performance. The method includes the following steps:

[0040] 1. After the image is input into the backbone network, feature encoding is performed on the input image to extract semantic information;

[0041] Feature maps with different resolutions usually have different representation capabilities, and both spatial details and semantic features are crucial for the accuracy of semantic segmentation. In this application, a lightweight backbone network is used as the encoder. During the process of backbone network feature extraction, the low-dimensional stage of the network often contains richer spatial details, and the high-dimensional stage of the network contains more semantic information. To perform feature encoding on the original input image and extract semantic information, as Figure 1As shown, first, the backbone network is empirically divided into eight stages, respectively called stage0 to stage7. Among them, stage0 to stage3 are defined as low-dimensional stages (spatial details), standard convolution operations are set, the stride is set to 2, and the output stride (OS) of the encoder for image feature extraction is set to 8. High-resolution features can retain more detailed information and spatial dimensions to extract image features. When the feature map enters the backbone network, at each level, the resolution of the feature map is reduced to 1 / 2 of the previous level. The original image resolution changes from 1 / 2, 1 / 4, until it becomes 1 / 8 of the original image resolution, that is, by continuously downsampling and convolution operations, semantic information is extracted. Stage4 to stage7 are defined as high-dimensional stages. To ensure the final segmentation effect, by introducing a dilation rate in the standard convolution, the receptive field is expanded without reducing the image resolution, and dilated convolution is performed to maintain the high resolution of the feature map. Starting from stage4, the image resolution will no longer decrease, and the image resolution is uniformly maintained at 1 / 8 of the original image resolution, and semantic extraction continues. The original input image completes feature encoding at stage7 at the tail of the backbone network.

[0042] 2. Input the feature map processed by the backbone network into the context cascade module, perform a cascade operation on the feature maps with different receptive fields at each level, and obtain a feature map with multi-scale context information with global features;

[0043] In order to capture rich multi-scale context information and improve the segmentation effect while maintaining a low computational cost and the number of parameters, a context cascade module (CCM) with a single-branch cascade structure is proposed. As Figure 3 shown, the context cascade module (CCM) includes three dense cascade dilated convolution modules (Dense Cascade Dilated Convolution Module, DCDM) with different dilation rate combinations, namely DCDM-A, DCDM-B, and DCDM-C. First, the feature map with a size of 1 / 8 of the original image resolution output by stage7 at the tail of the backbone network is denoted as There are four operations after the feature map is input into the CCM.

[0044] Specifically, the first-level operation: Compress the number of channels of the input feature map through a 1×1 standard convolution, aiming to reduce the computational cost. Multiply through the multiplication instruction and the input feature map and then add and fuse with the bias vector to prevent the model from overfitting, that is, attach a random value to it, and then follow batch normalization (Batch Normalization, BN) and the PReLU activation function to improve the accuracy at an additional negligible computational cost, and obtain the compressed feature map. Among them, the number of channels after compression is denoted as Cr , the compressed feature map is denoted as The above operation can be expressed by Equation 1:

[0045]

[0046] w in Equation 1 1×1 represents a 1×1 convolution, b represents a bias vector, represents the operations of batch normalization (BN) and activation function (PReLU).

[0047] Next, the compressed feature map enters three DCDMs (DCDM-A, DCDM-B, DCDM-C), and each DCDM has a set of dilation rates, and the set of their average dilation rates is D′ = {D a , D b , D c}, D a < D b < D c , and the dilation rates are set from small to large to obtain multi-scale context information. It should be noted that the number of input channels and output channels of the feature map passing through the DCDM is the same. The set of output feature maps at each level in the CCM is denoted as M′,

[0048] The second-level operation: The compressed feature map enters DCDM-A, which has a relatively small dilation rate and is designed to extract the features of small-sized targets. Through depthwise separable convolutions and grouped convolutions with different dilation rates, a feature map for extracting the semantic information of small-sized targets is obtained, denoted as The specific operation can be expressed by Equation 2:

[0049]

[0050] The third-level operation: The feature map enters DCDM-B, whose internal dilation rate is larger than that of DCDM-A and is designed to extract the features of medium-sized targets. Through depthwise separable convolutions and grouped convolutions with different dilation rates, a feature map for extracting the semantic information of medium-sized targets is obtained, denoted as The specific operation can be expressed by Equation 3:

[0051]

[0052] The fourth-level operation: The feature map Enter DCDM-C, whose internal dilation rate is set to be further enlarged compared with that of DCDM-B, aiming to extract the features of large-size targets. After depthwise separable convolution and grouped convolution with different dilation rates, a feature map for extracting the semantic information of large-size targets is obtained, denoted as The specific operation can be expressed by Formula 4:

[0053]

[0054] Among them, DCDM_A(*), DCDM_B(*), and DCDM_C(*) in Formulas 2, 3, and 4 respectively represent DCDM with dilation rate combinations A, B, and C.

[0055] That is, as Figure 3 shown, after the feature map enters CCM, when passing through DCDM-A, DCDM-B, and DCDM-C, the receptive field increases gradually. Through the way of short-term dense cascading, skip connections are made, and the feature maps with different receptive fields at each level are concatenated in the channel dimension. The output feature maps of DCDM-A, DCDM-B, and DCDM-C are concatenated and fused to obtain a multi-scale feature map with local features, and its number of channels is 3Cr.

[0056] Furthermore, four operations are also set inside each DCDM for the feature map. As Figure 2 shown, after the feature map compressed in channels enters DCDM-A, the dilation rate is small, so its receptive field is small. Inside DCDM-A, following the principle of gradually expanding the receptive field, four 3×3 depthwise separable convolutions with different dilation rates will be performed in sequence. The dilation rate set is D = {d1, d2, d3, d4}, where d1 < d2 < d3 < d4; at the same time, the number of channels decreases at a decreasing rate of 1 / 2. It should be noted that the last-level operation does not perform channel reduction, aiming to greatly reduce the computational amount and maintain a lightweight structure. The set of output feature maps at each level is:[[]]

[0057] That is, the first-level operation: the feature map compressed in channels when passing through the first 3×3 depthwise separable convolution, the number of channels is compressed at a decreasing rate of 1 / 2. That is, when the dilation rate is d1, the number of channels of the feature map becomes Cr / 2, and the output feature map of the first level is obtained The operation of the first level can be expressed by the following Formula 5 as:[[]]

[0058]

[0059] The second-level operation: the output feature map of the first level After the second 3×3 depth-separable convolution, the expansion rate increases compared to the previous one, and the receptive field becomes larger accordingly. That is, when the expansion rate is d2, the number of feature map channels becomes Cr / 4, and the output feature map of the second level is obtained. The operation of the second level can be expressed as follows:

[0060]

[0061] The third level operation: the output feature map of the second level After the third 3×3 depth-separable convolution, the expansion rate is enlarged again compared with the previous one, and the receptive field continues to increase. That is, when the expansion rate is d3, the number of feature map channels becomes Cr / 8, and the output feature map of the third level is obtained. The operation of the third level can be expressed by the following formula 7:

[0062]

[0063] Fourth level operation: output feature map of the third level After the fourth 3×3 depth-wise separable convolution, the expansion rate is increased compared to the previous one, and the receptive field is further enlarged, that is, the expansion rate is d4. However, in order to avoid losing too much feature information, the fourth level does not perform channel reduction operation, and the number of feature map channels remains unchanged at Cr / 8. The output feature map of the fourth level is recorded as The operation of the fourth level can be expressed as follows:

[0064]

[0065] In the above formulas The expansion rate is d z 3×3 depthwise separable convolution.

[0066] In actual application scenarios, there are objects of different sizes. If the network model learns with a single-view receptive field, it is difficult to effectively extract features for objects of different sizes. Therefore, it is necessary to use a multi-view receptive field to enable the network model to capture multi-scale contextual information and improve the prediction accuracy of the network model. First, the four-level operation in DCDM is a depth-separable convolution with different dilation rates. The dilation rate set D = {d1, d2, d3, d4}, d1 < d2 < d3 < d4. This configuration is to make the four-level operation in DCDM have a receptive field with a different perspective. Then, through a short-term dense cascade method, the output feature map of the four-level operation is channel-fused, so that DCDM can capture multi-scale contextual information and enhance information representation. That is, after the four-level operation of DCDM, the output feature map of each level obtained is channel-joined to obtain the feature map after the four-level operation cascade. The above operation can be expressed by formula 9:

[0067]

[0068] Among them, C(*) represents channel concatenation.

[0069] Since DCDM is also a single-branch cascaded structure, and CCM containing three DCDMs is also a single-branch cascaded structure, this deepens the depth of the network, which may lead to network degradation. To prevent network degradation, through another skip connection, the feature map after four-level operations and concatenation and the original input feature map of DCDM-A are added element-wise, and the multi-scale feature map with the number of channels being C r is obtained and output. The number of channels of its concatenated and fused multi-scale feature map is the same as that of the original input feature map. This enables DCDM to capture multi-scale context information and enhance information representation. The structures of DCDM_B and DCDM_C are the same as that of DCDM_A, except for the different combinations of internal dilation rates. Therefore, by the same token, no more elaboration is provided. The above operations can all be expressed by the following formula 10:

[0070]

[0071] where represents element-wise addition.

[0072] That is, as Figure 2 shown, the feature map after channel compression is input into DCDM_A, and after depthwise separable convolution and grouped convolution, the output feature map is obtained. The same applies to DCDM_B and DCDM_C.

[0073] Through the above process, as Figure 3 shown, CCM adopts a single-branch cascaded structure, cascading the outputs of different receptive fields at each level. Although it can capture multi-scale context information and enhance information representation, these context information are all local features and cannot obtain information with global feature representation, resulting in poor segmentation effect of large-size objects in the feature map. To solve this problem, the feature map after compressing the channels is subjected to global average pooling, and the pooled feature map is denoted as The specific operation is expressed by formula 11:

[0074]

[0075] The pooled feature map is restored to the size of the original feature map through upsampling operation, and the feature map with global feature representation is obtained and denoted as The specific operation can be expressed by the following formula 12:

[0076]

[0077] Among them, GAP(*) in Formulas 11 and 12 represents global average pooling, and S(*) in the formula represents upsampling.

[0078] To obtain more abundant multi-scale context information, CCM adopts a short-term dense cascade structure, which is cascaded with the output feature maps of the subsequent three-level operations, and the obtained feature map is denoted as This feature map not only has abundant multi-scale context information but also integrates information with global feature representation. The above operation can be expressed by Formula 13:

[0079]

[0080] DCDM is the core component of CCM. Both modules adopt a single-branch cascade structure, depthwise separable convolution, and grouped convolution to decompose the standard convolution. By introducing a dilation rate in the standard convolution, dilated convolution is performed to expand the receptive field without reducing the image resolution to maintain the high resolution of the feature map. The depthwise separable convolution and grouped convolution are generally used together. The depthwise separable convolution decomposes the 1×1 standard convolution into two steps. The first step is pointwise convolution, which groups the number of channels, and only the corresponding group of channels in each group performs convolution operations, which is grouped convolution. Compared with the standard convolution, it can avoid cross-channel convolution and greatly reduce the number of parameters and computational complexity. The second step is pointwise convolution, which integrates the information in the channel dimension through pointwise convolution; then, an operation of decreasing the number of channels is performed to maintain the continuity and relevance of the information. Through the above operations, CCM can obtain more abundant multi-scale context information with fewer computational resources and higher accuracy, improving the segmentation effect.

[0081] 3. Input the feature map processed by the backbone network into the multi-scale feature refinement module. Through channel splitting and convolution, low-dimensional multi-scale spatial information is obtained, and after being guided by attention, a low-dimensional multi-scale spatial feature map is obtained;

[0082] Since the multiple upsamplings of the symmetric encoder-decoder structure enlarge the size of the feature map and there are multiple feature fusion operations, it increases the computational overhead and memory occupancy, resulting in a slow inference speed. As Figure 1 shown in the black box part, first, the number of output channels of CCM is adjusted. The output feature map of CCM is compressed in channels through 1×1 convolution to reduce the computational complexity. The adjusted number of channels is denoted as C l, the specific operation is similar to Formula 1 and will not be elaborated here; mark the feature map output by the backbone network stage3 as Meanwhile, divide C l into four branches through channel splitting, and let the feature map enter the multi-branch parallel structure. These branches are depthwise separable convolutions with different dilation rates, and the set of dilation rates is denoted as R = {r1, r2, r3, r4}. The set of low-dimensional multi-scale spatial features captured by the above operations is denoted as where i is 1, 2, 3, 4. The above operations can be expressed by Formula 14:

[0083]

[0084] where, represents the depthwise separable convolution with a dilation rate of r i .

[0085] At this time, MFRM captures the low-dimensional multi-scale spatial feature map, such as Figure 1 in (a) of, thus enriching the feature space of the low-dimensional stage.

[0086] The low-dimensional stage of the network is filled with a large amount of noise information. If the low-dimensional features and high-dimensional features are directly fused, this noise information will interfere with the final prediction. Set MFRM to focus on the relationship between channels, and obtain the set of attention vectors of the CCM output feature map through SE channel attention, denoted as V l , which contains l attention vectors. Use the attention to constrain the low-dimensional multi-scale spatial feature map, so that it can autonomously select and suppress the noise information, thereby obtaining high-quality and multi-scale spatial detail information. Specifically, the Softmax function performs a weighted operation on each channel attention vector obtained above, and divides V l into four groups on average, that is Multiply V l and the low-dimensional multi-scale spatial feature set element by element according to the corresponding groups, and finally splice them according to the channel dimension to obtain the low-dimensional multi-scale spatial feature map guided by attention, denoted as F se , such as Figure 1 in (b) of. The above operations can be expressed by Formula 15:

[0087]

[0088] where represents element multiplication.

[0089] 4. Deeply fuse the attention-guided low-dimensional multi-scale spatial feature map and the feature map with multi-scale context information with global features in the multi-scale feature refinement module;

[0090] To more effectively and efficiently refine the spatial detail information, according to the above operations, continue to add the attention-guided low-dimensional multi-scale spatial feature map F se , and the output of the CCM after channel adjustment element-wise to obtain a deeply fused feature map, achieving more effective and efficient refinement of spatial detail information and improving the segmentation effect. The above operations can be expressed by Equation 16:

[0091]

[0092] The MFRM uses depthwise separable convolutions and a simple long skip connection, reducing memory occupancy and computational overhead. It can effectively refine spatial information with only less computational resources; adding an attention mechanism improves the segmentation performance and the prediction accuracy of small-sized objects; at the same time, the MFRM has a more lightweight structure and achieves considerable accuracy, enabling high-dimensional features to guide the captured low-dimensional multi-scale spatial features, enriching the feature space in the low-dimensional stage, promoting deep feature fusion between low-dimensional and high-dimensional features, achieving effective and efficient refinement of spatial detail information, and making up for the lack of spatial detail information in the high-dimensional stage.

[0093] 5. Input the deeply fused image into the upsampling module, and obtain a feature map with the same size as the original image after upsampling.

[0094] Finally, directly upsample the output deeply fused feature map to the size of the original input image through the upsampling layer (UL) to obtain a feature map with the same size as the original input image.

[0095] The present invention discloses a semantic segmentation method based on context cascading and multi-scale feature refinement. The method is applied to a convolutional neural network, which includes an efficient real-time semantic segmentation network (CCMFRNet) with context cascading and multi-scale feature refinement, and further includes: a backbone network, a context cascading module (CCM), a multi-scale feature refinement module (MFRM), and an upsampling module. The method includes: The CCM uses dense cascaded dilated convolution modules (DCDM) with 3 different dilation rate combinations to perform channel fusion in a short-term dense cascading manner, capture rich multi-scale context information, and improve the segmentation effect. The MFRM adopts SE channel attention to enable high-dimensional features to guide the captured low-dimensional multi-scale spatial features, enrich the feature space in the low-dimensional stage, promote deep feature fusion between low-dimensional and high-dimensional features, and effectively and efficiently refine spatial detail information. Both the CCM and MFRM can effectively improve the network learning ability, enable the CCMFRNet to have better convergence and higher accuracy, and achieve a better balance between accuracy and efficiency.

[0096] As Figure 4 shown, based on the same inventive concept, corresponding to the method of the present application, a device structure block diagram of a semantic segmentation method based on context cascading and multi-scale feature refinement of the present application is shown. The device can be applied to a convolutional neural network, which includes an efficient real-time semantic segmentation network (CCMFRNet) with context cascading and multi-scale feature refinement. The device includes:

[0097] A backbone network module 101, configured to perform feature encoding on an input image and extract semantic information;

[0098] A context cascading module 201, configured to perform channel fusion on the output feature maps of different receptive fields at each level through a cascading operation on the feature map processed by the backbone network module, and obtain a feature map with global feature multi-scale context information;

[0099] A multi-scale feature refinement module 301, configured to enable high-dimensional features to guide the captured low-dimensional multi-scale spatial features according to an attention mechanism, integrate the high-dimensional and low-dimensional features of the network, and refine spatial detail information;

[0100] An upsampling module 501, configured to restore the size of the result feature map to the size of the original input image.

[0101] Corresponding to the device of the present application, the efficient real-time semantic segmentation network (CCMFRNet) with context cascading and multi-scale feature refinement specifically further includes:

[0102] A dense cascaded dilated convolution module 401, configured to extract target features of different sizes in an image according to multi-perspective receptive fields.

[0103] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, an apparatus, or a computer program product. Therefore, the embodiments of the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to read-only memory, magnetic disks, or optical discs, etc.) that contain computer-usable program code.

[0104] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A processing method for image segmentation, characterized in that, The method is applied to a convolutional neural network, which includes a semantic segmentation network based on context cascading and multi-scale feature refinement. Among them, the semantic segmentation network based on context cascading and multi-scale feature refinement further includes: a backbone network, a context cascading module, a multi-scale feature refinement module, and an upsampling module; the context cascading module includes: a dense cascading dilated convolution module with multiple different dilation rate combinations; the method includes: After inputting an image into the backbone network, feature encoding is performed on the semantic information in the image; Input the feature map processed by the backbone network into the context cascading module, perform a cascading operation on the feature maps with different receptive fields at each level, and obtain a feature map with multi-scale context information with global features; Input the feature map processed by the backbone network into the multi-scale feature refinement module, obtain low-dimensional multi-scale spatial information through channel splitting and convolution, and obtain a low-dimensional multi-scale spatial feature map after being guided by attention; Deeply fuse the low-dimensional multi-scale spatial feature map and the feature map with multi-scale context information with global features in the multi-scale feature refinement module; Input the deeply fused feature map into the upsampling module, and obtain a feature map with the same size as the original image through upsampling; Among them, the step of inputting the feature map processed by the backbone network into the context cascading module, performing a cascading operation on the feature maps with different receptive fields at each level, and obtaining a feature map with multi-scale context information with global features includes: After the feature map processed by the backbone network enters the context cascading module and undergoes channel compression, it is sequentially input into the dense cascading dilated convolution modules with different dilation rate combinations. In each dense cascading dilated convolution module, depthwise separable convolutions with different dilation rates and channel reduction are sequentially performed to extract target features of different sizes in the feature map, obtain multi-scale context information of different-sized targets in the feature map, and after channel concatenation, fuse with the original input feature map. The dense cascading dilated convolution modules with different dilation rate combinations obtain their respective multi-scale feature maps with different receptive fields; The context cascading module cascades the multi-scale feature maps with different receptive fields at each level to obtain a multi-scale context information feature map with local features; Perform global pooling operation on the feature map processed by the backbone network after channel compression in the context cascading module, and obtain a feature map with global feature representation through upsampling; The context cascading module performs a cascading operation on the feature map with global feature representation and the multi-scale context information feature map with local features in a short-term dense cascading manner to obtain a feature map with multi-scale context information with global features.

2. The method according to claim 1, wherein There are three dense cascading dilated convolution modules, and the different dilation rates of the three dense cascading dilated convolution modules are set from small to large.

3. The method according to claim 1, wherein The step of inputting the feature map processed by the backbone network into the multi-scale feature refinement module, obtaining a low-dimensional multi-scale spatial feature map through channel splitting and convolution, includes: The feature map in the low-dimensional stage is set with four branches through channel splitting; The four branches respectively pass through depthwise separable convolutions with different expansion rates in parallel to enrich the multi-scale spatial information in the low-dimensional stage.

4. The method according to claim 3, characterized in that, The multi-scale spatial information in the low-dimensional stage is guided by an attention mechanism and constraints are added to obtain a low-dimensional multi-scale spatial feature map guided by attention.

5. The method according to claim 4, wherein The attention mechanism adopts channel attention.

6. The method according to claim 1 or 4, characterized in that The multi-scale feature refinement module integrates the low-dimensional multi-scale spatial feature map guided by attention and the feature map of the multi-scale context information with global features, integrates the high-dimensional features and the low-dimensional features, and refines the spatial detail information.

7. The method according to claim 1, wherein The upsampling module restores the resulting feature map to the size of the original image.

8. An image segmentation processing device, which is used to implement the method described in claim 1, characterized in that, The device can be applied to a convolutional neural network, and the convolutional neural network includes a semantic segmentation network based on context cascading and multi-scale feature refinement. The device includes: A backbone network module for feature encoding of the input image to extract semantic information; A context cascading module for channel fusion of the output feature maps with different receptive fields at each level of the feature map processed by the backbone network module through a cascading operation to obtain a feature map of multi-scale context information with global features; A multi-scale feature refinement module that, according to the attention mechanism, enables high-dimensional features to guide the captured low-dimensional multi-scale spatial features, integrates the high-dimensional and low-dimensional features of the network, and refines the spatial detail information; An upsampling module for restoring the size of the resulting feature map to the size of the original input image.

9. The device according to claim 8, characterized in that The context cascading module further includes: A dense cascading expansion convolution module for extracting target features of different sizes in the image according to multi-perspective receptive fields.

Citation Information

Patent Citations

  • Semantic segmentation method of attention mechanism based on deep learning

    CN112287940A

  • Semantic segmentation method and device based on lightweight multi-scale information fusion network

    CN114332094A