Boundary Enhancement Network Image Semantic Segmentation Method, System, Device and Medium

By extracting multi-scale target features and performing hollow convolution feature enhancement, combined with multi-scale feature aggregation of boundary enhancement networks, the problems of feature information discontinuity and large calculation amount in the prior art are solved, and the precise segmentation of the target object is achieved.

CN120014278BActive Publication Date: 2025-06-27ZENMORN (HEFEI) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510360727.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-27
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

When the existing semantic segmentation method retains high-resolution feature maps, there are problems of feature information discontinuity and high computational cost, and it is also difficult to effectively integrate global context and local detailed information.

Method used

By acquiring the target image, the target features of multiple different scales are extracted and inputted into the progressive semantic recognition network, and the multi-scale feature enhancement is performed using hollow convolutions of different expansion rates. Then, the enhanced features are input into the boundary enhancement network, and multi-scale feature aggregation is performed to generate a target mask to achieve segmentation of the target object.

Benefits of technology

It effectively improves the model's ability to identify the target area, optimizes the boundary integrity of the target area, reduces edge blur, and ultimately achieves accurate segmentation of the target object.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014278B_ABST
    Figure CN120014278B_ABST
Patent Text Reader

Abstract

The present invention relates to a method, system, device, and medium for semantic segmentation of boundary-enhanced network images. The method includes: obtaining a target image to be segmented; inputting the target image into the backbone network of an image segmentation model to extract target features of multiple different scales from the target image; selecting at least three scales of target features, and for each selected scale of target features: inputting the target features into the progressive semantic recognition network of the image segmentation model, performing multi-scale feature enhancement on the target features based on dilated convolutions with different dilation rates to obtain dilated convolution features corresponding to the target features of this scale; inputting the dilated convolution features of each scale into the boundary enhancement network of the image segmentation model to perform multi-scale feature aggregation on the dilated convolution features of each scale to generate a target mask; and segmenting the target object in the target image based on the target mask. The present invention improves the accuracy of image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a method, system, device and medium for semantic segmentation of boundary enhancement network images. Background Art

[0002] Semantic segmentation is a fundamental problem in computer vision, aiming to assign predefined class labels to each pixel in an image according to its respective semantic category, so as to segment the image into different regions. Semantic segmentation plays a crucial role in many fields such as autonomous driving, computational photography, and human-computer interaction. Currently, most semantic segmentation methods are based on the Fully Convolutional Network (FCN). However, the continuous downsampling operations (such as pooling and strided convolution) in the FCN structure will cause a significant reduction in the spatial details of the original input image.

[0003] To alleviate the above problems, currently, it is usually an FCN model based on dilated convolution and a U-shaped network based on the encoder-decoder architecture. However, both of the above two methods have certain limitations: The FCN model based on dilated convolution reduces the downsampling operation to retain a relatively high-resolution feature map and introduces dilated convolution to expand the receptive field, thereby enhancing the feature extraction ability of dilated convolution. But this method will lead to the discontinuity of feature information and has a large computational amount. The U-shaped network based on the encoder-decoder architecture designs a decoder module on the basis of a classification network and restores the high-resolution segmentation result by gradually upsampling. Compared with the FCN model based on dilated convolution, this method benefits from the low-resolution intermediate feature map, has a lower computational amount, and consumes less memory. But this method has problems of loss of boundary details and insufficient fusion of semantic information. Therefore, it is necessary to provide a method, system, device and medium for semantic segmentation of boundary enhancement network images. Summary of the Invention

[0004] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a method, system, device and medium for semantic segmentation of boundary enhancement network images, which improves the problem of low accuracy of image segmentation in the prior art.

[0005] To achieve the above and other related objectives, the present invention provides a method for semantic segmentation of boundary-enhanced network images, including: obtaining a target image to be segmented; inputting the target image into the backbone network of the image segmentation model to extract target features of multiple different scales from the target image; wherein, the backbone network is a convolutional neural network; selecting at least three scales of target features, and for each selected scale of target features: inputting the target features into the progressive semantic recognition network of the image segmentation model, and performing multi-scale feature enhancement on the target features based on dilated convolutions with different dilation rates to obtain the dilated convolution features corresponding to the target features of this scale; wherein, the progressive semantic recognition network is a multi-scale feature learning network; inputting the dilated convolution features of each scale into the boundary enhancement network of the image segmentation model, performing multi-scale feature aggregation on each dilated convolution feature, and generating a target mask; wherein, the boundary enhancement network is a global context modeling network; segmenting the target object in the target image based on the target mask.

[0006] In an embodiment of the present invention, the step of inputting the target features into the progressive semantic recognition network of the image segmentation model, performing multi-scale feature enhancement on the target features, and obtaining the dilated convolution features corresponding to the target features of this scale includes: inputting the target features into the progressive semantic recognition network, processing the target features based on a first dilated convolution to generate a first visual feature; fusing the first visual feature and the target features to generate a first fused feature; processing the first fused feature based on a second dilated convolution to generate a second visual feature, and fusing the second visual feature with the target features to obtain a second fused feature; wherein, the dilation rate of the second dilated convolution is less than the dilation rate of the first dilated convolution; processing the second fused feature based on a third dilated convolution to generate a third visual feature, and fusing the third visual feature with the second fused feature to obtain the dilated convolution features; wherein, the dilation rate of the third dilated convolution is less than the dilation rate of the second dilated convolution.

[0007] In an embodiment of the present invention, the step of inputting the dilated convolution features of each scale into the boundary enhancement network of the image segmentation model, performing multi-scale feature aggregation on each dilated convolution feature, and generating a target mask includes: arranging the dilated convolution features of each scale in sequence, passing them through the normalization module of the boundary enhancement network to align the dilated convolution features of adjacent scales, aggregating the aligned dilated convolution features of adjacent scales, and generating enhanced features corresponding to each scale; inputting the enhanced features of each scale into the multi-scale fusion module of the boundary enhancement network, and fusing each enhanced feature based on the multi-scale feature aggregation mechanism to generate a target mask.

[0008] In one embodiment of the present invention, arranging the atrous convolution features of each scale in sequence, passing through the normalization module of the boundary enhancement network to align the atrous convolution features of adjacent scales, and aggregating the aligned atrous convolution features of adjacent scales to generate enhanced features corresponding to the corresponding scale, includes: arranging the atrous convolution features of different scales in ascending / descending order of scale, and inputting the arranged atrous convolution features of each scale into the normalization module; for the atrous convolution features of each scale, taking it as the target feature, and performing the following process: taking the scale of the target feature as the benchmark, aligning the atrous convolution features of its adjacent scales to obtain the aligned atrous convolution features; aggregating the target feature and the aligned atrous convolution features corresponding thereto to generate enhanced features corresponding to the target feature of the current scale.

[0009] In one embodiment of the present invention, inputting the enhanced features of each scale into the multi-scale parallel sampling module of the boundary enhancement network, and fusing the enhanced features based on the multi-scale feature aggregation mechanism to generate a target mask, includes: arranging the enhanced features of each scale in sequence, and inputting the enhanced features of the smallest scale into the dual-branch parallel sampling unit of the multi-scale fusion module, performing average pooling processing of multiple preset scales on the enhanced features of the smallest scale through each branch to obtain feature maps of multiple different scales corresponding to each branch; splicing the feature maps of multiple different scales obtained by each branch through the splicing unit of the multi-scale fusion module to obtain the spliced features corresponding to each branch; inputting the spliced features corresponding to each branch into the adaptive fusion unit of the multi-scale fusion module, and adaptively adjusting and fusing each spliced feature based on the attention mechanism to obtain the adaptive fusion feature; fusing the adaptive fusion feature with the enhanced features of the remaining scales through the multi-scale integration unit of the multi-scale fusion module to obtain the target mask.

[0010] In one embodiment of the present invention, segmenting the target object in the target image based on the target mask includes: aligning the target mask with the coordinate space of the target image; segmenting the target object from the target image according to the aligned target mask.

[0011] In one embodiment of the present invention, the image segmentation model is trained by minimizing the target loss, and the target loss includes: calculating the binary cross-entropy loss between the target mask and the true mask based on the difference between the target mask and the corresponding true mask at each pixel; calculating the loss between the target mask and the corresponding true mask based on the degree of regional overlap between the target mask and the true mask; obtaining the target loss based on the binary cross-entropy loss and the loss.

[0012] In one embodiment of the present invention, a boundary enhancement network image semantic segmentation system is further provided. The system includes: a data acquisition module for acquiring a target image to be segmented; a feature extraction module for inputting the target image into the backbone network of an image segmentation model to extract target features of multiple different scales from the target image, wherein the backbone network is a convolutional neural network; a feature enhancement module for selecting at least three scales of target features, and for each selected scale of target feature: inputting the target feature into the progressive semantic recognition network of the image segmentation model, and performing multi-scale feature enhancement on the target feature based on dilated convolutions with different dilation rates to obtain the dilated convolution feature corresponding to the target feature of this scale, wherein the progressive semantic recognition network is a multi-scale feature learning network; a mask generation module for inputting the dilated convolution features of each scale into the boundary enhancement network of the image segmentation model, performing multi-scale feature aggregation on each dilated convolution feature, and generating a target mask, wherein the boundary enhancement network is a global context modeling network; and a segmentation module for segmenting the target object in the target image based on the target mask.

[0013] In one embodiment of the present invention, an electronic device is further provided, including: one or more processors; a storage device for storing one or more programs, which when executed by the one or more processors, cause the electronic device to implement the boundary enhancement network image semantic segmentation method described in any one of the above.

[0014] In one embodiment of the present invention, a computer-readable storage medium is further provided, on which a computer program is stored, which when executed by a processor of a computer, causes the computer to execute the boundary enhancement network image semantic segmentation method described in any one of the above.

[0015] As described above, a boundary enhancement network image semantic segmentation method, system, device, and medium of the present invention have the following beneficial effects: extracting target features of different scales from the target image so that the model takes into account both global context information and local details. Enhancing multi-scale features through dilated convolutions with different dilation rates, enabling the features to have a larger receptive field while retaining fine edge information, thereby effectively improving the model's recognition ability for target regions. Using the global context modeling mechanism of the boundary enhancement network to optimize the boundary integrity of the target region to reduce the edge blur phenomenon. Finally, accurately segmenting the target object through the target mask. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a schematic flowchart of a boundary enhancement network image semantic segmentation method provided by an embodiment of the present invention;

[0017] Figure 2 It is the processing flowchart of the progressive semantic recognition network provided by the present invention;

[0018] Figure 3 It is the processing flowchart of the boundary enhancement network provided by the present invention;

[0019] Figure 4 It is shown as the structural block diagram of the boundary enhancement network image semantic segmentation system provided by an embodiment of the present invention;

[0020] Figure 5 It is shown as a structural schematic diagram of an electronic device according to an embodiment of the present invention. Specific embodiments

[0021] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0022] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and ratios of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0023] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0024] Compared with some backbone networks with complex structures and large computational amounts, such as ResNet50 and VGG19, MobileNetV2 is relatively weak in feature representation ability. Generally speaking, high-level features mainly carry global semantic information, which helps in the localization of target objects, while low-level features contain rich edge and texture details and can provide precise target contours. Therefore, in order to more comprehensively represent target objects, it is usually necessary to fuse multi-level features to enhance the model's perception ability of different scales and different levels of information. Specifically, 1) Low-level features and high-level features can complement each other, enhancing both the fine boundary information of the target area and the semantic information of the target area. 2) Intermediate features fuse local details and global semantics and can more accurately capture the shape and structure of target objects. Gradually extracting and fusing semantic information at different feature levels can significantly improve the detection accuracy. However, relying solely on the features extracted by the encoder layer often makes it difficult to provide sufficient information for accurate boundary detection, and boundary information is particularly crucial for semantic segmentation tasks in complex scenarios such as traffic environments. Specifically, since traditional backbone networks usually use conventional sliding windows to traverse all local regions of an image, these windows can only capture local context information. The diversity and substantial noise in these local representations hinder the recognition of semantic boundaries.

[0025] To address the above problems, the present invention provides a method for semantic segmentation of boundary-enhanced network images, which extracts target features of different scales from a target image so that the model takes into account both global context information and local details. The multi-scale features are enhanced by dilated convolutions with different dilation rates, enabling the features to have a larger receptive field while retaining fine edge information, thereby effectively improving the model's recognition ability for target regions. The boundary integrity of the target region is optimized using the global context modeling mechanism of the boundary-enhanced network to reduce the edge blurring phenomenon. Finally, precise segmentation of the target object is achieved through a target mask.

[0026] Please refer to Figure 1 , the method for semantic segmentation of boundary-enhanced network images provided by the present invention includes the following steps:

[0027] S1. Obtain a target image to be segmented.

[0028] The size of the target image is , where represents the height of the image, Let \(W\) represent the width of the image, and \(C\) represent the number of channels of the image. Exemplarily, for an RGB image, \(C = 3\). The target image contains the target object to be recognized. For different application scenarios, the corresponding target images are different. Exemplarily, for a traffic scenario, the target image can be a road environment image collected by an in-vehicle camera or a road camera, where the target objects to be segmented can be objects such as pedestrians, vehicles, road signs, and traffic lights captured by the camera. In the present invention, the target image is not limited. In order to adapt to the subsequent model processing process, it is necessary to preprocess the original target image to form a standardized target image. Among them, the preprocessing process includes, but is not limited to, adjusting the size of the target image, converting the color channels of the target image, and denoising the target image, etc. Those skilled in the art can adaptively select the corresponding preprocessing method based on the actual task requirements, which is not limited here.

[0029] S2. Input the target image into the backbone network of the image segmentation model, and extract target features of multiple different scales from the target image; wherein, the backbone network is a convolutional neural network.

[0030] Input the target image into the backbone network of the pre-trained image segmentation model. Through layer-by-layer convolution and downsampling operations, multi-level feature extraction is performed on the target image, thereby generating target features of multiple different scales. Specifically, input the target image into the backbone network, and extract large-scale target features in the shallow part of the backbone network to capture coarse-grained features such as edges and textures. As the number of network layers deepens, more attention will be paid to the fine-grained information in the target image. By extracting deeper semantic information in the target image, target features with smaller scales are gradually generated. Among them, the target features represent the relevant information of the target object to be segmented in the target image.

[0031] It can be understood that the backbone network in the present invention can be any neural network model capable of performing multi-scale feature extraction, including but not limited to related models with a feature pyramid structure (such as FPN, PSPNet) or models that can perform downsampling at different scales (such as ResNet, EfficientNet, MobileNet), etc. Optionally, in order to better balance the parameter scale and hardware consumption, the backbone network is MobileNetV2. Further, in order to be more suitable for the multi-scale feature extraction task, the global average pooling layer and the last fully connected layer are also removed from MobileNetV2 to optimize the structure of MobileNetV2. Extract the features generated in each stage from the backbone network, and the obtained multi-scale target features can be expressed as , where M is the total amount of multi-scale target features. Exemplarily, M is 5. Each stage of the backbone network includes a convolutional layer with a stride of 2, so that the feature map of each stage has a resolution half that of the previous stage to obtain multi-scale target features.

[0032] S3. Select at least three scales of target features. For each selected scale of target features: input the target features into the progressive semantic recognition network of the image segmentation model, and perform multi-scale feature enhancement on the target features based on dilated convolutions with different dilation rates to obtain the dilated convolution features corresponding to the target features of this scale; where the progressive semantic recognition network is a multi-scale feature learning network.

[0033] From the multi-scale target features extracted by the backbone network, select at least three scales of target features and perform feature enhancement through the progressive semantic recognition network. Optionally, to balance the requirements of computational complexity and segmentation accuracy, the backbone network extracts 5 different scales of target features. Except for the target features of the largest scale, the remaining target features of each scale are input into the progressive semantic recognition network for processing. The progressive semantic recognition network can be any model capable of multi-scale feature learning, including but not limited to DeepLab, HRNet, etc. Optionally, the progressive semantic recognition network is obtained through a number of dilated convolutions and residual connections. The dilated convolution can capture a larger range of context information without increasing the computational amount, and the residual connection realizes the information combination between features of different scales, so that the final feature representation is more accurate.

[0034] Specifically, S3 includes S31 to S34 (not shown in the figure):

[0035] S31. Input the target features into the progressive semantic recognition network and process the target features based on the first dilated convolution to generate the first visual features.

[0036] For each selected target feature ( ), input the target features into the progressive semantic recognition network. First, perform feature transformation through two 1×1 convolutional layers to obtain the first transformed target feature and the second target feature . The process of feature transformation is shown in formulas (1) and (2):

[0037] (1)

[0038] (2)

[0039] where is a 1×1 convolutional layer. It should be noted that the parameters of the above two 1×1 convolutional layers can be different, so as to independently transform the input target features for enriching the semantic information of the target features after transformation. By * and the dilation rate is the first dilated convolution of the target features is processed. Since the dilation rate of the first dilated convolution is large, it can capture a larger range of context information to improve the model's understanding of global features and retain local fine-grained information, thereby obtaining the first visual feature. The relevant parameters of the first dilated convolution can be adaptively set. Exemplarily, the first dilated convolution is a 3×3 dilated convolution with a dilation rate of 5.

[0040] S32. Fuse the first visual feature and the target feature to generate a first fused feature.

[0041] Fuse the first visual feature and the first target feature obtained above to obtain a first fused feature , as shown in formula (3):

[0042] (3)

[0043] where represents × the first dilated convolution with a dilation rate of . The first target feature serves as a residual connection and is element-wise added to the first visual feature obtained after the dilated convolution. The obtained first fused feature has both global context information and retains local detailed information, thereby improving the accuracy of subsequent segmentation.

[0044] S33. Process the first fused feature based on a second dilated convolution to generate a second visual feature, and fuse the second visual feature with the target feature to obtain a second fused feature; where the dilation rate of the second dilated convolution is less than that of the first dilated convolution.

[0045] Since the dilation rate of the second dilated convolution is less than that of the first dilated convolution, it has a smaller receptive field and can capture the detailed features lost in the first dilated convolution stage, enabling the model to balance information expression at different scales. Similarly, the second target feature serves as a residual connection and is element-wise added to the second visual feature obtained after the dilated convolution. The obtained second fused feature , as shown in formula (4):

[0046] (4)

[0047] Among them, represents * the second dilated convolution, and the dilation rate is . The relevant parameters of the second dilated convolution can be adaptively set. Exemplarily, the second dilated convolution is a 3×3 dilated convolution with a dilation rate of 3.

[0048] S34. Process the second fused feature based on the third dilated convolution to generate a third visual feature, and fuse the third visual feature with the second fused feature to obtain a dilated convolution feature; among them, the dilation rate of the third dilated convolution is smaller than that of the second dilated convolution.

[0049] Since the dilation rate of the third dilated convolution is smaller than that of the second dilated convolution, its receptive field is further reduced, enabling it to more precisely capture local fine-grained information to supplement the edges and subtle features that may be lost during the feature extraction process of the first two layers, thereby obtaining the third visual feature. Similarly, the second fused feature serves as a residual connection and is processed by element-wise addition with the third visual feature obtained after dilated convolution to obtain the dilated convolution feature , as shown in formula (5):

[0050] (5)

[0051] Among them, represents * the third dilated convolution, and the dilation rate is . The relevant parameters of the third dilated convolution can be adaptively set. Exemplarily, the third dilated convolution is a 3×3 dilated convolution with a dilation rate of 1.

[0052] Through the above processing process, in the multi-branch architecture, the progressive semantic recognition network can gradually adjust the receptive field, thereby realizing the feature extraction process from a large receptive field (with strong global perception ability) to a small receptive field (with the ability to finely capture local details), so that the image segmentation model can dynamically adapt to semantic information of different scales. In addition, in the multi-branch architecture, different branches are interconnected and cooperate with each other, enabling each branch to share information and complement each other, thereby enhancing the overall expression ability and segmentation accuracy of the model.

[0053] As Figure 2 shown, the present invention uses the target features extracted from the backbone network , and Taking the progressive semantic recognition network processing process as an example for illustration: the target features extracted by the backbone network , and are input into the progressive semantic recognition network, and the feature is subjected to 1×1 convolution (Conv1×1) to reduce the number of channels, and feature extraction is performed through 3×3 dilated convolutions (Conv3×3) with three different dilation rates (d = 1, d = 3, d = 5) to ensure that the feature information of different receptive fields is fully captured, and the corresponding dilated convolution features are obtained. The dilated convolution features undergo 1×1 convolution for channel dimensionality reduction, and their scales are adjusted through upsampling to align them with . undergoes 1×1 convolution for channel dimensionality reduction, and is subjected to max pooling to reduce the resolution, then undergoes 1×1 convolution for channel dimensionality reduction, the processed data are fused, and a 3×3 convolution is used to restore them to the original channels, and finally the enhanced feature is obtained. Among them, the process of obtaining the enhanced feature can be seen in the subsequent introduction.

[0054] S4. Input the dilated convolution features of each scale into the boundary enhancement network of the image segmentation model, perform multi-scale feature aggregation on each dilated convolution feature, and generate a target mask; wherein, the boundary enhancement network is a global context modeling network.

[0055] Input the dilated convolution features of each scale generated by the progressive semantic recognition network into the boundary enhancement network, perform multi-scale feature aggregation on features of different scales to fuse global and local information, thereby optimizing the edge integrity of the target region and reducing the phenomenon of missegmentation. Preferably, in order to balance the requirements of computational complexity and segmentation accuracy, in the present invention, the dilated convolution features of each obtained scale can also be sorted, and the dilated convolution features with the top three scales are selected and input into the boundary enhancement network for feature aggregation. The boundary enhancement network can be any network with global context modeling function, including but not limited to FPN, PSANet, CCNet, etc.

[0056] Specifically, S4 includes S41 and S42 (not shown in the figure):

[0057] S41. Arrange the dilated convolution features of each scale in sequence, pass them through the normalization module of the boundary enhancement network, align the dilated convolution features of adjacent scales, and aggregate the aligned dilated convolution features of adjacent scales to generate enhanced features corresponding to the scales.

[0058] Arrange the dilated convolutional features at each scale in sequence, pass them through a normalization module, and perform scale alignment processing on the dilated convolutional features at adjacent scales by upsampling or downsampling. Further aggregate the features of adjacent scales after alignment to fuse information at different scales and generate enhanced features corresponding to each scale. By aggregating the feature information of adjacent scales, gradually mine semantic information at different feature levels to accurately identify the target object.

[0059] In an optional embodiment, the enhanced features are obtained through the following steps:

[0060] First, arrange the dilated convolutional features at different scales in ascending / descending order of scale, and input the arranged dilated convolutional features at each scale into the normalization module.

[0061] Arrange the dilated convolutional features at each scale in ascending or descending order of scale, and input the arranged dilated convolutional features at each scale into the normalization module for feature alignment and aggregation processing. It can be understood that if the three dilated convolutional features at the highest scale are selected for processing, only the selected three dilated convolutions need to be arranged in ascending or descending order of scale here and jointly sent to the normalization module for subsequent processing.

[0062] Then, for each dilated convolutional feature at a certain scale, take it as the target feature and perform the following process: Based on the scale of the target feature, align the dilated convolutional features at its adjacent scales to obtain the aligned dilated convolutional features; Aggregate the target feature and the corresponding aligned dilated convolutional features to generate enhanced features corresponding to the target feature at the current scale.

[0063] For all the dilated convolutional features input to the normalization module, sequentially select each dilated convolutional feature as the target feature, and for each target feature : Take its own scale as the reference scale, and for its adjacent scales, perform upsampling or downsampling processing respectively to unify the adjacent scales to its own scale. Exemplarily, as Figure 2 shown, after arranging in descending order of scale, for the dilated convolutional feature greater than the target feature , first adjust the resolution of to be the same as through a max pooling operation, and then reduce its channels through a 1×1 convolutional layer to obtain the aligned dilated convolutional feature , that is, , where MP is max pooling. Similarly, for the dilated convolutional feature smaller than the target feature , a 1×1 convolutional layer can be used first to reduce its number of channels, and through bilinear upsampling operation, the resolution is adjusted to be the same as to obtain the aligned dilated convolutional features , that is, , where Up is upsampling. Further, for the target feature , a 1×1 convolutional layer is also used to reduce its number of channels, obtaining the target feature with reduced channels , that is, . After the above feature transformation, the target feature and the aligned dilated convolutional features at adjacent scales are aggregated to obtain the enhanced feature corresponding to the target feature. Further, considering that the channel information still needs to be optimized after multi-scale feature aggregation, the aggregated features are also passed through a 3×3 convolutional layer to adjust the number of channels, obtaining the final enhanced feature , that is, , where represents the feature concatenation operation.

[0064] It can be understood that if the target feature is the first dilated convolutional feature input to the normalization module, only the target feature and its adjacent features need to be processed as described above.

[0065] S42. Input the enhanced features at each scale into the multi-scale fusion module of the boundary enhancement network, and fuse the enhanced features based on the multi-scale feature aggregation mechanism to generate a target mask.

[0066] Input the enhanced features at each scale into the multi-scale fusion module, adopt the multi-scale feature aggregation mechanism, and fuse the feature information at different scales, so as to simultaneously capture the global context information and local details of the target object, thereby optimizing the boundary expression of the target region to generate a target mask.

[0067] In an alternative embodiment, the target mask is generated through the following process:

[0068] First, arrange the enhanced features at each scale in sequence, and input the enhanced feature at the smallest scale into the dual-branch parallel sampling unit of the multi-scale fusion module. Through each branch, perform average pooling processing on the enhanced feature at the smallest scale with multiple preset scales to obtain multiple different-scale feature maps corresponding to each branch.

[0069] In the multi-scale fusion module, the input enhanced feature at the smallest scale is divided into three branches, which are denoted as Q, K, V, where Arrange each enhanced feature in order of scale size, and take the enhanced feature with the smallest scale Input it into the dual-branch parallel sampling unit of the multi-scale fusion module. This unit contains two parallel branches, denoted as branch Q and branch K respectively. In the dual-branch parallel sampling unit, for branch Q and branch K, respectively perform average pooling processing on the input enhanced feature with the smallest scale at multiple preset scales, so that both of these branches will generate multiple feature maps of different scales for subsequent feature fusion.

[0070] Specifically, after the enhanced feature with the smallest scale undergoes multi-scale average pooling, global information can be effectively extracted to enhance the correlation between features. Compared with the max pooling that only focuses on local features, this average pooling in the present invention can capture global relationships more comprehensively, ensuring that the image segmentation model has stronger spatial perception ability at different scales. The sampling process is shown in formula (6):

[0071] (6)

[0072] where, represents the pooling scale, represents the feature map after sampling. Optionally, four different pooling scales ( ) can be selected to perform parallel sampling on branches Q and K, so that each branch can obtain 4 feature maps of different scales. These feature maps can effectively represent global spatial information, thus providing multi-scale information support for subsequent feature fusion and segmentation.

[0073] Secondly, the splicing unit of the multi-scale fusion module splices the multiple feature maps of different scales obtained by each branch to obtain the corresponding splicing features of each branch.

[0074] For the feature maps of different scales obtained by branch K and branch Q, use the splicing unit to map and reconstruct each obtained feature map to the same size to make it conform to the format of, where is the number of pixels of the feature map obtained after pooling. Concatenate these reshaped feature maps in the second dimension to form the final splicing feature, which is the output result of the splicing unit. Taking four different pooling scales as an example, the splicing features obtained by each branch are shown in formula (7):

[0075]

[0076] (7)

[0077] where, represents the splicing feature, Represents a Reshape operation to convert feature maps of different scales into a consistent size. Denotes a feature concatenation operation to fuse the features aligned at all scales along the channel dimension. Through this concatenation process, the image segmentation model can simultaneously retain the feature information of different scales to enhance the feature expression ability, thereby providing richer multi-scale information for subsequent feature fusion and global relationship modeling. It should be noted that for branch Q and branch K, their concatenated features and are both calculated as shown in formula (7). However, due to different parameters of multi-scale pooling, the resulting concatenated features and will be different.

[0078] Then, the concatenated features corresponding to each branch are input into the adaptive fusion unit of the multi-scale fusion module. Based on the attention mechanism, the concatenated features are adaptively adjusted and fused to obtain the adaptive fusion features.

[0079] Specifically, through the adaptive fusion unit, the concatenated features obtained by branch Q and branch K through the concatenation unit are multiplied, and the matrix multiplication is used to obtain the mutual relationship between the global features. The resulting global feature relationship is shown in formula (8):

[0080] (8)

[0081] Multiply the enhanced feature with the smallest scale in branch and to further optimize the feature fusion effect. The result of multiplying and is subjected to a Reshape operation to convert the result into the format to generate the adaptive fusion feature , as shown in formula (9):

[0082] = (9)

[0083] where softmax(·) is the softmax function used to calculate the normalized attention weights to enhance the feature expression of key regions, and R(·) is the Reshape operation to ensure that the final feature shape matches the requirements of subsequent tasks. Through this adaptive fusion process, features of different scales can be effectively combined, enabling the full utilization of global context information, thereby improving the accuracy of the segmentation model.

[0084] Finally, the multi-scale integration unit of the multi-scale fusion module fuses the adaptive fusion feature with the enhanced features of the remaining scales to obtain the target mask.

[0085] The feature after boundary enhancement is gradually upsampled back to the original resolution to obtain the target mask , taking the pooling of the above four different scales as an example, the obtained target mask is shown in formula (10):

[0086] (10)

[0087] where 、 and are the enhanced features that are one level, two levels, and three levels larger than the smallest scale respectively. Exemplarily, for the enhanced feature of the smallest scale being , at this time the target mask .

[0088] As Figure 3 shown, the input feature passes through the Q branch and the K branch, and generates feature maps of different scales through average pooling of various different scales to adapt to targets of different sizes. The features after each pooling are reshaped into the same format and fused through feature concatenation Cat to obtain the final multi-scale fusion feature C S, where C is the number of channels of the final multi-scale fusion feature, and S is the step dimension of the final multi-scale fusion feature. Further, after the branches Q and K pass through multi-scale fusion, their global feature relationship (C C) is calculated and normalized by Softmax. After weighted calculation by the adaptive attention mechanism, the feature of branch V is multiplied by the weight through feature fusion, and the fusion result is reshaped to obtain the final output feature .

[0089] S5. Segment the target object in the target image based on the target mask.

[0090] Using the target mask as the region guidance, perform pixel-level screening and region division on the target image. Specifically, the foreground region in the target mask corresponds to the target object to be segmented, while the background region corresponds to irrelevant information. By performing pixel-by-pixel mapping of the target mask and the original target image, the target region can be accurately extracted and the background part can be masked, thus completing the segmentation of the target object.

[0091] In an optional embodiment, S5 includes the following process:

[0092] First, align the target mask with the coordinate space of the target image.

[0093] Determine whether the target mask is consistent with the resolution of the target image. If not, perform scale adjustment on the mask using over-interpolation upsampling or downsampling to match the spatial resolution of the target image, so that the coordinates of the target mask are aligned with the coordinates of the target image.

[0094] Then, based on the aligned target mask, segment the target object from the target image.

[0095] Specifically, traverse each pixel in the target image and perform screening according to the foreground / background information of the target mask, only retaining the pixels belonging to the target object, thereby realizing the segmentation of the target object.

[0096] It can be understood that the image segmentation model is obtained through training. During the training process, the Adam optimizer is used, the momentum is set to 0.9, and the maximum number of iterations is fixed at 160000. The initial learning rate is set to 0.0001, and the batch size is set to 8. In an optional embodiment, the image segmentation model is trained by minimizing the target loss, and the target loss is obtained through the following process:

[0097] First, calculate the binary cross-entropy loss between the target mask and the corresponding ground truth mask based on the difference between them at each pixel. As shown in formula (11), calculate the binary cross-entropy loss, where is the binary cross-entropy loss between the target mask Y and the ground truth mask G:

[0098] (11)

[0099] Then, calculate the loss between the target mask and the corresponding ground truth mask based on the degree of regional overlap between them. As shown in formula (12), calculate the loss, where is the loss between the target mask Y and the ground truth mask G, represents norm:

[0100] (12)

[0101] Finally, obtain the target loss based on the binary cross-entropy loss and the loss. As shown in formula (13), calculate the target loss :

[0102] (13)

[0103] Please refer toFigure 4 , the boundary enhancement network image semantic segmentation system 100 includes: a data acquisition module 110, a feature extraction module 120, a feature enhancement module 130, a mask generation module 140, and a segmentation module 150. The data acquisition module 110 is used to acquire a target image to be segmented. The feature extraction module 120 is used to input the target image into the backbone network of the image segmentation model, and extract target features of multiple different scales from the target image; wherein, the backbone network is a convolutional neural network. The feature enhancement module 130 is used to select at least three scales of target features, and for each selected scale of target features: input the target features into the progressive semantic recognition network of the image segmentation model, and perform multi-scale feature enhancement on the target features based on dilated convolutions with different dilation rates to obtain the dilated convolution features corresponding to the target features of this scale; wherein, the progressive semantic recognition network is a multi-scale feature learning network. The mask generation module 140 is used to input the dilated convolution features of each scale into the boundary enhancement network of the image segmentation model, perform multi-scale feature aggregation on each dilated convolution feature, and generate a target mask; wherein, the boundary enhancement network is a global context modeling network. The segmentation module 150 is used to segment the target object in the target image based on the target mask.

[0104] For the specific limitations of the boundary enhancement network image semantic segmentation system, reference can be made to the limitations on the boundary enhancement network image semantic segmentation method in the above text, which will not be elaborated here. Each module in the above image segmentation system can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of the processor in the computer device in hardware format, or stored in the memory of the computer device in software format, so that the processor can call the corresponding operations of the above modules.

[0105] It should be noted that, in order to highlight the innovative part of the present invention, modules that are not closely related to solving the technical problems proposed by the present invention are not introduced in this embodiment, but this does not mean that there are no other modules in this embodiment.

[0106] Please refer to Figure 5 , the electronic device 1 may include a memory 11, a processor 12, and a bus, and may also include a computer program stored in the memory 11 and executable on the processor 12, such as a boundary enhancement network image semantic segmentation program.

[0107] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 11 can be an internal storage unit of the electronic device 1, such as the mobile hard disk of the electronic device 1. In other embodiments, the memory 11 can also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the electronic device 1. Further, the memory 11 can also include both the internal storage unit and the external storage device of the electronic device 1. The memory 11 can be used not only to store application software installed in the electronic device 1 and various types of data, such as the code for boundary-enhanced network image semantic segmentation, etc., but also to temporarily store data that has been output or will be output.

[0108] In some embodiments, the processor 12 can be composed of integrated circuits. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions packaged together, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 12 is the control core (Control Unit) of the electronic device 1, connecting various components of the entire electronic device 1 through various interfaces and circuits, and by running or executing programs or modules (such as image segmentation programs, etc.) stored in the memory 11, and calling data stored in the memory 11, to perform various functions of the electronic device 1 and process data.

[0109] The processor 12 executes the operating system of the electronic device 1 and various installed application programs. The processor 12 executes the application program to implement the steps in the above-mentioned boundary-enhanced network image semantic segmentation method.

[0110] Exemplarily, the computer program can be divided into one or more modules, and the one or more modules are stored in the memory 11 and executed by the processor 12 to complete this application. The one or more modules can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device 1. For example, the computer program can be divided into a data acquisition module 110, a feature extraction module 120, a feature enhancement module 130, a mask generation module 140, and a segmentation module 150.

[0111] The integrated unit implemented in the form of software functional modules can be stored in a computer-readable storage medium, which can be non-volatile or volatile. The above-mentioned software functional modules are stored in a storage medium and include several instructions for causing a computer device (which can be a personal computer, a computer device, or a network device, etc.) or a processor to execute part of the functions of the boundary enhancement network image semantic segmentation method described in various embodiments of the present application.

[0112] In summary, a boundary enhancement network image semantic segmentation method, system, device, and medium disclosed by the present invention extract target features of different scales from a target image, so that the model takes into account both global context information and local details. The multi-scale features are enhanced through dilated convolutions with different dilation rates, so that the features have a larger receptive field while retaining fine edge information, thereby effectively improving the model's recognition ability for target regions. The boundary integrity of the target region is optimized by using the global context modeling mechanism of the boundary enhancement network to reduce the edge blur phenomenon. Finally, accurate segmentation of the target object is achieved through the target mask. Therefore, the present invention effectively overcomes various disadvantages in the prior art and has high industrial utilization value.

[0113] The above embodiments are only illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A boundary-enhanced network image semantic segmentation method, characterized in that: The method comprises: Obtain a target image to be segmented; Inputting the target image into a backbone network of an image segmentation model, and extracting target features of multiple different scales from the target image; wherein the backbone network is a convolutional neural network; Select target features of at least three scales, and for each selected target feature of the scale: input the target feature into the progressive semantic recognition network of the image segmentation model, perform multi-scale feature enhancement on the target feature based on dilated convolutions with different dilation rates, and obtain dilated convolution features corresponding to the target feature of the scale; wherein the progressive semantic recognition network is a multi-scale feature learning network; Inputting the hole convolution features of each scale into the boundary enhancement network of the image segmentation model, performing multi-scale feature aggregation on each hole convolution feature, and generating a target mask; wherein the boundary enhancement network is a global context modeling network; Segmenting the target object in the target image based on the target mask; The step of inputting the target feature into the progressive semantic recognition network of the image segmentation model, performing multi-scale feature enhancement on the target feature based on dilated convolutions with different dilation rates, and obtaining dilated convolution features corresponding to the target feature of the scale includes: Inputting the target feature into the progressive semantic recognition network, processing the target feature based on a first dilated convolution to generate a first visual feature; Fusing the first visual feature and the target feature to generate a first fused feature; Processing the first fused feature based on a second dilated convolution to generate a second visual feature, and fusing the second visual feature with the target feature to obtain a second fused feature; wherein the expansion rate of the second dilated convolution is smaller than the expansion rate of the first dilated convolution; The second fused feature is processed based on the third dilated convolution to generate a third visual feature, and the third visual feature is fused with the second fused feature to obtain a dilated convolution feature; wherein the expansion rate of the third dilated convolution is smaller than the expansion rate of the second dilated convolution.

2. The boundary-enhanced network image semantic segmentation method according to claim 1, characterized in that: The hole convolution features of each scale are input into the boundary enhancement network of the image segmentation model, and multi-scale feature aggregation is performed on each hole convolution feature to generate a target mask, including: Arrange the hole convolution features of each scale in order, align the scales of the hole convolution features of adjacent scales through the normalization module of the boundary enhancement network, aggregate the hole convolution features of adjacent scales after alignment, and generate enhanced features of the corresponding scale; The enhanced features of each scale are input into the multi-scale fusion module of the boundary enhancement network, and the enhanced features are fused based on the multi-scale feature aggregation mechanism to generate a target mask.

3. The boundary-enhanced network image semantic segmentation method according to claim 2, characterized in that: The hole convolution features of each scale are arranged in sequence, the hole convolution features of adjacent scales are scale-aligned through the normalization module of the boundary enhancement network, and the hole convolution features of adjacent scales after alignment are aggregated to generate enhanced features of corresponding scales, including: Arranging the atrous convolution features of different scales in an order of increasing / decreasing scales, and inputting the arranged atrous convolution features of each scale into the normalization module; For each scale of the hole convolution feature, take it as the target feature and perform the following process: Taking the scale of the target feature as a reference, the scales of the dilated convolution features of adjacent scales are aligned to obtain the aligned dilated convolution features; The target feature and the corresponding aligned dilated convolutional feature are aggregated to generate an enhanced feature corresponding to the target feature of the current scale.

4. The boundary-enhanced network image semantic segmentation method according to claim 2, characterized in that: The step of inputting the enhanced features of each scale into the multi-scale fusion module of the boundary enhancement network, fusing each enhanced feature based on a multi-scale feature aggregation mechanism, and generating a target mask includes: The enhanced features of each scale are arranged in order, and the enhanced features of the minimum scale are input into the dual-branch parallel sampling unit of the multi-scale fusion module, and the enhanced features of the minimum scale are respectively subjected to average pooling processing of multiple preset scales through each branch to obtain feature maps of multiple different scales corresponding to each branch; The feature maps of various scales obtained by each branch are spliced ​​together by the splicing unit of the multi-scale fusion module to obtain splicing features corresponding to each branch; The splicing features corresponding to each branch are input into the adaptive fusion unit of the multi-scale fusion module, and based on the attention mechanism, each splicing feature is adaptively adjusted and fused to obtain an adaptive fusion feature; The adaptive fusion feature is fused with the enhanced features of the remaining scales through the multi-scale integration unit of the multi-scale fusion module to obtain a target mask.

5. The boundary-enhanced network image semantic segmentation method according to claim 1, characterized in that: The segmenting the target object in the target image based on the target mask includes: Aligning the target mask with the coordinate space of the target image; The target object is segmented from the target image according to the aligned target mask.

6. The boundary-enhanced network image semantic segmentation method according to claim 1, characterized in that: The image segmentation model is trained by minimizing the target loss, and the target loss includes: Calculating a binary cross entropy loss between the target mask and the true mask based on the difference between the target mask and the corresponding true mask at each pixel; Based on the degree of regional overlap between the target mask and the real mask, the distance between the target mask and the corresponding real mask is calculated. loss; Based on the binary cross entropy loss and the The loss obtains the target loss.

7. A boundary-enhanced network image semantic segmentation system, characterized in that: The segmentation system comprises: A data acquisition module, used for acquiring a target image to be segmented; A feature extraction module, used to input the target image into a backbone network of an image segmentation model, and extract target features of multiple different scales from the target image; wherein the backbone network is a convolutional neural network; A feature enhancement module is used to select target features of at least three scales, and for each selected target feature of a scale: input the target feature into the progressive semantic recognition network of the image segmentation model, perform multi-scale feature enhancement on the target feature based on dilated convolutions with different dilation rates, and obtain a dilated convolution feature corresponding to the target feature of the scale; wherein the progressive semantic recognition network is a multi-scale feature learning network; A mask generation module, used to input the hole convolution features of each scale into the boundary enhancement network of the image segmentation model, perform multi-scale feature aggregation on each hole convolution feature, and generate a target mask; wherein the boundary enhancement network is a global context modeling network; A segmentation module, used for segmenting the target object in the target image based on the target mask; The step of inputting the target feature into the progressive semantic recognition network of the image segmentation model, performing multi-scale feature enhancement on the target feature based on dilated convolutions with different dilation rates, and obtaining dilated convolution features corresponding to the target feature of the scale includes: Inputting the target feature into the progressive semantic recognition network, processing the target feature based on a first dilated convolution to generate a first visual feature; Fusing the first visual feature and the target feature to generate a first fused feature; Processing the first fused feature based on a second dilated convolution to generate a second visual feature, and fusing the second visual feature with the target feature to obtain a second fused feature; wherein the expansion rate of the second dilated convolution is smaller than the expansion rate of the first dilated convolution; The second fused feature is processed based on the third dilated convolution to generate a third visual feature, and the third visual feature is fused with the second fused feature to obtain a dilated convolution feature; wherein the expansion rate of the third dilated convolution is smaller than the expansion rate of the second dilated convolution.

8. An electronic device, characterized in that: The electronic device comprises: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the electronic device to implement the boundary-enhanced network image semantic segmentation method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor of a computer, the computer is caused to execute the boundary-enhanced network image semantic segmentation method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Icon semantic segmentation method and system based on double-model prediction and automatic screening mechanism and storage medium

    CN118968072A