Boundary enhanced network image semantic segmentation method, system, equipment and medium

By extracting multi-scale target features and performing hollow convolution feature enhancement, combined with multi-scale feature aggregation of boundary enhancement networks, the problems of feature information discontinuity and boundary details loss in the prior art are solved, and precise segmentation of the target object is achieved.

CN120014278AActive Publication Date: 2025-05-16ZENMORN (HEFEI) TECH CO LTD

Patent Information

Application Number
CN202510360727.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-05-16
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

The existing semantic segmentation methods have dilemmas between retaining high-resolution feature maps and reducing the amount of computation, resulting in discontinuity of feature information and loss of boundary details.

Method used

By acquiring the target image, extracting the target features of multiple different scales, and inputting them into the progressive semantic recognition network for multi-scale feature enhancement, using the boundary enhancement network for multi-scale feature aggregation, and generating the target mask to achieve accurate segmentation.

Benefits of technology

It effectively improves the model's ability to identify the target area, optimizes the boundary integrity of the target area, reduces edge blur, and ultimately achieves accurate segmentation of the target object.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014278A_ABST
    Figure CN120014278A_ABST
Patent Text Reader

Abstract

The invention relates to a boundary enhanced network image semantic segmentation method, system and device and a medium. The method comprises the following steps: acquiring a to-be-segmented target image; inputting the target image into a backbone network of an image segmentation model, and extracting multiple target features of different scales from the target image; selecting target features of at least three scales, and for the selected target features of each scale, inputting the target features into a progressive semantic recognition network of an image segmentation model, and performing multi-scale feature enhancement on the target features based on cavity convolution of different expansion rates to obtain cavity convolution features corresponding to the target features of the scale; inputting the dilated convolution features of each scale into a boundary enhancement network of an image segmentation model, and performing multi-scale feature aggregation on each dilated convolution feature to generate a target mask; and segmenting the target object in the target image based on the target mask. According to the invention, the accuracy of image segmentation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a method, system, device and medium for semantic segmentation of boundary-enhanced network images. Background Art

[0002] Semantic segmentation is a fundamental problem in computer vision, which aims to divide an image into different regions by assigning a predefined class label to each pixel in the image according to its respective semantic category. Semantic segmentation plays a vital role in many fields such as autonomous driving, computational photography, and human-computer interaction. Currently, most semantic segmentation methods are based on fully convolutional networks (FCN). However, the continuous downsampling operations in the FCN structure (such as pooling and strided convolution) will cause a significant reduction in the spatial details of the original input image.

[0003] In order to alleviate the above problems, the current methods are usually based on the FCN model of dilated convolution and the U-shaped network based on the encoder-decoder architecture. However, the above two methods have certain limitations: the FCN model based on dilated convolution reduces the downsampling operation to retain the relatively high-resolution feature map, and introduces dilated convolution to expand the receptive field, thereby reducing the feature extraction capability of the hollow convolution. However, this method will lead to discontinuity of feature information and has a large amount of calculation. The U-shaped network based on the encoder-decoder architecture is a decoder module designed on the basis of the classification network, which restores the high-resolution segmentation results by gradually upsampling. Compared with the FCN model based on dilated convolution, this method benefits from the low-resolution intermediate feature map, has lower computational complexity and less memory consumption. However, this method has the problem of loss of boundary details and insufficient fusion of semantic information. Therefore, it is necessary to provide a boundary enhancement network image semantic segmentation method, system, device and medium. Summary of the invention

[0004] In view of the above-mentioned shortcomings of the prior art, an object of the present invention is to provide a method, system, device and medium for semantic segmentation of boundary-enhanced network images, which improves the problem of low accuracy of image segmentation in the prior art.

[0005] To achieve the above-mentioned purpose and other related purposes, the present invention provides a boundary-enhanced network image semantic segmentation method, comprising: obtaining a target image to be segmented; inputting the target image into a backbone network of an image segmentation model, and extracting target features of multiple scales from the target image; wherein the backbone network is a convolutional neural network; selecting target features of at least three scales, and for each selected target feature of a scale: inputting the target feature into a progressive semantic recognition network of the image segmentation model, performing multi-scale feature enhancement on the target feature based on hole convolutions of different expansion rates, and obtaining hole convolution features corresponding to the target feature of the scale; wherein the progressive semantic recognition network is a multi-scale feature learning network; inputting hole convolution features of each scale into a boundary enhancement network of the image segmentation model, performing multi-scale feature aggregation on each hole convolution feature, and generating a target mask; wherein the boundary enhancement network is a global context modeling network; and segmenting the target object in the target image based on the target mask.

[0006] In one embodiment of the present invention, the target feature is input into the progressive semantic recognition network of the image segmentation model, multi-scale feature enhancement is performed on the target feature, and a hole convolution feature corresponding to the target feature of the scale is obtained, including: inputting the target feature into the progressive semantic recognition network, processing the target feature based on a first hole convolution to generate a first visual feature; fusing the first visual feature and the target feature to generate a first fused feature; processing the first fused feature based on a second hole convolution to generate a second visual feature, and fusing the second visual feature with the target feature to obtain a second fused feature; wherein the expansion rate of the second hole convolution is less than the expansion rate of the first hole convolution; processing the second fused feature based on a third hole convolution to generate a third visual feature, and fusing the third visual feature with the second fused feature to obtain a hole convolution feature; wherein the expansion rate of the third hole convolution is less than the expansion rate of the second hole convolution.

[0007] In one embodiment of the present invention, the hole convolution features of each scale are input into the boundary enhancement network of the image segmentation model, and multi-scale feature aggregation is performed on each hole convolution feature to generate a target mask, including: arranging the hole convolution features of each scale in sequence, aligning the scales of the hole convolution features of adjacent scales through the normalization module of the boundary enhancement network, aggregating the aligned hole convolution features of adjacent scales, and generating enhanced features of the corresponding scale; inputting the enhanced features of each scale into the multi-scale fusion module of the boundary enhancement network, fusing each enhanced feature based on the multi-scale feature aggregation mechanism, and generating a target mask.

[0008] In one embodiment of the present invention, the hole convolution features of each scale are arranged in sequence, the hole convolution features of adjacent scales are scale-aligned through the normalization module of the boundary enhancement network, the hole convolution features of adjacent scales after alignment are aggregated, and the enhanced features of the corresponding scale are generated, including: arranging the hole convolution features of different scales in the order of increasing / decreasing scales, and inputting the arranged hole convolution features of each scale into the normalization module; for each hole convolution feature of each scale, taking it as the target feature, performing the following process: taking the scale of the target feature as a reference, aligning the hole convolution features of the adjacent scales to obtain the aligned hole convolution features; aggregating the target feature and the corresponding aligned hole convolution feature to generate an enhanced feature corresponding to the target feature of the current scale.

[0009] In one embodiment of the present invention, the enhanced features of each scale are input into the multi-scale parallel sampling module of the boundary enhancement network, and the enhanced features are fused based on the multi-scale feature aggregation mechanism to generate a target mask, including: arranging the enhanced features of each scale in sequence, and inputting the enhanced features of the minimum scale into the dual-branch parallel sampling unit of the multi-scale fusion module, performing average pooling processing of multiple preset scales on the enhanced features of the minimum scale through each branch, and obtaining feature maps of multiple different scales corresponding to each branch; splicing the feature maps of multiple different scales obtained by each branch through the splicing unit of the multi-scale fusion module to obtain the splicing features corresponding to each branch; inputting the splicing features corresponding to each branch into the adaptive fusion unit of the multi-scale fusion module, and adaptively adjusting and fusion of each splicing feature based on the attention mechanism to obtain an adaptive fusion feature; fusing the adaptive fusion feature with the enhanced features of the remaining scales through the multi-scale integration unit of the multi-scale fusion module to obtain a target mask.

[0010] In one embodiment of the present invention, segmenting the target object in the target image based on the target mask includes: aligning the target mask with the coordinate space of the target image; and segmenting the target object from the target image according to the aligned target mask.

[0011] In one embodiment of the present invention, the image segmentation model is obtained by training by minimizing the target loss, wherein the target loss includes: calculating the binary cross entropy loss between the target mask and the real mask based on the difference between each pixel of the target mask and the corresponding real mask; calculating the cross entropy loss between the target mask and the corresponding real mask based on the degree of regional overlap between the target mask and the real mask; loss; based on the binary cross entropy loss and the The loss obtains the target loss.

[0012] In one embodiment of the present invention, a boundary enhancement network image semantic segmentation system is also provided, the system comprising: a data acquisition module, used to acquire a target image to be segmented; a feature extraction module, used to input the target image into a backbone network of an image segmentation model, and extract target features of multiple scales from the target image; wherein the backbone network is a convolutional neural network; a feature enhancement module, used to select target features of at least three scales, and for each selected target feature of a scale: input the target feature into a progressive semantic recognition network of the image segmentation model, perform multi-scale feature enhancement on the target feature based on hole convolutions with different expansion rates, and obtain hole convolution features corresponding to the target feature of the scale; wherein the progressive semantic recognition network is a multi-scale feature learning network; a mask generation module, used to input hole convolution features of each scale into a boundary enhancement network of the image segmentation model, perform multi-scale feature aggregation on each hole convolution feature, and generate a target mask; wherein the boundary enhancement network is a global context modeling network; and a segmentation module, used to segment the target object in the target image based on the target mask.

[0013] In one embodiment of the present invention, an electronic device is also provided, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the electronic device implements any of the above-mentioned boundary-enhanced network image semantic segmentation methods.

[0014] In one embodiment of the present invention, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed by a processor of a computer, the computer executes any of the above-mentioned boundary-enhanced network image semantic segmentation methods.

[0015] As described above, a boundary-enhanced network image semantic segmentation method, system, device and medium of the present invention have the following beneficial effects: extracting target features of different scales from the target image so that the model takes into account both global context information and local details. Multi-scale features are enhanced by dilated convolutions with different expansion rates so that the features have a larger receptive field while retaining fine edge information, thereby effectively improving the model's ability to recognize the target area. The boundary integrity of the target area is optimized using the global context modeling mechanism of the boundary-enhanced network to reduce edge blurring. Finally, accurate segmentation of the target object is achieved through the target mask. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A schematic flow chart of a method for semantic segmentation of a network image with enhanced boundary provided by an embodiment of the present invention; Figure 2A processing flow chart of the progressive semantic recognition network provided by the present invention; Figure 3 A processing flow chart of the boundary enhancement network provided by the present invention; Figure 4 Shown is a structural block diagram of a boundary-enhanced network image semantic segmentation system provided by an embodiment of the present invention; Figure 5 Shown is a structural schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0017] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0018] It should be noted that the illustrations provided in the following embodiments are only used to illustrate the basic concept of the present invention in a schematic manner, and thus the illustrations only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.

[0019] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0020] Compared with some backbone networks with complex structures and large computational complexity (such as ResNet50 and VGG19), MobileNetV2 is relatively weak in feature representation capabilities. Generally speaking, high-level features mainly carry global semantic information and help locate the target object, while low-level features contain rich edge and texture details and can provide accurate target contours. Therefore, in order to more comprehensively characterize the target object, it is usually necessary to fuse multi-level features to enhance the model's perception of information at different scales and levels. Specifically, 1) low-level features and high-level features can complement each other, both enhancing the fine boundary information of the target area and enriching the semantic information of the target area. 2) Intermediate features combine local details with global semantics, and can more accurately capture the morphology and structure of the target object. Gradually extracting and fusing semantic information at different feature levels can significantly improve detection accuracy. However, relying solely on features extracted by the encoder layer often fails to provide enough information to achieve accurate boundary detection, and boundary information is particularly critical for semantic segmentation tasks in complex scenes (such as traffic environments). Specifically, since traditional backbone networks usually use regular sliding windows to traverse all local regions of an image, these windows can only capture local contextual information. The diversity and substantial noise in these local representations hinder the identification of semantic boundaries.

[0021] In view of the above problems, the present invention provides a boundary-enhanced network image semantic segmentation method, which extracts target features of different scales from the target image so that the model can take into account both global context information and local details. Multi-scale features are enhanced by dilated convolutions with different expansion rates, so that the features have a larger receptive field while retaining fine edge information, thereby effectively improving the model's recognition ability of the target area. The boundary integrity of the target area is optimized by using the global context modeling mechanism of the boundary-enhanced network to reduce edge blurring. Finally, the target object is accurately segmented through the target mask.

[0022] See also Figure 1 The boundary-enhanced network image semantic segmentation method provided by the present invention comprises the following steps: S1. Obtain a target image to be segmented.

[0023] The target image size is ,in Indicates the height of the image. Represents the width of the image, C represents the number of channels of the image, and illustratively, for RGB images, C is 3. The target image contains the target object to be identified, and the corresponding target image is different for different application scenarios. Exemplarily, for a traffic scene, the target image can be a road environment image captured by a vehicle-mounted camera or a road camera, wherein the target object to be segmented can be objects such as pedestrians, vehicles, road signs, and traffic lights captured by the camera. The target image is not limited in the present invention. In order to adapt to the subsequent model processing process, it is necessary to preprocess the original target image to form a standardized target image, wherein the preprocessing process includes but is not limited to adjusting the size of the target image, converting the color channels of the target image, and denoising the target image, etc. Those skilled in the art can adaptively select the corresponding preprocessing method based on the actual task needs, which is not limited here.

[0024] S2. Input the target image into a backbone network of an image segmentation model, and extract target features of multiple scales from the target image; wherein the backbone network is a convolutional neural network.

[0025] The target image is input into the backbone network of the pre-trained image segmentation model, and multi-level feature extraction is performed on the target image through layer-by-layer convolution and downsampling operations, thereby generating target features of various scales. Specifically, the target image is input into the backbone network, and large-scale target features are extracted in the shallow part of the backbone network to capture coarse-grained features such as edges and textures. As the number of network layers increases, more attention will be paid to the fine-grained information in the target image, and by extracting deeper semantic information in the target image, smaller-scale target features will be gradually generated. Among them, the target feature represents the relevant information of the target object to be segmented in the target image.

[0026] It is understandable that the backbone network in the present invention can be any neural network model capable of multi-scale feature extraction, including but not limited to related models with a feature pyramid structure (such as FPN, PSPNet) or models that can perform downsampling at different scales (such as ResNet, EfficientNet, MobileNet), etc. Optionally, in order to better strike a balance between parameter scale and hardware consumption, the backbone network is MobileNetV2. Furthermore, in order to be more suitable for multi-scale feature extraction tasks, the global average pooling layer and the last fully connected layer are removed from MobileNetV2 to achieve structural optimization of MobileNetV2. The features generated at each stage of the backbone network are extracted, and the obtained multi-scale target features can be expressed as , where M is the total amount of multi-scale target features, and illustratively, M is 5. Each stage of the backbone network contains a convolutional layer with a stride of 2, so that the resolution of the feature map of each stage is reduced by half compared to the previous stage to obtain multi-scale target features.

[0027] S3. Select target features of at least three scales, and for each selected target feature of the scale: input the target feature into the progressive semantic recognition network of the image segmentation model, perform multi-scale feature enhancement on the target feature based on dilated convolutions with different dilation rates, and obtain dilated convolution features corresponding to the target feature of the scale; wherein the progressive semantic recognition network is a multi-scale feature learning network.

[0028] From the multi-scale target features extracted by the backbone network, select target features of at least three scales, and perform feature enhancement through the progressive semantic recognition network. Optionally, in order to take into account the requirements of both computational complexity and segmentation accuracy, the backbone network extracts target features of 5 different scales. Except for the target features of the largest scale, the target features of the remaining scales are input into the progressive semantic recognition network for processing. The progressive semantic recognition network can be any model capable of multi-scale feature learning, including but not limited to DeepLab, HRNet, etc. Optionally, the progressive semantic recognition network is obtained through a number of dilated convolutions and residual connections. The dilated convolution can capture a wider range of contextual information without increasing the amount of computation, and the residual connection can be used to combine information between features of different scales, so that the final feature expression is more accurate.

[0029] Specifically, S3 includes S31 to S34 (not shown in the figure): S31. Input the target feature into the progressive semantic recognition network, process the target feature based on the first dilated convolution, and generate a first visual feature.

[0030] For each target feature selected ( ), the target feature Input into the progressive semantic recognition network, first pass through two 1 × 1 convolutional layers for feature conversion, and obtain the first target feature after conversion and the second target feature , the process of feature conversion is shown in formulas (1) and (2): (1) (2) in, It should be noted that the parameters of the above two 1×1 convolutional layers can be different, so that the input target features can be Perform independent transformation to enrich the semantic information of the transformed target features. * And the expansion rate is The first dilated convolution on the target feature Due to the large expansion rate of the first dilated convolution, a wider range of contextual information can be captured to improve the model's understanding of global features and retain local fine-grained information, thereby obtaining the first visual feature. The relevant parameters of the first dilated convolution can be adaptively set. For example, the first dilated convolution is a 3×3 dilated convolution with a dilation rate of 5.

[0031] S32: Fusing the first visual feature and the target feature to generate a first fused feature.

[0032] The first visual feature and the first target feature obtained above are Fusion is performed to obtain the first fusion feature , as shown in formula (3): (3) in, express × The first dilated convolution is . First target feature As a residual connection, it is element-by-element added with the first visual feature obtained after the dilated convolution. The obtained first fused feature has both global context information and local detail information, thereby improving the accuracy of subsequent segmentation.

[0033] S33. Process the first fused feature based on a second dilated convolution to generate a second visual feature, and fuse the second visual feature with the target feature to obtain a second fused feature; wherein the expansion rate of the second dilated convolution is smaller than the expansion rate of the first dilated convolution.

[0034] Since the expansion rate of the second dilated convolution is smaller than that of the first dilated convolution, it has a smaller receptive field and can capture the detailed features lost in the first dilated convolution stage, so that the model can take into account the information expression of different scales. As a residual connection, the second visual feature obtained after the dilated convolution is added element by element to obtain the second fusion feature. , as shown in formula (4): (4) in, express * The second hole convolution with a dilation rate of The relevant parameters of the second dilated convolution can be adaptively set. For example, the second dilated convolution is a 3×3 dilated convolution, and the dilation rate is 3.

[0035] S34. Process the second fused feature based on the third dilated convolution to generate a third visual feature, and fuse the third visual feature with the second fused feature to obtain a dilated convolution feature; wherein the expansion rate of the third dilated convolution is smaller than the expansion rate of the second dilated convolution.

[0036] Since the expansion rate of the third dilated convolution is smaller than that of the second dilated convolution, its receptive field is further reduced, enabling it to more accurately capture local fine-grained information to supplement the edges and subtle features that may be lost in the feature extraction process of the first two layers, thereby obtaining the third visual feature. As a residual connection, it is element-wise added with the third visual feature obtained after the dilated convolution to obtain the dilated convolution feature , as shown in formula (5): (5) in, express * The third hole convolution with a dilation rate of The relevant parameters of the third dilated convolution can be adaptively set. Exemplarily, the third dilated convolution is a 3×3 dilated convolution, and the dilation rate is 1.

[0037] Through the above processing, under the multi-branch architecture, the progressive semantic recognition network can gradually adjust the receptive field, thereby realizing the feature extraction process from a large receptive field (with strong global perception ability) to a small receptive field (with the ability to finely capture local details), so that the image segmentation model can dynamically adapt to semantic information of different scales. In addition, in the multi-branch architecture, different branches are interconnected and synergistic, so that each branch can share information and complement each other, thereby improving the overall expression ability and segmentation accuracy of the model.

[0038] like Figure 2 As shown, the present invention extracts target features from the backbone network , and Take the progressive semantic recognition network processing as an example to illustrate: the target features extracted by the backbone network , and Input to the progressive semantic recognition network, the features A 1×1 convolution (Conv1×1) is performed to reduce the number of channels, and feature extraction is performed through 3×3 dilated convolutions (Conv3×3) with three different dilation rates (d=1, d=3, d=5) to ensure that the feature information of different receptive fields is fully captured and the corresponding dilated convolution features are obtained. . Dilated convolution features The channel dimension is reduced by 1×1 convolution, and its scale is adjusted by upsampling to make it consistent with Align, After a 1×1 convolution, channel dimension reduction is performed and Perform maximum pooling to reduce the resolution, then perform channel dimension reduction through a 1×1 convolution, fuse the processed data, and restore it to the original channel through a 3×3 convolution, and finally obtain enhanced features , where the process of obtaining enhanced features is described in the subsequent section.

[0039] S4. Input the hole convolution features of each scale into the boundary enhancement network of the image segmentation model, perform multi-scale feature aggregation on each hole convolution feature, and generate a target mask; wherein the boundary enhancement network is a global context modeling network.

[0040] The hole convolution features of each scale generated by the progressive semantic recognition network are input into the boundary enhancement network, and multi-scale feature aggregation is performed on features of different scales to fuse global and local information, thereby optimizing the edge integrity of the target area and reducing mis-segmentation. Preferably, in order to take into account the requirements of both computational complexity and segmentation accuracy, the hole convolution features of each scale obtained in the present invention can also be sorted, and the hole convolution features with the top three scales are selected and input into the boundary enhancement network for feature aggregation. The boundary enhancement network can be any network with global context modeling function, including but not limited to FPN, PSANet, CCNet, etc.

[0041] Specifically, S4 includes S41 and S42 (not shown in the figure): S41, arranging the hole convolution features of each scale in order, aligning the scales of the hole convolution features of adjacent scales through the normalization module of the boundary enhancement network, aggregating the aligned hole convolution features of adjacent scales, and generating enhanced features of the corresponding scale.

[0042] The hole convolution features of each scale are arranged in order, and the hole convolution features of adjacent scales are aligned by upsampling or downsampling through the normalization module. The aligned adjacent scale features are further aggregated to fuse information of different scales and generate enhanced features of the corresponding scale. By aggregating the feature information of adjacent scales, semantic information is gradually mined at different feature levels to accurately identify the target object.

[0043] In an optional embodiment, the enhanced features are obtained by the following steps: First, the hole convolution features of different scales are arranged in the order of increasing / decreasing scales, and the arranged hole convolution features of each scale are input into the normalization module.

[0044] The hole convolution features of each scale are arranged in order of increasing or decreasing scale, and the arranged hole convolution features of each scale are input into the normalization module for feature alignment and aggregation. It can be understood that if the three hole convolution features of the highest scale are selected for processing, then only the three selected hole convolutions need to be arranged in order of increasing or decreasing scale, and sent to the normalization module for subsequent processing.

[0045] Then, for each scale of the dilated convolution feature, take it as the target feature and perform the following process: take the scale of the target feature as the benchmark, align the scales of the dilated convolution features of the adjacent scales to obtain the aligned dilated convolution features; aggregate the target feature and the corresponding aligned dilated convolution features to generate an enhanced feature corresponding to the target feature of the current scale.

[0046] For all the hole convolution features input to the normalization module, each hole convolution feature is selected as the target feature in sequence. : Take its own scale as the reference scale, and perform upsampling or downsampling on the adjacent scales to unify the adjacent scales to its own scale. For example, Figure 2 As shown, after arranging in descending order of scale, for features larger than the target The dilated convolution feature , we can first use the maximum pooling operation to The resolution is adjusted to Similarly, a 1×1 convolution layer is used to reduce the channels and obtain the aligned hole convolution features. ,Right now , where MP is the maximum pooling. Similarly, for features smaller than the target The dilated convolution feature , we can first reduce its channels through a 1×1 convolution layer and then perform a bilinear upsampling operation. The resolution is adjusted to Same, get the aligned hole convolution features ,Right now , where Up is upsampling. Further, for the target feature , and also reduce its channels through a 1×1 convolution layer to obtain the target features after reducing the channels ,Right now After the above feature conversion, the target feature The aligned hole convolution features of the adjacent scales are aggregated to obtain the enhanced features corresponding to the target features. Furthermore, considering that channel information still needs to be optimized after multi-scale feature aggregation, the aggregated features are also passed through a 3×3 convolution layer to adjust the number of channels to obtain the final enhanced features. ,Right now ,in Represents a feature concatenation operation.

[0047] It is understandable that if the target feature is the first hole convolution feature input to the normalization module, then only the target feature and its adjacent features Just carry out the above processing.

[0048] S42, inputting the enhanced features of each scale into the multi-scale fusion module of the boundary enhancement network, fusing each enhanced feature based on the multi-scale feature aggregation mechanism, and generating a target mask.

[0049] The enhanced features of each scale are input into the multi-scale fusion module, and a multi-scale feature aggregation mechanism is adopted to fuse feature information of different scales so as to capture the global context information and local details of the target object at the same time, thereby optimizing the boundary expression of the target area to generate a target mask.

[0050] In an optional embodiment, the target mask is generated by the following process: First, the enhanced features of each scale are arranged in sequence, and the enhanced features of the minimum scale are input into the dual-branch parallel sampling unit of the multi-scale fusion module. The enhanced features of the minimum scale are subjected to average pooling processing of multiple preset scales through each branch to obtain feature maps of multiple different scales corresponding to each branch.

[0051] In the multi-scale fusion module, the minimum scale of the input enhanced features It is divided into three branches, which are denoted as Q, K, and V. Arrange the enhanced features in order according to their scale, and put the smallest scale enhanced features The dual-branch parallel sampling unit is input to the multi-scale fusion module. The unit contains two parallel branches, which are denoted as branch Q and branch K. In the dual-branch parallel sampling unit, for branch Q and branch K, the minimum scale of the input is enhanced. Average pooling processing of multiple preset scales is performed so that both branches can generate multiple feature maps of different scales for subsequent feature fusion.

[0052] Specifically, the smallest scale enhancement feature After multi-scale average pooling, global information can be effectively extracted to enhance the correlation between features. Compared with the maximum pooling that only focuses on local features, the average pooling of the present invention can capture global relationships more comprehensively and ensure that the image segmentation model has stronger spatial perception capabilities at different scales. The sampling process is shown in formula (6): (6) in, represents the pooling scale, represents the sampled feature map. Optionally, four different pooling scales can be selected ( ) The branches Q and K are sampled in parallel, so that each branch can obtain four feature maps of different scales. These feature maps can effectively represent the global spatial information, thereby providing multi-scale information support for subsequent feature fusion and segmentation.

[0053] Secondly, the feature maps of various scales obtained by each branch are spliced ​​together through the splicing unit of the multi-scale fusion module to obtain the splicing features corresponding to each branch.

[0054] For branch K and branch Q, feature maps of different scales are obtained, and each feature map is mapped and reconstructed to the same size using the splicing unit to make it conform to The format is is the number of pixels of the feature map obtained after pooling. These reshaped feature maps are concatenated in the second dimension to form the final splicing feature as the output result of the splicing unit. Taking four different pooling scales as an example, the splicing features obtained by each branch are shown in formula (7):

[0055] (7) in, represents the splicing feature, Reshape represents a reshape operation to convert feature maps of different scales into a consistent size. represents the feature concatenation operation to fuse all scale-aligned features in the channel dimension. Through this concatenation process, the image segmentation model can simultaneously retain feature information of different scales to enhance feature expression capabilities, thereby providing richer multi-scale information for subsequent feature fusion and global relationship modeling. It should be noted that for branches Q and K, their concatenated features are and The calculation method of is as shown in formula (7), but due to the different parameters of multi-scale pooling, the obtained splicing feature and There will be differences.

[0056] Then, the splicing features corresponding to each branch are input into the adaptive fusion unit of the multi-scale fusion module, and based on the attention mechanism, each splicing feature is adaptively adjusted and fused to obtain an adaptive fusion feature.

[0057] Specifically, the splicing features of branch Q and branch K obtained by the splicing unit are multiplied by the adaptive fusion unit, and the relationship between the global features is obtained by matrix multiplication. The global feature relationship is obtained As shown in formula (8): (8) Branch The smallest scale enhancement feature in and Multiply to further optimize the feature fusion effect. and The result of the multiplication is reshaped to convert the result into Format to generate adaptive fusion features , as shown in formula (9): = (9) Among them, softmax(·) is the softmax function, which is used to calculate the normalized attention weight to enhance the feature expression of the key area, and R(·) is the reshaping operation to ensure that the final feature shape matches the requirements of subsequent tasks. Through this adaptive fusion process, features of different scales can be effectively combined, so that the global context information can be fully utilized, thereby improving the accuracy of the segmentation model.

[0058] Finally, the adaptive fusion feature is fused with the enhanced features of the remaining scales through the multi-scale integration unit of the multi-scale fusion module to obtain the target mask.

[0059] The features after boundary enhancement are gradually upsampled to the original resolution to obtain the target mask. , taking the above four poolings of different scales as an example, the target mask obtained is shown in formula (10): (10) in, , and are enhanced features one level, two levels, and three levels larger than the minimum scale, respectively. For example, the enhanced features of the minimum scale are When the target mask .

[0060] like Figure 3 As shown, the input features After the Q branch and the K branch, through average pooling of multiple scales, feature maps of different scales are generated to adapt to objects of different sizes. The features after each pooling are converted to the same format through reshaping and fused through feature splicing Cat to obtain the final multi-scale fusion feature C S, where C is the number of channels of the final multi-scale fusion feature and S is the step size of the final multi-scale fusion feature. Furthermore, after multi-scale fusion of branches Q and K, their global feature relationship (C C) and normalized by Softmax. After weighted calculation by the adaptive attention mechanism, the features of branch V are multiplied by weights through feature fusion, and the fusion result is reshaped to obtain the final output feature .

[0061] S5. Segment the target object in the target image based on the target mask.

[0062] The target mask is used as a region guide to perform pixel-level screening and region segmentation on the target image. Specifically, the foreground area in the target mask corresponds to the target object to be segmented, while the background area corresponds to irrelevant information. By mapping the target mask with the original target image pixel by pixel, the target area can be accurately extracted and the background part can be shielded, thereby completing the segmentation of the target object.

[0063] In an optional embodiment, S5 includes the following process: First, the target mask is aligned with the coordinate space of the target image.

[0064] It is determined whether the resolution of the target mask is consistent with that of the target image. If not, the mask is resized by over-interpolation upsampling or downsampling to match the spatial resolution of the target image, so that the coordinates of the target mask are aligned with those of the target image.

[0065] Then, the target object is segmented from the target image according to the aligned target mask.

[0066] Specifically, each pixel in the target image is traversed and filtered according to the foreground / background information of the target mask, and only the pixels belonging to the target object are retained, thereby achieving the segmentation of the target object.

[0067] It can be understood that the image segmentation model is obtained by training. During the training process, the Adam optimizer is used, the momentum is set to 0.9, and the maximum number of iterations is fixed to 160000. The initial learning rate is set to 0.0001 and the batch size is set to 8. In an optional embodiment, the image segmentation model is obtained by training by minimizing the target loss, and the target loss is obtained by the following process: First, based on the difference between the target mask and the corresponding true mask at each pixel, the binary cross entropy loss between the target mask and the true mask is calculated. As shown in formula (11), the binary cross entropy loss is calculated, where: is the binary cross entropy loss between the target mask Y and the true mask G: (11) Then, based on the degree of regional overlap between the target mask and the real mask, the distance between the target mask and the corresponding real mask is calculated. As shown in formula (12), calculate Losses, of which, is the distance between the target mask Y and the real mask G loss, express Norm: (12) Finally, based on the binary cross entropy loss and the The target loss is obtained by calculating the target loss. As shown in formula (13), the target loss is calculated : (13) See also Figure 4The boundary enhancement network image semantic segmentation system 100 includes: a data acquisition module 110, a feature extraction module 120, a feature enhancement module 130, a mask generation module 140 and a segmentation module 150. The data acquisition module 110 is used to acquire a target image to be segmented. The feature extraction module 120 is used to input the target image into the backbone network of the image segmentation model, and extract target features of multiple scales from the target image; wherein the backbone network is a convolutional neural network. The feature enhancement module 130 is used to select target features of at least three scales, and for each selected target feature of a scale: the target feature is input into the progressive semantic recognition network of the image segmentation model, and multi-scale feature enhancement is performed on the target feature based on dilated convolutions with different dilation rates to obtain dilated convolution features corresponding to the target feature of the scale; wherein the progressive semantic recognition network is a multi-scale feature learning network. The mask generation module 140 is used to input dilated convolution features of each scale into the boundary enhancement network of the image segmentation model, and multi-scale feature aggregation is performed on each dilated convolution feature to generate a target mask; wherein the boundary enhancement network is a global context modeling network. The segmentation module 150 is used to segment the target object in the target image based on the target mask.

[0068] For the specific limitations of the boundary-enhanced network image semantic segmentation system, please refer to the limitations of the boundary-enhanced network image semantic segmentation method above, which will not be repeated here. Each module in the above-mentioned image segmentation system can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware format, or can be stored in the memory of the computer device in software format, so that the processor can call the operations corresponding to the above modules.

[0069] It should be noted that, in order to highlight the innovative part of the present invention, the present embodiment does not introduce modules that are not closely related to solving the technical problem proposed by the present invention, but this does not mean that there are no other modules in the present embodiment.

[0070] See also Figure 5 The electronic device 1 may include a memory 11, a processor 12 and a bus, and may also include a computer program stored in the memory 11 and executable on the processor 12, such as a boundary-enhanced network image semantic segmentation program.

[0071] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a mobile hard disk, a multimedia card, a card-type memory (for example, SD or DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device 1, such as a mobile hard disk of the electronic device 1. In other embodiments, the memory 11 may also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device 1. Further, the memory 11 may also include both an internal storage unit of the electronic device 1 and an external storage device. The memory 11 can not only be used to store application software and various types of data installed in the electronic device 1, such as codes for semantic segmentation of boundary-enhanced network images, etc., but also can be used to temporarily store data that has been output or is to be output.

[0072] In some embodiments, the processor 12 may be composed of an integrated circuit, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and a combination of various control chips. The processor 12 is the control core (Control Unit) of the electronic device 1, and uses various interfaces and lines to connect various components of the entire electronic device 1, and executes various functions and processes data of the electronic device 1 by running or executing programs or modules (such as image segmentation programs, etc.) stored in the memory 11, and calling data stored in the memory 11.

[0073] The processor 12 executes the operating system and various installed applications of the electronic device 1. The processor 12 executes the applications to implement the steps in the above-mentioned boundary-enhanced network image semantic segmentation method.

[0074] Exemplarily, the computer program may be divided into one or more modules, which are stored in the memory 11 and executed by the processor 12 to complete the present application. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device 1. For example, the computer program may be divided into a data acquisition module 110, a feature extraction module 120, a feature enhancement module 130, a mask generation module 140, and a segmentation module 150.

[0075] The above-mentioned integrated unit implemented in the form of a software function module can be stored in a computer-readable storage medium, and the computer-readable storage medium can be non-volatile or volatile. The above-mentioned software function module is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a computer device, or a network device, etc.) or a processor to perform part of the functions of the boundary-enhanced network image semantic segmentation method described in each embodiment of the present application.

[0076] In summary, the present invention discloses a method, system, device and medium for semantic segmentation of boundary-enhanced network images, which extract target features of different scales from the target image so that the model can take into account both global context information and local details. Multi-scale features are enhanced by dilated convolutions with different expansion rates, so that the features have a larger receptive field while retaining fine edge information, thereby effectively improving the model's recognition ability for the target area. The boundary integrity of the target area is optimized by using the global context modeling mechanism of the boundary-enhanced network to reduce edge blurring. Ultimately, accurate segmentation of the target object is achieved through the target mask. Therefore, the present invention effectively overcomes the various shortcomings of the prior art and has a high industrial utilization value.

[0077] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Anyone familiar with the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by a person of ordinary skill in the art without departing from the spirit and technical concept disclosed by the present invention shall still be covered by the claims of the present invention.

Claims

1. A boundary-enhanced network image semantic segmentation method, characterized in that: The method comprises: Obtain a target image to be segmented; Inputting the target image into a backbone network of an image segmentation model, and extracting target features of multiple different scales from the target image; wherein the backbone network is a convolutional neural network; Select target features of at least three scales, and for each selected target feature of the scale: input the target feature into the progressive semantic recognition network of the image segmentation model, perform multi-scale feature enhancement on the target feature based on dilated convolutions with different dilation rates, and obtain dilated convolution features corresponding to the target feature of the scale; wherein the progressive semantic recognition network is a multi-scale feature learning network; Inputting the hole convolution features of each scale into the boundary enhancement network of the image segmentation model, performing multi-scale feature aggregation on each hole convolution feature, and generating a target mask; wherein the boundary enhancement network is a global context modeling network; The target object in the target image is segmented based on the target mask.

2. The boundary-enhanced network image semantic segmentation method according to claim 1, characterized in that: The target feature is input into the progressive semantic recognition network of the image segmentation model, and multi-scale feature enhancement is performed on the target feature based on the dilated convolution with different dilation rates to obtain the dilated convolution feature corresponding to the target feature of the scale, including: Inputting the target feature into the progressive semantic recognition network, processing the target feature based on a first dilated convolution to generate a first visual feature; Fusing the first visual feature and the target feature to generate a first fused feature; Processing the first fused feature based on a second dilated convolution to generate a second visual feature, and fusing the second visual feature with the target feature to obtain a second fused feature; wherein the expansion rate of the second dilated convolution is smaller than the expansion rate of the first dilated convolution; The second fused feature is processed based on the third dilated convolution to generate a third visual feature, and the third visual feature is fused with the second fused feature to obtain a dilated convolution feature; wherein the expansion rate of the third dilated convolution is smaller than the expansion rate of the second dilated convolution.

3. The boundary-enhanced network image semantic segmentation method according to claim 1, characterized in that: The hole convolution features of each scale are input into the boundary enhancement network of the image segmentation model, and multi-scale feature aggregation is performed on each hole convolution feature to generate a target mask, including: Arrange the hole convolution features of each scale in order, align the scales of the hole convolution features of adjacent scales through the normalization module of the boundary enhancement network, aggregate the hole convolution features of adjacent scales after alignment, and generate enhanced features of the corresponding scale; The enhanced features of each scale are input into the multi-scale fusion module of the boundary enhancement network, and the enhanced features are fused based on the multi-scale feature aggregation mechanism to generate a target mask.

4. The boundary-enhanced network image semantic segmentation method according to claim 3, characterized in that: The hole convolution features of each scale are arranged in sequence, the hole convolution features of adjacent scales are scale-aligned through the normalization module of the boundary enhancement network, and the hole convolution features of adjacent scales after alignment are aggregated to generate enhanced features of corresponding scales, including: Arranging the atrous convolution features of different scales in an order of increasing / decreasing scales, and inputting the arranged atrous convolution features of each scale into the normalization module; For each scale of the hole convolution feature, take it as the target feature and perform the following process: Taking the scale of the target feature as a reference, the scales of the dilated convolution features of adjacent scales are aligned to obtain the aligned dilated convolution features; The target feature and the corresponding aligned dilated convolutional feature are aggregated to generate an enhanced feature corresponding to the target feature of the current scale.

5. The boundary-enhanced network image semantic segmentation method according to claim 3, characterized in that: The step of inputting the enhanced features of each scale into the multi-scale fusion module of the boundary enhancement network, fusing each enhanced feature based on a multi-scale feature aggregation mechanism, and generating a target mask includes: The enhanced features of each scale are arranged in order, and the enhanced features of the minimum scale are input into the dual-branch parallel sampling unit of the multi-scale fusion module, and the enhanced features of the minimum scale are respectively subjected to average pooling processing of multiple preset scales through each branch to obtain feature maps of multiple different scales corresponding to each branch; The feature maps of various scales obtained by each branch are spliced ​​together by the splicing unit of the multi-scale fusion module to obtain splicing features corresponding to each branch; The splicing features corresponding to each branch are input into the adaptive fusion unit of the multi-scale fusion module, and based on the attention mechanism, each splicing feature is adaptively adjusted and fused to obtain an adaptive fusion feature; The adaptive fusion feature is fused with the enhanced features of the remaining scales through the multi-scale integration unit of the multi-scale fusion module to obtain a target mask.

6. The boundary-enhanced network image semantic segmentation method according to claim 1, characterized in that: The segmenting the target object in the target image based on the target mask includes: Aligning the target mask with the coordinate space of the target image; The target object is segmented from the target image according to the aligned target mask.

7. The boundary-enhanced network image semantic segmentation method according to claim 1, characterized in that: The image segmentation model is trained by minimizing the target loss, and the target loss includes: Calculating a binary cross entropy loss between the target mask and the true mask based on the difference between the target mask and the corresponding true mask at each pixel; Based on the degree of regional overlap between the target mask and the real mask, the distance between the target mask and the corresponding real mask is calculated. loss; Based on the binary cross entropy loss and the The loss obtains the target loss.

8. A boundary-enhanced network image semantic segmentation system, characterized in that: The segmentation system comprises: A data acquisition module, used for acquiring a target image to be segmented; A feature extraction module, used to input the target image into a backbone network of an image segmentation model, and extract target features of multiple different scales from the target image; wherein the backbone network is a convolutional neural network; A feature enhancement module is used to select target features of at least three scales, and for each selected target feature of a scale: input the target feature into the progressive semantic recognition network of the image segmentation model, perform multi-scale feature enhancement on the target feature based on dilated convolutions with different dilation rates, and obtain a dilated convolution feature corresponding to the target feature of the scale; wherein the progressive semantic recognition network is a multi-scale feature learning network; A mask generation module, used to input the hole convolution features of each scale into the boundary enhancement network of the image segmentation model, perform multi-scale feature aggregation on each hole convolution feature, and generate a target mask; wherein the boundary enhancement network is a global context modeling network; A segmentation module is used to segment the target object in the target image based on the target mask.

9. An electronic device, characterized in that: The electronic device comprises: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the electronic device to implement the boundary-enhanced network image semantic segmentation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor of a computer, the computer is enabled to execute the boundary-enhanced network image semantic segmentation method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Three-dimensional heart image segmentation method based on multi-scale edge perception

    CN114663445A

  • Blood vessel image segmentation method and system based on three-dimensional deep network

    CN115546570A

  • SAR image oil spill detection method

    CN115984697A

  • Remote sensing image semantic segmentation method for guiding multi-information fusion based on boundary information

    CN116797792A

  • PCB small target defect detection method based on improved YOLOv8s

    CN118587188A

Cited By

  • Target segmentation method for boundary learning optimization

    CN120672783A

  • Anaphora image segmentation method and system based on edge enhancement and bottleneck vector

    CN120852767A

  • Feature alignment system and training method of large-small model suitable for target detection

    CN122435246A