Target segmentation method and device, electronic equipment and storage medium
By introducing a multi-scale attention mechanism into the target segmentation network, the oversegment problem under limited data sets and complex background conditions is solved, and more accurate target segmentation results are achieved, improving the robustness and applicability of the segmentation network.
Patent Information
- Application Number
- CN202510397857.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-18
AI Technical Summary
Under the conditions of limited data sets and complex background, the target segmentation network is prone to oversegment, making it difficult to accurately distinguish target boundaries, resulting in a decrease in segmentation accuracy.
A multi-scale attention mechanism is introduced, through feature fusion and mask generation, and convolution operations and upsampling are performed in combination with the multi-scale attention mechanism, the basic mask is generated and linearly combined with the mask coefficient to improve segmentation accuracy.
It effectively suppresses oversegment, improves the accuracy and robustness of target segmentation, is suitable for data scarcity and complex scenarios, and provides more efficient and accurate segmentation solutions.
Smart Images

Figure CN120339610A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and particularly to a method, device, electronic device and storage medium for object segmentation. Background Art
[0002] Object detection and object segmentation are two core tasks in the field of computer vision. They aim to accurately identify and locate target objects of interest from images or videos, and at the same time perform fine boundary division on the targets. However, in some specific application scenarios, especially when the sample size of the dataset is small, the generalization ability of object detection and segmentation networks will be limited, and it is difficult to fully learn the diverse features of target objects, thus affecting the accuracy of detection and segmentation. In addition, in the face of complex backgrounds or densely arranged objects, interfering elements in the background, such as similar textures, colors or shapes, may mislead the model, resulting in false detections or missed detections.
[0003] The over-segmentation phenomenon is particularly prominent in object segmentation tasks. Over-segmentation refers to the situation where the model wrongly divides parts that originally belong to the same object into multiple independent regions, which is usually due to inaccurate recognition of the object boundary by the model. In complex backgrounds, the over-segmentation phenomenon is more serious because the model is more easily interfered by background noise, making it difficult to accurately distinguish object boundaries. Therefore, under the conditions of limited dataset and complex background, how to effectively suppress the over-segmentation phenomenon has become an urgent problem to be solved in the current field of computer vision. Summary of the Invention
[0004] In view of the above deficiencies of the prior art, the present invention provides a method, device, electronic device and storage medium for object segmentation, which effectively solves the problem of over-segmentation in object segmentation tasks under the conditions of limited dataset and complex background.
[0005] In a first aspect, the present invention provides a method for object segmentation, the method comprising:
[0006] Performing feature extraction on the image to be processed to obtain an initial feature image;
[0007] Performing feature fusion on the initial feature image to obtain a feature fusion image;
[0008] Performing convolution operation and upsampling operation according to the feature fusion image in combination with a multi-scale attention mechanism to generate a basic mask;
[0009] Generating a mask coefficient according to the feature information of the object in the feature fusion image;
[0010] Performing a linear combination of the basic mask and the mask coefficient to obtain an object segmentation result.
[0011] In an alternative embodiment, the convolution operation and the upsampling operation are performed on the feature fusion image in combination with the multi-scale attention mechanism to generate a basic mask, including:
[0012] Extract features from the feature fusion image through a first convolutional layer, and increase channel attention through the multi-scale attention mechanism to obtain a first output image;
[0013] Perform upsampling on the first output image through an upsampling layer to obtain a second output image, and the resolution of the second output image is higher than that of the first output image;
[0014] Extract features from the second output image through a second convolutional layer, and increase channel attention through the multi-scale attention mechanism to obtain the basic mask.
[0015] In an alternative embodiment, increasing the channel attention through the multi-scale attention mechanism includes:
[0016] Divide the input feature image into multiple sub-feature maps in the channel dimension direction;
[0017] Perform branch processing on the multiple sub-feature maps using parallel lines to obtain a first spatial attention map and a second spatial attention map;
[0018] Perform cross-dimensional interaction and feature aggregation on the first spatial attention map and the second spatial attention map to obtain an output feature image.
[0019] In an alternative embodiment, performing branch processing on the multiple sub-feature maps using parallel lines includes:
[0020] Encode the sub-feature maps through two global average pooling operations along the height direction and the width direction respectively to obtain two one-dimensional vectors;
[0021] Perform feature transformation and activation on the two one-dimensional vectors to obtain two activation feature maps;
[0022] Multiply the two activation feature maps with the corresponding sub-feature maps to obtain the first spatial attention map.
[0023] In an alternative embodiment, performing branch processing on the multiple sub-feature maps using parallel lines further includes:
[0024] Extract local features of the sub-feature maps through a convolutional kernel, and process the local features using batch normalization and an activation function to obtain the second spatial attention map.
[0025] In an alternative embodiment, the feature extraction of the image to be processed to obtain an initial feature image includes:
[0026] Performing shallow feature extraction on the image to be processed to obtain a first feature image;
[0027] Performing deep feature extraction on the first feature image to obtain a second feature image;
[0028] Performing multi-scale feature fusion on the second feature image to obtain the initial feature image.
[0029] In an alternative embodiment, the feature fusion of the initial feature image to obtain a feature fusion image includes:
[0030] Performing multi-scale feature extraction and sampling operations on the initial feature image to obtain an initial fusion image;
[0031] Performing feature fusion on the initial fusion image through a path aggregation network and a feature pyramid network to obtain the feature fusion image.
[0032] In a second aspect, the present invention provides an apparatus for object segmentation, the apparatus includes:
[0033] A feature extraction module, configured to perform feature extraction on an image to be processed to obtain an initial feature image;
[0034] A feature fusion module, configured to perform feature fusion on the initial feature image to obtain a feature fusion image;
[0035] An object segmentation module, configured to perform a convolution operation and an upsampling operation on the feature fusion image in combination with a multi-scale attention mechanism to generate a basic mask;
[0036] A coefficient generation module, configured to generate a mask coefficient according to the feature information of an object in the feature fusion image;
[0037] A result acquisition module, configured to perform a linear combination of the basic mask and the mask coefficient to obtain an object segmentation result.
[0038] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the object segmentation method according to any one of the first aspects of the present invention.
[0039] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the object segmentation method according to any one of the first aspects of the present invention.
[0040] The method, apparatus, electronic device, and storage medium for object segmentation provided by the present invention introduce an attention mechanism into the object segmentation network, enabling the object segmentation network to focus on different target objects according to different channels, thereby avoiding over-segmentation. By establishing the relationship between the basic mask and the object category, the segmentation accuracy is improved, and a more accurate segmentation result is obtained. At the same time, it can still achieve high-precision segmentation results in cases such as insufficient training data and non-overlapping objects, enhancing the robustness and accuracy of the object segmentation network, expanding its applicability in data-scarce and complex scenarios, and providing a more efficient and accurate solution for the object segmentation task. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0042] Figure 1 It is the first schematic diagram of the method flow for object segmentation provided by the embodiment of the present invention;
[0043] Figure 2 It is the second schematic diagram of the method flow for object segmentation provided by the embodiment of the present invention;
[0044] Figure 3 It is the third schematic diagram of the method flow for object segmentation provided by the embodiment of the present invention;
[0045] Figure 4 It is the schematic diagram of the segmentation network module structure of the YOLOv8 network model in the embodiment of the present invention;
[0046] Figure 5 It is the fourth schematic diagram of the method flow for object segmentation provided by the embodiment of the present invention;
[0047] Figure 6 It is the schematic diagram of the object segmentation result of the original YOLOv8 network model in the embodiment of the present invention;
[0048] Figure 7 It is the schematic diagram of the result of the 32-channel global MASK of the original YOLOv8 network model in the embodiment of the present invention;
[0049] Figure 8 It is the schematic diagram of the object segmentation result of the improved YOLOv8 network model in the embodiment of the present invention;
[0050] Figure 9 It is the schematic diagram of the structure of the apparatus for object segmentation provided by the embodiment of the present invention;
[0051] Figure 10 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention.
[0052] Main component symbol description:
[0053] 200, target segmentation device; 210, feature extraction module; 220, feature fusion module; 230, target segmentation module; 240, coefficient generation module; 250, result acquisition module; 300, electronic device; 310, processor; 320, communication interface; 330, memory; 340, communication bus. Detailed implementation manners
[0054] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be further described clearly and completely below with reference to the accompanying drawings in the embodiments of the present invention. It should be noted that the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0055] In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality of" means two or more, unless otherwise specifically defined.
[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0057] The over-segmentation phenomenon is a particularly prominent problem in the target segmentation task. The generation of the over-segmentation phenomenon is mainly because when the target segmentation network processes the segmentation task, it directly generates multiple full-image masks and generates the mask within the detection box of the target object according to the linear combination of multiple full-image masks. This easily causes the target segmentation network to only focus on the mask within the detection box and ignores the shape and boundary of the target object itself. Thus, when there are multiple target objects within the detection box, the target objects are wrongly divided into multiple regions, especially significantly in the case of obvious overlap.
[0058] More specifically, when the training dataset is insufficient, the segmentation ability of the target segmentation network is easily affected. As a result, when the target segmentation network generates segmentation results, it often produces unnecessary segmentations, leading to a decrease in accuracy. In a complex background, the over-segmentation phenomenon is more severe because the model is more easily interfered by background noise, making it difficult to accurately distinguish the target boundaries. Therefore, under the conditions of limited dataset and complex background, how to effectively suppress the over-segmentation phenomenon has become an urgent problem to be solved in the field of computer vision today.
[0059] Embodiment 1
[0060] The embodiment of the present invention provides a method for target segmentation, which effectively solves the problem of over-segmentation in the target segmentation task under the conditions of limited dataset and complex background. In the embodiment of the present invention, the YOLOv8 network model is used as the main network model for target segmentation. Since the segmentation module in the YOLOv8 network model adopts the strategy of detecting first and then segmenting, this strategy separates the target detection and instance segmentation tasks. Under this strategy, the YOLOv8 network model first performs target detection, identifies each target object and its bounding box in the image to be detected, and then performs a fine segmentation operation within the regions of these detected target objects, using the already located bounding boxes to accurately delimit the specific contours of the target objects, which is more efficient than the traditional end-to-end segmentation method.
[0061] For the training of the YOLOv8 network model, there is the following formula:
[0062] l i = f(mask i , gt_mask i ) / area
[0063] L seg = mean(l)
[0064] In the above formula, l i represents the loss of instance i, f represents the regional error calculation function, mask i represents the predicted segmentation region of instance i, gt_mask i represents the true segmentation region of instance i, area represents the detection region of instance i, and L seg represents the segmentation loss, and mean(l) represents the average loss of all instances.
[0065] As can be seen from the above formula, when a large amount of simple, unoccluded, and scattered data is used in the training set, during the training process of the YOLOv8 network model, since there is only one target object in the detection box, when all the foregrounds in the detection box are used as the segmentation result, the condition of low loss is satisfied, resulting in the YOLOv8 network model not having the ability to distinguish multiple target instances in the detection box. When detecting overlapping object images, over-segmentation will occur. In the embodiments of the present invention, by improving the YOLOv8 network model and adding a multi-scale attention mechanism as the target segmentation model, different channels of the target segmentation model focus on different objects, reducing the over-segmentation phenomenon.
[0066] Figure 1 It is the first schematic diagram of the method flow for target segmentation provided by the embodiments of the present invention. As Figure 1 shown, the method includes the following processes:
[0067] S100. Extract features from the image to be processed to obtain an initial feature image.
[0068] Optionally, before extracting features from the image to be processed, the image to be processed can be preprocessed, such as resizing and normalizing the image, etc., so as to meet the input requirements of the target segmentation network model.
[0069] In the embodiments of the present invention, feature extraction is performed through the backbone network of the YOLOv8 network model. Figure 2 It is the second schematic diagram of the method flow for target segmentation provided by the embodiments of the present invention. As Figure 2 shown, the feature extraction specifically includes the following steps:
[0070] S110. Perform shallow feature extraction on the image to be processed to obtain a first feature image.
[0071] The image to be processed first passes through the initial convolutional layers in the backbone network. These initial convolutional layers usually have larger convolutional kernels and more output channels, and are used to capture low-level features in the image to be processed, such as edge and texture features. After the initial convolutional layers, a downsampling operation is usually performed, such as max pooling or average pooling. The downsampling operation can reduce the spatial dimension of the feature map, thereby reducing the computational amount and increasing the receptive field of the subsequent layers, and finally obtaining the first feature image.
[0072] S120. Perform deep feature extraction on the first feature image to obtain a second feature image.
[0073] After shallow feature extraction, the first feature image enters deep convolutional layers, which use larger convolutional kernels and more output channels to further extract high-level features in the first feature image. Residual Connections and Bottleneck Structures can be adopted in the deep part of the backbone network. These structures help alleviate the vanishing gradient problem in deep networks, improving the training efficiency and performance of the network. Structures such as CSPDarknet can also be adopted in the backbone network to improve computational efficiency and enhance feature representation ability by introducing cross-stage partial connections. The CSPDarknet structure effectively extracts deep features and obtains the second feature image by splitting the first feature image into two parts for processing and merging them in subsequent stages.
[0074] S130. Perform multi-scale feature fusion on the second feature image to obtain the initial feature image.
[0075] In the deep part of the backbone network, a Feature Pyramid Network can be introduced for multi-scale feature fusion. The Feature Pyramid Network can fuse information from different levels, which helps improve the detection accuracy of small objects and the overall performance. Through the Feature Pyramid Network, the YOLOv8 network model can learn feature representations at different scales, thus better adapting to object detection of different sizes.
[0076] After deep feature extraction and multi-scale feature fusion in the backbone network of the YOLOv8 network model, a series of initial feature images are finally output. These initial feature images contain feature information at different levels in the image, providing rich feature representations for subsequent object segmentation detection.
[0077] S200. Perform feature fusion on the initial feature image to obtain the feature fusion image.
[0078] In the embodiment of the present invention, feature extraction is performed through the neck network of the YOLOv8 network model. Figure 3 It is the third schematic diagram of the method flow for object segmentation provided by the embodiment of the present invention. As Figure 3 shown, the feature fusion specifically includes the following steps:
[0079] S210. Perform multi-scale feature extraction and sampling operations on the initial feature image to obtain the initial fusion image.
[0080] The initial feature image first undergoes multi-scale feature extraction through the Spatial Pyramid Pooling (SPP) structure. The SPP structure splices feature maps of different scales together through pooling operations at different scales, thereby enhancing the detection ability of the YOLOv8 network model for targets of different sizes. At the same time, in order to achieve the fusion of feature maps of different scales, the neck network can perform upsampling or downsampling operations. The upsampling operation enlarges the size of the feature map through an interpolation algorithm, while the downsampling operation reduces the size of the feature map through convolutional layers or pooling layers.
[0081] S220. Pass the initial fusion image through a path aggregation network and a feature pyramid network for feature fusion to obtain a feature fusion image.
[0082] In the embodiment of the present invention, the initial fusion image undergoes feature fusion through a Path Aggregation Network and a Feature Pyramid Network. The neck network of the YOLOv8 network model is set with an optimized PAN-FPN structure, which has bottom-up and top-down feature fusion methods. The bottom-up path starts from the bottom-layer feature map and gradually fuses higher-level feature maps upward. This path helps to transmit bottom-layer detail information to higher-level feature maps. The top-down path starts from the highest-level feature map and gradually fuses lower-level feature maps downward. This path helps to transmit high-level semantic information to lower-level feature maps.
[0083] In the two paths of bottom-up and top-down, the feature maps are gradually fused and enhanced, forming a richer and more accurate feature representation. The feature fusion image output by the neck network contains feature information from different scales and different levels, providing a more reliable and accurate basis for subsequent target segmentation detection.
[0084] S300. Perform convolutional operations and upsampling operations on the feature fusion image in combination with a multi-scale attention mechanism to generate a basic mask.
[0085] In the embodiment of the present invention, by adding an EMA multi-scale attention mechanism, the YOLOv8 network model pays more attention to channel attention. The core idea of the EMA multi-scale attention mechanism is to combine the advantages of cross-space learning and multi-size learning to construct an efficient multi-scale attention module. Through information integration at different scales, the efficiency and accuracy of attention calculation are improved. Figure 4 It is a schematic diagram of the segmentation network module structure of the YOLOv8 network model in the embodiment of the present invention, as Figure 4As shown, after each convolutional and activation function unit in the segmentation network module, a channel attention mechanism is added to enhance the channel sensitivity of the YOLOv8 network model without disrupting the original structure.
[0086] Figure 5 It is the fourth schematic diagram of the method flow for object segmentation provided by an embodiment of the present invention. As Figure 5 shown, object segmentation specifically includes the following steps:
[0087] S310. Extract features from the feature fusion image through the first convolutional layer, and increase channel attention through the multi-scale attention mechanism to obtain the first output image.
[0088] In the embodiment of the present invention, increasing channel attention through the multi-scale attention mechanism specifically includes the following steps:
[0089] First, divide the input feature image into multiple sub-feature maps in the channel dimension direction. For any given input feature Figure X ∈R C×H×W , where C represents the number of channels, H represents the height, and W represents the width. Divide the input feature Figure X into G sub-features in the channel dimension Figure X i , denoted as X = [X0, X1,..., X G-1 , and each sub-feature map
[0090] Then, use parallel lines to perform branch processing on multiple sub-feature maps to obtain the first spatial attention map and the second spatial attention map. In the embodiment of the present invention, three parallel routes are used to process the grouped feature maps, including a 1x1 branch (two parallel sub-branches) and a 3x3 branch.
[0091] In the 1x1 branch, use two 1D global average pooling operations to encode the sub-feature map along the height direction and the width direction respectively to obtain two one-dimensional vectors. Concatenate the two one-dimensional vectors, perform feature transformation through a 1x1 convolution, apply a non-linear activation function, such as the Sigmoid function, to activate the transformed features, and obtain two activated feature maps. Multiply the activated features by the corresponding sub-feature maps to obtain the weighted first spatial attention map, thereby capturing global information and recalibrating the channel weights.
[0092] In the 3x3 branch, use a 3x3 convolution kernel to extract the local features of the sub-feature map, apply batch normalization and a non-linear activation function, such as the ReLU function, to process the convolved features, and obtain the second spatial attention map.
[0093] Finally, cross-dimensional interaction and feature aggregation are performed on the first spatial attention map and the second spatial attention map to obtain an output feature image. Specifically, the first spatial attention map and the second spatial attention map output by the 1x1 branch and the 3x3 branch are subjected to cross-dimensional interaction, and information of different dimensions is fused through operations such as matrix multiplication. At the same time, the Softmax function is used to calculate the attention weights to ensure that the sum of the weights is 1, which helps to highlight important features and suppress irrelevant features.
[0094] The first spatial attention map and the second spatial attention map are aggregated, and features are fused by element-wise addition or concatenation to obtain the final output feature image. The output feature image within each group is obtained by aggregating the two generated spatial attention weight values, which can capture pixel-level pairwise relationships and highlight the global context of all pixels.
[0095] S320. Upsample the first output image through an upsampling layer to obtain a second output image, and the resolution of the second output image is higher than that of the first output image.
[0096] S330. Extract features from the second output image through a second convolutional layer, and increase channel attention through a multi-scale attention mechanism to obtain a basic mask.
[0097] In the traditional self-attention mechanism, the computational complexity is O(N 2 ). The EMA multi-scale attention mechanism avoids the overhead of global calculation and localizes the calculation by introducing multi-scale attention, that is, calculating attention within each local region for feature maps of each scale. The calculated attention weights are used to weight the feature maps of each scale and fuse them into a unified feature, thereby reducing the complexity. Once the attention calculations for multiple scales are completed, the feature maps of these different scales are fused to obtain the final basic mask.
[0098] S400. Generate a mask coefficient according to the feature information of the object in the feature fusion image.
[0099] In the embodiment of the present invention, the YOLOv8 network model first performs object detection on the feature fusion image to identify the target objects in the feature fusion image and their position and category information. For each detected target, the YOLOv8 network model generates a blank mask with the same size as the original image. According to the results of object detection, the YOLOv8 network model fills the corresponding positions in the mask with white to represent the exact positions of the target objects in the image.
[0100] Before generating the mask, the YOLOv8 network model needs to obtain the size information of the image to be detected, including the width and height of the image. If the image to be detected has been processed such as scaled or cropped during the object detection process, the YOLOv8 network model also needs to obtain the size of the processed image. To ensure that the mask can accurately cover the corresponding area of the image to be detected, the YOLOv8 network model needs to calculate the scaling factor, which is usually obtained by comparing the size of the processed image with the size of the image to be detected, and the smaller value of the ratio of the corresponding dimensions (width or height) of the two can be taken to obtain the mask coefficient.
[0101] S500. Linearly combine the basic mask and the mask coefficient to obtain the target segmentation result.
[0102] In the embodiment of the present invention, the basic mask and the mask coefficient are linearly combined to generate the final segmentation result F fuse , that is:
[0103]
[0104] In the above formula, F fuse represents the target segmentation result, N represents the total number of basic masks, α i represents the mask coefficient, and F i represents the basic mask.
[0105] To verify the effectiveness of the target segmentation method provided in the embodiment of the present invention, the original YOLOv8 network model and the improved YOLOv8 network model in the embodiment of the present invention are used to perform target segmentation on the same image to be detected. Figure 6 is a schematic diagram of the target segmentation result of the original YOLOv8 network model in the embodiment of the present invention. As Figure 6 shown, for the two target objects on the far left of the segmentation result graph, over-segmentation occurs due to overlap. Figure 7 is a schematic diagram of the result of the original YOLOv8 network model in the embodiment of the present invention performing 32-channel global MASK. As Figure 7 shown, the original YOLOv8 network model performs 32 MASKs on the test image and linearly combines these 32 channels as the final segmentation result. These 32 channels respectively focus on different features, including features such as background and edges. Figure 8 is a schematic diagram of the target segmentation result of the improved YOLOv8 network model in the embodiment of the present invention. As Figure 8 shown, for the over-segmentation phenomenon that occurs for the two target objects on the far left of the target segmentation by the original YOLOv8 network model, the improved YOLOv8 network model does not have such a phenomenon, thus obtaining a more accurate segmentation result.
[0106] The method for target segmentation provided by the embodiments of the present invention introduces an attention mechanism into the target segmentation network, enabling the target segmentation network to focus on different target objects according to different channels, thereby avoiding the over-segmentation phenomenon. By establishing the relationship between the basic mask and the object category, the segmentation accuracy is improved, and a more accurate segmentation result is obtained.
[0107] Embodiment 2
[0108] Based on the same technical concept as in Embodiment 1 of the above method, the embodiments of the present invention provide a device for target segmentation. Figure 9 It is a schematic structural diagram of the device for target segmentation provided by the embodiments of the present invention, as Figure 9 shown. The device 200 for target segmentation includes:
[0109] A feature extraction module 210, configured to extract features from the image to be processed to obtain an initial feature image.
[0110] A feature fusion module 220, configured to perform feature fusion on the initial feature image to obtain a feature fusion image.
[0111] A target segmentation module 230, configured to perform a convolution operation and an upsampling operation according to the feature fusion image in combination with a multi-scale attention mechanism to generate a basic mask.
[0112] A coefficient generation module 240, configured to generate a mask coefficient according to the feature information of the object in the feature fusion image.
[0113] A result acquisition module 250, configured to perform a linear combination of the basic mask and the mask coefficient to obtain a target segmentation result.
[0114] The device for target segmentation provided by the embodiments of the present invention can still achieve a high-precision segmentation result in cases such as insufficient training data and non-overlapping objects, improving the robustness and accuracy of the target segmentation network, expanding its applicability in data-scarce and complex scenarios, and providing a more efficient and accurate solution for the target segmentation task.
[0115] It can be understood that the implementation manners in the method for target segmentation described in Embodiment 1 above are equally applicable to this embodiment and can achieve the same technical effects, so they will not be repeated here.
[0116] Embodiment 3
[0117] Based on the same concept, the embodiments of the present invention further provide an electronic device. Figure 10 It is a schematic structural diagram of an electronic device provided by the embodiments of the present invention, as Figure 10As shown in the figure, the electronic device 300 may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communication interface 320, and the memory 330 complete mutual communication through the communication bus 340. The processor 310 may call the logic instructions in the memory 330 to execute the steps of the method for target segmentation described in the above embodiments. For example, it includes:
[0118] S100. Extract features from the image to be processed to obtain an initial feature image;
[0119] S200. Perform feature fusion on the initial feature image to obtain a feature fusion image;
[0120] S300. Perform convolution operations and upsampling operations according to the feature fusion image in combination with a multi-scale attention mechanism to generate a basic mask;
[0121] S400. Generate a mask coefficient according to the feature information of the object in the feature fusion image;
[0122] S500. Perform a linear combination of the basic mask and the mask coefficient to obtain a target segmentation result.
[0123] Among them, the processor 310 may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or a combination of the above types of chips.
[0124] In addition, when the logical instructions in the above-mentioned memory 330 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0125] The memory 330 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created by the processor, etc. In addition, the memory may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and their combinations.
[0126] Embodiment 4
[0127] Based on the same concept, an embodiment of the present invention also provides a computer-readable storage medium. This computer-readable storage medium stores a computer program, and this computer program includes at least one segment of code. This at least one segment of code can be executed by a main control device to control the main control device to implement the steps of the method for target segmentation as described in the above-mentioned various embodiments. For example, it includes:
[0128] S100. Extract features from the image to be processed to obtain an initial feature image;
[0129] S200. Perform feature fusion on the initial feature image to obtain a feature fusion image;
[0130] S300. Perform convolution operations and upsampling operations according to the feature fusion image in combination with a multi-scale attention mechanism to generate a basic mask;
[0131] S400. Generate a mask coefficient according to the feature information of the object in the feature fusion image;
[0132] S500. Linearly combine the base mask and the mask coefficient to obtain the target segmentation result.
[0133] Based on the same technical concept, an embodiment of the present invention further provides a computer program, which, when executed by a main control device, is used to implement the above method embodiment.
[0134] The computer program can be stored in whole or in part on a computer-readable storage medium packaged together with the processor, or can be stored in whole or in part on a memory not packaged together with the processor.
[0135] Based on the same technical concept, an embodiment of the present invention further provides a processor, which is used to implement the above method embodiment. The above processor can be a chip.
[0136] In summary, the method, device, electronic device and storage medium for target segmentation provided by the present invention introduce an attention mechanism into the target segmentation network, enabling the target segmentation network to focus on different target objects according to different channels, thereby avoiding the over-segmentation phenomenon. By establishing the relationship between the base mask and the object category, the segmentation accuracy is improved, and a more accurate segmentation result is obtained. At the same time, it can still achieve a high-precision segmentation result in cases such as insufficient training data and non-overlapping objects, improving the robustness and accuracy of the target segmentation network, expanding its applicability in data-scarce and complex scenarios, and providing a more efficient and accurate solution for the target segmentation task.
[0137] Referring to "embodiment" in this article means that the specific features, structures or characteristics described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0138] The above-described embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent should be subject to the appended claims.
[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for target segmentation, characterized in that, The method includes: Performing feature extraction on the image to be processed to obtain an initial feature image; Performing feature fusion on the initial feature image to obtain a feature fusion image; Performing a convolution operation and an upsampling operation on the feature fusion image in combination with a multi-scale attention mechanism to generate a basic mask; Generating a mask coefficient according to the feature information of the object in the feature fusion image; Performing a linear combination of the basic mask and the mask coefficient to obtain a target segmentation result.
2. The method for target segmentation according to claim 1, wherein, The performing a convolution operation and an upsampling operation on the feature fusion image in combination with a multi-scale attention mechanism to generate a basic mask includes: Extracting features of the feature fusion image through a first convolutional layer and increasing channel attention through the multi-scale attention mechanism to obtain a first output image; Performing an upsampling operation on the first output image through an upsampling layer to obtain a second output image, where the resolution of the second output image is higher than that of the first output image; Extracting features of the second output image through a second convolutional layer and increasing channel attention through the multi-scale attention mechanism to obtain the basic mask.
3. The method for target segmentation according to claim 2, wherein The increasing channel attention through the multi-scale attention mechanism includes: Dividing the input feature image into multiple sub-feature maps in the channel dimension direction; Performing branch processing on the multiple sub-feature maps using parallel lines to obtain a first spatial attention map and a second spatial attention map; Performing cross-dimensional interaction and feature aggregation on the first spatial attention map and the second spatial attention map to obtain an output feature image.
4. The method for target segmentation according to claim 3, wherein The performing branch processing on the multiple sub-feature maps using parallel lines includes: Encoding the sub-feature maps through two global average pooling operations along the height direction and the width direction respectively to obtain two one-dimensional vectors; Performing feature transformation and activation on the two one-dimensional vectors to obtain two activation feature maps; Multiplying the two activation feature maps by the corresponding sub-feature maps to obtain the first spatial attention map.
5. The method for target segmentation according to claim 4, wherein The performing branch processing on the multiple sub-feature maps using parallel lines further includes: Extracting local features of the sub-feature maps through a convolution kernel and processing the local features using batch normalization and an activation function to obtain the second spatial attention map.
6. The method for target segmentation according to claim 1, wherein The performing feature extraction on the image to be processed to obtain an initial feature image includes: Performing shallow feature extraction on the image to be processed to obtain a first feature image; Performing deep feature extraction on the first feature image to obtain a second feature image; Performing multi-scale feature fusion on the second feature image to obtain the initial feature image.
7. The method for target segmentation according to claim 1, wherein The performing feature fusion on the initial feature image to obtain a feature fusion image includes: Performing multi-scale feature extraction and sampling operations on the initial feature image to obtain an initial fusion image; Performing feature fusion on the initial fusion image through a path aggregation network and a feature pyramid network to obtain the feature fusion image.
8. An apparatus for target segmentation, characterized in that, The device includes: A feature extraction module for performing feature extraction on the image to be processed to obtain an initial feature image; A feature fusion module for performing feature fusion on the initial feature image to obtain a feature fusion image; A target segmentation module, configured to perform convolution operations and upsampling operations based on the feature fusion image in combination with a multi-scale attention mechanism to generate a basic mask; A coefficient generation module, configured to generate a mask coefficient according to the feature information of an object in the feature fusion image; A result acquisition module, configured to perform a linear combination of the basic mask and the mask coefficient to obtain a target segmentation result.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the target segmentation method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the target segmentation method according to any one of claims 1 to 7.