Image instance segmentation method and device

By combining the feature fusion method of visible light images and binary edge images, the problem of segmentation incompleteness in the prior art when dealing with discontinuous and irregular targets is solved, and a more accurate and robust image instance segmentation effect is achieved.

CN119941762AActive Publication Date: 2025-05-06SHENZHEN UNIV

Patent Information

Application Number
CN202311448102.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-02
Publication Date
2025-05-06
Estimated Expiration
2043-11-02

AI Technical Summary

Technical Problem

When the prior art deals with discontinuous and irregular edges, it is easy to cause incomplete segmentation or missed detection, which limits the accurate acquisition of edge features and affects the overall segmentation effect.

Method used

A method of image instance segmentation is proposed. By acquiring visible light images and corresponding binary edge images, a feature extraction network and binary edge image guidance module are used to perform feature fusion, and combined with a joint attention module and a prediction module, an instance segmentation of visible light images is realized.

Benefits of technology

It effectively improves the network's perception of boundary features, reduces the impact of noise in the image, improves the accuracy and robustness of segmentation, and alleviates the problem of poor segmentation effect when multiple object targets overlap.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941762A_ABST
    Figure CN119941762A_ABST
Patent Text Reader

Abstract

The invention provides an image instance segmentation method and device. The method comprises the following steps: acquiring a visible light image to be subjected to object segmentation and a binary edge image corresponding to the visible light image; extracting a second image feature of the binary edge image through convolution operation; fusing the first image feature and the second image feature to obtain a third image feature; the third image features are input into a joint attention module, and the joint attention module is used for adding space attention and coordinate attention to the third image features to form fourth image features; and inputting the fourth image features into a prediction module, performing semantic prediction and mask prediction on the fourth image features by the prediction module, and outputting a segmentation scheme of the object in the visible light image. According to the invention, the perception capability of the network to the boundary features can be effectively improved, and the instance segmentation performance of the network is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition, and in particular to a method and device for image instance segmentation. Background Art

[0002] Different types of images contain different feature information. In recent years, computer vision tasks that use one image to guide another image have become increasingly widespread. Among them, a binary edge image refers to the edge of the target in a binary image. It has a certain spatial continuity and is the dividing line between the target and the background. The edge part is an important source of information about the object, which can determine the characteristics of the image. Therefore, using binary edge images to guide object segmentation in visible light images has become an effective method to improve performance. By using binary edge images, more accurate edge features can be obtained, thereby improving the performance of object segmentation. In instance segmentation tasks, the edges of the target need to be accurately extracted. Using binary edge images to guide object segmentation can reduce the impact of noise in the image to a certain extent and improve the accuracy of segmentation. At the same time, edge structures and characteristics can also help the model better understand the shape and structure of the object, thereby improving the robustness of segmentation.

[0003] However, the disadvantage of the prior art is that when processing targets with discontinuous or irregular edges, incomplete segmentation or missed detection may occur, which limits the accurate acquisition of edge features in some practical applications, thereby affecting the overall segmentation effect. Summary of the invention

[0004] In order to solve the above technical problems, the present invention proposes a method and device for image instance segmentation, which are used to solve the technical problem of the limitations of the prior art in processing targets with discontinuous edges and irregular shapes.

[0005] According to a first aspect of the present invention, a method for image instance segmentation is provided, the method comprising the following steps:

[0006] Step S1: obtaining a visible light image to be segmented and a binary edge image corresponding to the visible light image; extracting a first image feature of the visible light image by a feature extraction network, wherein the first image feature is a multi-layer feature map having num levels of image features; extracting a second image feature of the binary edge image by a convolution operation;

[0007] Step S2: inputting the first image feature and the second image feature into a binary edge image guiding module, wherein the binary edge image guiding module comprises a top-down fusion path and a bottom-down fusion path, and is used to fuse the first image feature and the second image feature to obtain a third image feature;

[0008] Step S3: inputting the third image feature into a joint attention module, wherein the joint attention module is used to add spatial attention and coordinate attention to the third image feature to form a fourth image feature;

[0009] Step S4: inputting the fourth image feature into a prediction module, the prediction module performs semantic prediction and mask prediction on the fourth image feature, and outputs a segmentation scheme for the object in the visible light image.

[0010] Preferably, the step S2 comprises:

[0011] Select the hierarchical features of level 2 to level num from the first image features in descending order of size, input the hierarchical features of level 2 to level num and the second image features into the binary edge image guiding module, and perform the binary edge image guiding on the hierarchical feature C of level num. i Perform 1×1 convolution to get P num , according to the order of i decreasing from num-1 to 2, for P i+1 Up-sample and then perform the i-th level feature C i Perform 1×1 convolution and transform P i+1 The upsampling result is similar to C i The result of 1×1 convolution is added according to the corresponding elements of the pixel position to get P i ; Fuse P2 with the second image feature to obtain the fused image feature M2; in the order of i increasing from 2 to num-1, i Downsampling, for P i+1 Perform 1×1 convolution and transform M i The downsampling result is similar to P i+1 The result of 1×1 convolution is added according to the corresponding elements of the pixel position to get M i+1 ; For M2 to M num Do 3×3 convolution to get N2 to N num ; for N num Do coordinate convolution and get N num+1 ; Among them, N2 to N num+1 is the output of the binary edge image guiding module, that is, the third image feature; num is a natural number.

[0012] Preferably, in step S4, the prediction module includes a semantic category branch, a mask convolution kernel branch and a mask feature branch, the semantic category branch is used to perform an interpolation convolution operation on the input fourth image feature to obtain the probability that the content corresponding to each grid unit belongs to the target category; the mask convolution kernel branch is parallel to the semantic category branch, and is used to perform a dynamic convolution calculation operation on the input fourth image feature to obtain a prediction mask; the mask feature branch is used to perform convolution, batch normalization nonlinear activation and adaptive interpolation operation on the input fourth image feature to obtain a mask feature map, and based on the mask feature map, a segmentation scheme for the object in the visible light image is obtained.

[0013] Preferably, non-maximum suppression is performed on the output of the prediction module, including: obtaining all prediction scores output by the prediction module, subjecting the top N prediction scores with the highest scores to non-maximum suppression processing to obtain an IoU matrix of size N×N; performing a maximum value operation on the column direction of the IoU matrix, and determining the mask to be retained based on the maximum value of each column of the IoU matrix; based on the prediction score of each prediction, calculating the penalty of the prediction for other predictions through a linear function, and then fitting the prediction with the largest overlap as the suppression probability, calculating the attenuation factor using the penalty and suppression probability of each prediction, and obtaining the final prediction based on the attenuation factor.

[0014] Preferably, in the training process of the visible light image instance segmentation network model guided by binary edge images, a positive and negative sample label assignment strategy based on minimum loss matching is introduced, and the positive and negative sample label assignment strategy based on minimum loss matching is:

[0015] Determine the level outputs N2 to N3 corresponding to the third image feature output by the binary edge image guiding module. num +1, determine the object size range corresponding to the output of each level, and the object size ranges corresponding to each level are different and have no intersection; when an object has multiple possible sizes, each possible size corresponds to a level, and the outputs of the multiple levels corresponding to the object are predicted respectively; in each corresponding level, by calculating the classification and segmentation losses of each grid in the true target box and the true target, find the grid with the minimum comprehensive loss, and take the grid with the minimum comprehensive loss as the positive sample.

[0016] According to a second aspect of the present invention, a device for image instance segmentation is provided, the device comprising:

[0017] A feature extraction module: configured to obtain a visible light image to be segmented and a binary edge image corresponding to the visible light image; extract a first image feature of the visible light image by a feature extraction network, wherein the first image feature is a multi-layer feature map having num levels of image features; and extract a second image feature of the binary edge image by a convolution operation;

[0018] A feature fusion module: configured to input the first image feature and the second image feature into a binary edge image guidance module, wherein the binary edge image guidance module includes a top-down fusion path and a bottom-down fusion path, and is used to fuse the first image feature and the second image feature to obtain a third image feature;

[0019] An attention adding module: configured to input the third image feature into a joint attention module, wherein the joint attention module is used to add spatial attention and coordinate attention to the third image feature to form a fourth image feature;

[0020] Segmentation module: configured to input the fourth image feature into a prediction module, the prediction module performs semantic prediction and mask prediction on the fourth image feature, and outputs a segmentation scheme for the object in the visible light image.

[0021] According to a third aspect of the present invention, a system for image instance segmentation is provided, comprising:

[0022] A processor, which is used to execute multiple instructions;

[0023] A memory for storing a plurality of instructions;

[0024] The plurality of instructions are used to be stored by the memory and loaded and executed by the processor to implement the method as described above.

[0025] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, wherein a plurality of instructions are stored in the storage medium; the plurality of instructions are used for a processor to load and execute the aforementioned method.

[0026] According to the above scheme of the present invention, the method of the present invention is implemented by a visible light image instance segmentation network model based on binary edge image guidance (Binary Edge Image Guide-Segmenting Objects by Locations, BEIG-SOLO). The visible light image instance segmentation network model is guided by a binary edge image guidance module, which fuses binary edge image feature information and visible light image feature information. The binary edge image guidance module includes a binary edge image information extraction module, a top-down and bottom-up fusion module, and after the binary edge image guidance module is embedded in the output of the feature extraction network, the network's perception of the boundary feature information of the visible light image is enhanced. In addition, a positive and negative sample label assignment strategy based on minimum loss matching is set. By calculating the sum of the semantic classification and mask losses of all grids, the grid with the smallest comprehensive loss is marked as a positive sample, which alleviates the problem of poor segmentation effect caused by the overlap of multiple object targets to the greatest extent. Experimental results show that on the visible light image dataset and the COCO dataset, the visible light image instance segmentation network model based on binary edge image guidance has good instance segmentation performance.

[0027] The technical effects of the present invention are:

[0028] (1) This paper proposes a network model for instance segmentation of visible light images based on binary edge image guidance. The network adds a binary edge image guidance module and a positive and negative sample label assignment strategy based on minimum loss matching, which effectively improves the instance segmentation performance of the network.

[0029] (2) The present invention proposes a binary edge image guidance module. The module integrates the feature information of the binary edge image into the multi-level feature fusion network and adds a bottom-up fusion path in the multi-level feature fusion network, which can effectively improve the network's perception of boundary features.

[0030] (3) The present invention proposes a negative sample label assignment strategy based on minimum loss matching. This strategy can effectively alleviate the problem of poor segmentation effect caused by the overlap of multiple instance targets and improve the instance segmentation performance of the network.

[0031] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present invention in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention, and the present invention is described by providing the following accompanying drawings. In the accompanying drawings:

[0033] Figure 1 A schematic diagram of a method flow for image instance segmentation according to an embodiment of the present invention;

[0034] Figure 2 This is a schematic diagram of a network model structure for instance segmentation of visible light images based on binary edge image guidance according to an embodiment of the present invention;

[0035] Figure 3 This is a schematic diagram of the structure of a binary edge image guiding module according to an embodiment of the present invention;

[0036] Figure 4 A structural block diagram of an apparatus for image instance segmentation according to an embodiment of the present invention. DETAILED DESCRIPTION

[0037] First combine Figure 1 The following describes an image instance segmentation method according to an embodiment of the present invention. Figure 1 As shown, the method comprises the following steps:

[0038] Step S1: obtaining a visible light image to be segmented and a binary edge image corresponding to the visible light image; extracting a first image feature of the visible light image by a feature extraction network, wherein the first image feature is a multi-layer feature map having num levels of image features; extracting a second image feature of the binary edge image by a convolution operation;

[0039] Step S2: inputting the first image feature and the second image feature into a binary edge image guiding module, wherein the binary edge image guiding module comprises a top-down fusion path and a bottom-down fusion path, and is used to fuse the first image feature and the second image feature to obtain a third image feature;

[0040] Step S3: inputting the third image feature into a joint attention module, wherein the joint attention module is used to add spatial attention and coordinate attention to the third image feature to form a fourth image feature;

[0041] Step S4: inputting the fourth image feature into a prediction module, the prediction module performs semantic prediction and mask prediction on the fourth image feature, and outputs a segmentation scheme for the object in the visible light image.

[0042] like Figure 2 As shown, the visible light image instance segmentation network model based on binary edge image guidance includes a feature extraction network, a binary edge image guidance module, a joint attention module and a prediction module connected in sequence.

[0043] The feature extraction network adopts a ResNet50 network, which is used to extract a first image feature from the visible light image. The first image feature is a multi-layer feature map with multiple levels of image features, and different levels focus on different objects and detail features.

[0044] In this embodiment, since it is impossible to obtain visible light images and binary edge images of the same scene at the same time in practical applications, the BEIG-SOLO network is divided into a training phase and a testing phase. In the training phase, based on the visible light image dataset, an edge extraction algorithm is used to generate a corresponding binary edge image dataset, and the instance segmentation performance of the BEIG-SOLO network for visible light images is improved by guiding the binary edge images. In the testing phase, there is no need for binary edge images to guide, and the trained BEIG-SOLO network can be used directly to complete the instance segmentation task of visible light images. In the training phase, assume that (C, H and W are the number of channels, height and width of the input image respectively) is the visible light image input by BEIG-SOLO, is the binary edge image input by BEIG-SOLO, I Seg is the output after instance segmentation. First, I Op The input feature extraction network extracts feature information, which is used to generate feature maps of different levels to obtain C1 to C5. The process can be expressed as:

[0045] C i =f fe (I Op ),i=1~5 (Formula 1)

[0046] Among them, f fe (·) represents the implementation of the feature extraction network, C i Represents the feature maps of different levels of the feature extraction network output, i=1~5 means that the output has 5 levels.

[0047] The step S2, inputting the first image feature and the second image feature into a binary edge image guiding module, wherein the binary edge image guiding module comprises a top-down fusion path and a bottom-down fusion path for fusing the first image feature and the second image feature to obtain a third image feature, comprises:

[0048] Select the hierarchical features of level 2 to level num from the first image features in descending order of size, input the hierarchical features of level 2 to level num and the second image features into the binary edge image guiding module, and perform the binary edge image guiding on the hierarchical feature C of level num. i Perform 1×1 convolution to get P num , according to the order of i decreasing from num-1 to 2, for P i+1Up-sample and then perform the i-th level feature C i Perform 1×1 convolution and transform P i+1 The upsampling result is similar to C i The result of 1×1 convolution is added according to the corresponding elements of the pixel position to get P i ; Fuse P2 with the second image feature to obtain the fused image feature M2; in the order of i increasing from 2 to num-1, i Downsampling, for P i+1 Perform 1×1 convolution and transform M i The downsampling result is similar to P i+1 The result of 1×1 convolution is added according to the corresponding elements of the pixel position to get M i+1 ; For M2 to M num Do 3×3 convolution to get N2 to N num ; for N num Do coordinate convolution and get N num+1 ; Among them, N2 to N num+1 It is the output of the binary edge image guiding module, that is, the third image feature.

[0049] In order to use binary edge images to guide the object segmentation of visible light images, reduce the influence of noise in images, and improve the accuracy of segmentation, a method such as Figure 3 The binary edge image guidance module is shown. The binary edge image refers to the part of the target edge in the binary image. It has a certain spatial continuity and is the dividing line between the target and the background. The edge part is an important source of information about the object, which can determine the characteristics of the image. By using the binary edge image, more accurate edge features can be obtained, thereby improving the performance of object segmentation. In the object segmentation task, the edge of the target needs to be accurately extracted. Using the binary edge image to guide the object segmentation can reduce the impact of noise in the image to a certain extent and improve the accuracy of segmentation. At the same time, the edge structure and characteristics can also help the model better understand the shape and structure of the object, thereby improving the robustness of the segmentation.

[0050] In this embodiment, in order to obtain better feature representation and integrate the feature information of the binary edge image, the generated multi-layer feature maps C1~C5 and the features of the binary edge image are input into the binary edge image guidance module for feature fusion to obtain N2~N5. Through a feature fusion method that combines top-down and bottom-up, the high-level and low-level feature maps containing different features are fused step by step, so as to combine semantic information and shallow information to obtain a more accurate feature representation. The binary edge image guidance module can realize multi-scale prediction of targets of different sizes, and can also alleviate the overlap problem of different instance targets to a certain extent. The process can be expressed as:

[0051] N j =f BEIG (C i ,I Be ),i=1~5,j=2~5 (Formula 2)

[0052] Among them, f BEIG (·) represents the binary edge image guidance module, C i Represents the feature maps of different levels of the binary edge image guidance module input, N j Represents the feature maps of different levels output by the binary edge image guidance module, i=1~5 means that the input has 5 levels, and j=2~5 means that the output has 4 levels.

[0053] In this embodiment, the highest level feature map N output by the binary edge image guidance module is num The purpose of coordinate convolution is to add x, y channel information to strengthen the processing of position information. The process can be expressed as:

[0054] N6=f CoordConv (N5) (Formula 3)

[0055] Among them, f CoordConv (·) denotes coordinate convolution.

[0056] In this embodiment, after the visible light image passes through the feature extraction network, the output can be divided into 5 levels according to the size of the feature map, and the sizes are C1 to C5 from large to small, which is also the order of the levels from low to high. According to the principle of image processing, low-level feature maps usually pay more attention to small targets and detail features in the image, as well as their position information in the image. High-level feature maps pay more attention to large targets in the image and the semantic information they contain. In order to reduce network performance overhead, {C2, C3, C4, C5} are selected as the input of the binary edge image guidance module, and their sizes are reduced to {1 / 4, 1 / 8, 1 / 16, 1 / 32} compared to the size of the input image. First, C5 is convolved 1×1 to obtain P5. The process can be expressed as:

[0057] P5=Conv 1×1 (C5)

[0058] Among them, Conv 1×1 (·) denotes a 1×1 convolution.

[0059] Next, upsample P5, perform 1×1 convolution on C4, and add and sum the resulting feature maps according to the corresponding elements of the pixel positions to obtain P4; upsample P4, perform 1×1 convolution on C3, and add and sum the resulting feature maps according to the corresponding elements of the pixel positions to obtain P3; upsample P3, perform 1×1 convolution on C2, and add and sum the resulting feature maps according to the corresponding elements of the pixel positions to obtain P2. This top-down process is to fuse feature maps at different levels to better extract and identify the features of the target. The process can be expressed as:

[0060]

[0061] in, Indicates that the two feature maps are added according to the corresponding elements of the pixel position. 2× (·) indicates 2x upsampling, Conv 1×1 (·) denotes a 1×1 convolution.

[0062] Then, 3×3 convolution is performed on P2~P5 to obtain M2~M5. The process can be expressed as:

[0063] M i =Conv 3×3 (P i ),i=2~5

[0064] Among them, Conv 3×3 (·) denotes a 3×3 convolution.

[0065] Due to the different spatial structures, the feature distribution of visible light images and binary edge images is different, and it is difficult to directly fuse the two images to better complete the instance segmentation of the target. Therefore, based on this consideration, the convolution method is first used to extract image features from binary edge images. Through the convolution operation, the edge features in the image can be extracted, making the subsequent image processing and analysis more accurate and precise. The process can be expressed as:

[0066] F Be =σ(BN(Conv 3×3 (I Be )))

[0067] Among them, F Be Represents the binary edge image features extracted by convolution, I Be represents the input binary edge image, σ(·) represents the nonlinear activation of the activation function, BN(·) represents batch normalization, Conv 3×3 (·) represents a 3×3 convolution.

[0068] Next, the visible light image feature P2 is combined with the binary boundary image feature F Be After the fusion operation is performed, the number of channels is aligned through 1×1 convolution to obtain the fused image feature M2. The process can be expressed as:

[0069]

[0070] Among them, Conv 1×1 (·) represents a 1×1 convolution, Represents concatenation in the channel dimension.

[0071] Then, downsample M2, perform 1×1 convolution on P3, add and sum the resulting feature maps according to the corresponding elements of the pixel positions to obtain M3; downsample M3, perform 1×1 convolution on P4, add and sum the resulting feature maps according to the corresponding elements of the pixel positions to obtain M4; downsample M4, perform 1×1 convolution on P5, add and sum the resulting feature maps according to the corresponding elements of the pixel positions to obtain M5. This bottom-up fusion process is to fuse the feature maps of different levels again to better extract and identify the features of the target. The process can be expressed as:

[0072]

[0073] in, Indicates that the two feature maps are added according to the corresponding elements of the pixel position. 2× (·) indicates 2 times downsampling, Conv 1×1 (·) denotes a 1×1 convolution.

[0074] Next, perform 3×3 convolution on M2~M5 to obtain N2~N5. The process can be expressed as:

[0075] N i =Conv 3×3 (M i ),i=2~5

[0076] Among them, Conv 3×3 (·) denotes a 3×3 convolution.

[0077] Then, in order to obtain more accurate location feature information, coordinate convolution is performed on N5 to obtain N6. The process can be expressed as:

[0078] N6=CoordConv(N5)

[0079] Wherein, CoordConv(·) represents coordinate convolution. So far, N2 to N6 are all outputs of the binary edge image guidance module.

[0080] The step S3: input the third image feature into the joint attention module, the joint attention module is used to add spatial attention and coordinate attention to the third image feature to form a fourth image feature, wherein the output feature map N2~N6 of the binary edge image guidance module is used as the input of the joint attention module to obtain I', and after the joint attention module applies spatial attention and coordinate attention to the input feature map, the sensitivity of the network to the position feature is enhanced. The process can be expressed as:

[0081] I'=f JAM (M i ),i=2~6

[0082] Among them, f JAM (·) represents the joint attention module, and i=2~6 means that the input has 5 levels.

[0083] In this embodiment, the joint attention module can apply spatial attention and coordinate attention at the same time. Embedding the joint attention module into the output of each level of the binary edge image module can improve the network's perception of position features.

[0084] The step S4: inputs the fourth image feature into a prediction module, the prediction module performs semantic prediction and mask prediction on the fourth image feature, and outputs a segmentation scheme for the object in the visible light image, wherein the prediction module includes a semantic category branch, a mask convolution kernel branch, and a mask feature branch, the semantic category branch is used to perform an interpolation convolution operation on the input fourth image feature to obtain the probability that the content corresponding to each grid unit belongs to the target category; the mask convolution kernel branch is parallel to the semantic category branch, and is used to perform a dynamic convolution calculation operation on the input fourth image feature to obtain a predicted mask; the mask feature branch is used to perform convolution, batch normalization nonlinear activation, and adaptive interpolation operation on the input fourth image feature to obtain a mask feature map, and based on the mask feature map, a segmentation scheme for the object in the visible light image is obtained.

[0085] In this embodiment, the multi-level feature map I' after attention is input into the prediction module for mask prediction. The semantic category branch and the mask convolution kernel branch in the prediction module are parallel. The mask feature branch is directly generated by the multi-level feature map after attention is applied. The prediction mask is dynamically generated, and the instance segmentation of the image is completed by matrix non-maximum suppression to obtain I Seg . The process can be expressed as:

[0086] I Seg =f M-NMS (f PM (I'))

[0087] Among them, f M-NMS(·) represents the implementation of matrix non-maximum suppression, f PM (·) indicates the implementation of the prediction module.

[0088] In order to reduce the computational cost of traditional non-maximum suppression (NMS) in the mask output process, a matrix non-maximum suppression (Matrix NMS) mechanism is adopted, which has the characteristics of hard deletion and sequential operation.

[0089] Furthermore, non-maximum suppression is performed on the output of the prediction module, including: obtaining all prediction scores output by the prediction module, subjecting the top N prediction scores with the highest scores to non-maximum suppression processing to obtain an IoU matrix of size N×N; performing a maximum value operation on the column direction of the IoU matrix, and determining the mask to be retained based on the maximum value of each column of the IoU matrix; based on the prediction score of each prediction, calculating the penalty of the prediction on other predictions through a linear function, and then fitting the most overlapping predictions as the suppression probability, and finally calculating the attenuation factor using the penalty and suppression probability of each prediction, and obtaining the final prediction based on the attenuation factor.

[0090] Furthermore, we predict m j The attenuation factor is affected by the following factors: (a) Each prediction m i For m j (s i >s j ) penalty, where s i and j is the confidence score; (b) m i The probability of being suppressed.

[0091] For (a), each prediction m i For m j (s i >s j ) can be calculated by f(iou i,j ) is calculated. For (b), the suppression probability is usually positively correlated with IoU. In order to simplify the calculation, m i The probability is approximated by the most overlapping predictions, i.e.

[0092]

[0093] For this purpose, the final attenuation factor can be expressed as:

[0094]

[0095] The updated score is s j =s j·decay j Calculate. Where f(iou i,j )=1-iou i,j .

[0096] The prediction module is the core part of BEIG-SOLO. It takes the feature map after attention as input to complete the instance segmentation task. At the same time, in order to solve the problem of poor segmentation results caused by multi-target overlap, a positive and negative sample label assignment strategy based on minimum loss matching is proposed.

[0097] Furthermore, in the training process of the visible light image instance segmentation network model guided by binary edge images, a positive and negative sample label assignment strategy based on minimum loss matching is introduced. The positive and negative sample label assignment strategy based on minimum loss matching is:

[0098] Determine the level outputs N2 to N3 corresponding to the third image feature output by the binary edge image guiding module. num +1, determine the object size range corresponding to the output of each level, and the object size ranges corresponding to each level are different and have no intersection; when an object has multiple possible sizes, each possible size corresponds to a level, and the outputs of the multiple levels corresponding to the object are predicted respectively; in each corresponding level, by calculating the classification and segmentation losses of each grid in the true target box and the true target, find the grid with the minimum comprehensive loss, and take the grid with the minimum comprehensive loss as the positive sample.

[0099] In this embodiment, label assignment is a very important step in the network training process. It has a great impact on the upper limit of model accuracy. The purpose of label assignment is to distinguish positive and negative samples during the training phase, thereby providing a suitable learning target for the network. The SOLO-based instance segmentation network divides the feature map into S×S grids of the same size in the prediction module, and determines which grids are to be marked as positive samples and which grids are to be marked as negative samples according to the centroid of the target's ground truth box (GT box).

[0100] Assume that the center of the GT box of a given target is (c x ,c y), with a width of w and a height of h, and a scaling factor of ε is set. The GT box of the target is reduced by ε to obtain a positive sample target box (Positive Sample box, PS box). When the grid at the position (i, j) falls within the PS box, the grid is marked as a positive sample. When there are two or more instance targets in the input image that overlap to a large extent, or when one instance target is completely surrounded by another instance target, the positive sample target boxes of these two or more instance targets may fall into the same grid, which causes the problem of different instance targets being predicted by the same grid, causing serious confusion in the contours of different instance targets, which greatly interferes with the prediction results.

[0101] In order to solve the above problems, the present invention proposes a positive and negative sample label assignment strategy based on minimum cost assignment (MCA), which is divided into positive and negative sample label assignment strategies at different levels and at the same level.

[0102] First, a positive and negative sample label assignment strategy for different levels is proposed. The role of this strategy is to make each level responsible for a target within a specific size range. The output of the binary edge image guidance module contains the output of 5 levels, from N2 to N6 from low to high, and the corresponding size ranges are ((384,2048), (192,768), (96,384), (48,192), (1,96)).

[0103] There are two cases for the size of the same instance target, one is in one size range and the other is in multiple size ranges. Assume that the size of instance target A is only in the range of (384, 2048), then the corresponding N2 layer will be responsible for predicting instance target A. Assuming that the size of instance target A is in both the range of (384, 2048) and (192, 768), then the corresponding N2 layer and N3 layer will be responsible for predicting instance target A. In this way, the network can more effectively predict targets of various sizes. When an instance target falls within multiple ranges at the same time, the instance target can be predicted in different ranges, thereby increasing the number of positive samples and alleviating the problem of a large gap between the number of positive and negative samples.

[0104] In special cases, when the size range of the level corresponding to N2 to N6 does not contain any instance target size, all grid samples of this level will be ignored and not counted as positive or negative samples. Assuming that the range of (384,2048) of the N2 level does not contain any instance target size, all grids of the N2 level will be ignored. This strategy can effectively improve the performance and robustness of the network, making it perform well on a variety of different target sizes and shapes.

[0105] Secondly, a positive and negative sample label assignment strategy for the same level is proposed. When a level is assigned the task of predicting an instance target, positive and negative sample labels must be assigned to the grids of this level. Assuming that level N2 is responsible for predicting instance target A, positive and negative sample labels must be assigned to the grids of level N2. MCA regards the assignment of positive and negative sample labels to the grid as an adaptive matching problem. By calculating the classification and segmentation losses of each grid in the true target box and the true target, the grid with the minimum comprehensive loss is found and marked as a positive sample. This method can better handle the overlap problem between instance targets and improve the accuracy and stability of the network. The process can be expressed as:

[0106] M = min(focalloss(c i,j ,c gt )+diceloss(m i,j ,m gt ))

[0107] Among them, M represents the grid with the smallest comprehensive loss in the real target box, which is marked as a positive sample. i,j is the classification prediction value of the grid at position i, j, c gt is the true category value of the sample, and focalloss is the classification loss function. i,j is the segmentation prediction value of the grid at position (i, j), m gt is the true mask of the sample, diceloss is the loss function of segmentation, and i,j is the range of the true target box mapped to the S×S grid.

[0108] During the training process, instance targets are arranged from small to large according to the area of ​​the true target box, and the positive sample grid of each instance target is calculated in turn. The grids that have been marked as positive samples will be ignored in the next calculation process to prevent overlapping instance targets from being marked in the same grid.

[0109] The overall loss function of the visible light image instance segmentation network model based on binary edge image guidance is:

[0110] L=L cate +λL mask

[0111] Among them, L cate is the loss function of the semantic branch, L mask is the loss function of the mask branch, λ is a hyperparameter, set to 3; L cate The focal loss is used. mask The expression is:

[0112]

[0113] Where k = i × S + j, N pos is the number of positive samples, p * ,m* are the category truth value and mask truth value respectively, is an indicator function, if If yes, it is 1, otherwise it is 0; m k To predict the mask, is the mask truth value;

[0114] d mask The implementation is as follows, using dice loss and boundary loss:

[0115] d mask =L Dice +L Boundary

[0116] Among them, L Dice is defined as follows:

[0117]

[0118] Among them, p x,y ,q x,y are the pixel values ​​of the predicted mask and the true mask at the (x, y) position respectively;

[0119] L Boundary is the boundary loss, which is calculated as:

[0120] According to the real mask m and the predicted mask m of the target * Construct the true boundary and predicted boundary of the target, and the calculation formula is:

[0121] b=pool(1-m,θ)-(1-m)

[0122] b * =pool(1-m * ,θ)-(1-m * )

[0123] Among them, b,b * denote the real boundary and predicted boundary respectively, pool(·,·) denotes the pixel-level maximum pooling operation, m,m * represent the true mask and the predicted mask respectively, and θ is a hyperparameter set to 3.

[0124] According to the obtained real boundary b and predicted boundary b * Construct boundary loss and calculate it as:

[0125]

[0126] Among them, i is the boundary index, b i is the i-th true boundary, is the i-th prediction boundary.

[0127] The embodiment of the present invention further provides a device for image instance segmentation, such as Figure 4 As shown, the device comprises:

[0128] A feature extraction module: configured to obtain a visible light image to be segmented and a binary edge image corresponding to the visible light image; extract a first image feature of the visible light image by a feature extraction network, wherein the first image feature is a multi-layer feature map having num levels of image features; and extract a second image feature of the binary edge image by a convolution operation;

[0129] A feature fusion module: configured to input the first image feature and the second image feature into a binary edge image guidance module, wherein the binary edge image guidance module includes a top-down fusion path and a bottom-down fusion path, and is used to fuse the first image feature and the second image feature to obtain a third image feature;

[0130] An attention adding module: configured to input the third image feature into a joint attention module, wherein the joint attention module is used to add spatial attention and coordinate attention to the third image feature to form a fourth image feature;

[0131] Segmentation module: configured to input the fourth image feature into a prediction module, the prediction module performs semantic prediction and mask prediction on the fourth image feature, and outputs a segmentation scheme for the object in the visible light image.

[0132] The embodiment of the present invention further provides a system for image instance segmentation, including:

[0133] A processor, which is used to execute multiple instructions;

[0134] A memory for storing a plurality of instructions;

[0135] The plurality of instructions are used to be stored by the memory and loaded and executed by the processor to implement the method as described above.

[0136] The embodiment of the present invention further provides a computer-readable storage medium, wherein a plurality of instructions are stored in the storage medium; the plurality of instructions are used for a processor to load and execute the method as described above.

[0137] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.

[0138] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0139] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0140] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0141] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a physical server, or a network cloud server, etc., and the Ubuntu operating system must be installed) to perform some steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), disk or optical disk and other media that can store program codes.

[0142] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention still falls within the scope of the technical solution of the present invention.

Claims

1. A method for image instance segmentation, characterized in that: The method comprises the following steps: Step S1: obtaining a visible light image to be segmented and a binary edge image corresponding to the visible light image; extracting a first image feature of the visible light image by a feature extraction network, wherein the first image feature is a multi-layer feature map having num levels of image features; extracting a second image feature of the binary edge image by a convolution operation; Step S2: inputting the first image feature and the second image feature into a binary edge image guiding module, wherein the binary edge image guiding module comprises a top-down fusion path and a bottom-down fusion path, and is used to fuse the first image feature and the second image feature to obtain a third image feature; Step S3: inputting the third image feature into a joint attention module, wherein the joint attention module is used to add spatial attention and coordinate attention to the third image feature to form a fourth image feature; Step S4: inputting the fourth image feature into a prediction module, the prediction module performs semantic prediction and mask prediction on the fourth image feature, and outputs a segmentation scheme for the object in the visible light image.

2. The method according to claim 1, characterized in that The step S2 comprises: Select the hierarchical features of level 2 to level num from the first image features in descending order of size, input the hierarchical features of level 2 to level num and the second image features into the binary edge image guiding module, and perform the binary edge image guiding on the hierarchical feature C of level num. i Perform 1×1 convolution to get P num , according to the order of i decreasing from num-1 to 2, for P i+1 Up-sample and then perform the i-th level feature C i Perform 1×1 convolution and transform P i+1 The upsampling result is similar to C i The result of 1×1 convolution is added according to the corresponding elements of the pixel position to get P i ; Fuse P2 with the second image feature to obtain the fused image feature M2; in the order of i increasing from 2 to num-1, i Downsampling, for P i+1 Perform 1×1 convolution and transform M i The downsampling result is similar to P i+1 The result of 1×1 convolution is added according to the corresponding elements of the pixel position to get M i+1 ; For M2 to M num Do 3×3 convolution to get N2 to N num ; for N num Do coordinate convolution and get N num+1 ; Among them, N2 to N num+1 is the output of the binary edge image guiding module, that is, the third image feature; num is a natural number.

3. The method as claimed in claim 2, characterized in that The step S4, wherein the prediction module includes a semantic category branch, a mask convolution kernel branch and a mask feature branch, the semantic category branch is used to perform an interpolation convolution operation on the input fourth image feature to obtain the probability that the content corresponding to each grid unit belongs to the target category; the mask convolution kernel branch is parallel to the semantic category branch, and is used to perform a dynamic convolution calculation operation on the input fourth image feature to obtain a prediction mask; the mask feature branch is used to perform convolution, batch normalization nonlinear activation and adaptive interpolation operation on the input fourth image feature to obtain a mask feature map, and based on the mask feature map, a segmentation scheme for the object in the visible light image is obtained.

4. The method according to claim 3, characterized in that The output of the prediction module is subjected to non-maximum suppression, including: obtaining all prediction scores output by the prediction module, subjecting the top N prediction scores with the highest scores to non-maximum suppression processing to obtain an IoU matrix of size N×N; performing a maximum value operation on the column direction of the IoU matrix, and determining a mask to be retained based on the maximum value of each column of the IoU matrix; based on the prediction score of each prediction, calculating the penalty of the prediction on other predictions through a linear function, and then fitting the prediction with the largest overlap as the suppression probability, calculating an attenuation factor using the penalty and suppression probability of each prediction, and obtaining a final prediction based on the attenuation factor.

5. The method according to claim 4, characterized in that In the training process of the visible light image instance segmentation network model guided by binary edge images, a positive and negative sample label assignment strategy based on minimum loss matching is introduced. The positive and negative sample label assignment strategy based on minimum loss matching is: Determine the level outputs N2 to N3 corresponding to the third image feature output by the binary edge image guiding module. num +1, determine the object size range corresponding to the output of each level, and the object size ranges corresponding to each level are different and have no intersection; when an object has multiple possible sizes, each possible size corresponds to a level, and the outputs of the multiple levels corresponding to the object are predicted respectively; in each corresponding level, by calculating the classification and segmentation losses of each grid in the true target box and the true target, find the grid with the minimum comprehensive loss, and take the grid with the minimum comprehensive loss as the positive sample.

6. A device for image instance segmentation, characterized in that: The device comprises: A feature extraction module: configured to obtain a visible light image to be segmented and a binary edge image corresponding to the visible light image; extract a first image feature of the visible light image by a feature extraction network, wherein the first image feature is a multi-layer feature map having num levels of image features; and extract a second image feature of the binary edge image by a convolution operation; A feature fusion module: configured to input the first image feature and the second image feature into a binary edge image guidance module, wherein the binary edge image guidance module includes a top-down fusion path and a bottom-down fusion path, and is used to fuse the first image feature and the second image feature to obtain a third image feature; An attention adding module: configured to input the third image feature into a joint attention module, wherein the joint attention module is used to add spatial attention and coordinate attention to the third image feature to form a fourth image feature; Segmentation module: configured to input the fourth image feature into a prediction module, the prediction module performs semantic prediction and mask prediction on the fourth image feature, and outputs a segmentation scheme for the object in the visible light image.

7. A system for image instance segmentation, comprising: A processor, which is used to execute multiple instructions; A memory for storing a plurality of instructions; The plurality of instructions are used to be stored in the memory and loaded and executed by the processor according to any one of claims 1 to 5.

8. A computer-readable storage medium, wherein a plurality of instructions are stored in the storage medium; the plurality of instructions are used for a processor to load and execute the method as claimed in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and computer readable medium

    CN112419342A

  • Automatic extraction method for ancient city battlement

    CN114511582A

  • Living fish weight estimation method and system based on instance segmentation

    CN114998375A

  • Ship instance segmentation algorithm based on global and local attention mechanisms

    CN115797626A

  • Aircraft target instance segmentation method and system and readable storage medium

    CN116310323A

Cited By

  • Image instance segmentation method and device based on boundary guidance, and related medium

    CN121639721A