An improved Mask R-CNN object segmentation method
By fusing global features in the image target segmentation method and using the coordination loss function to generate detection boxes, and combining the boundary strengthening method to generate target boundaries and masks, the problem of inaccurate target segmentation and lack of boundary details in the prior art is solved, and a higher accuracy of target segmentation is achieved.
Patent Information
- Application Number
- CN202210038272.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-13
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-01-13
AI Technical Summary
In the prior art, when the image target segmentation method processes complex scenes, there are problems such as low mask quality, missing target boundary details information, and inaccurate segmentation.
By fusing global features into the features of the region of interest, and using the coordination loss function to generate detection boxes under the premise that regression branches and classification branches are mutually supervised, the generation of target boundary and masks is achieved in combination with boundary enhancement methods.
The accuracy of target segmentation is improved, the problems of incomplete mask segmentation and rough target boundaries are avoided, and the target detection and segmentation performance in complex scenarios is enhanced.
Smart Images

Figure CN114445620B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and more specifically, to an object segmentation method for improving Mask R-CNN. Background Art
[0002] Image object segmentation is an important research topic in computer vision. It combines the characteristics of both object detection and semantic segmentation tasks. It can complete image understanding by predicting the category label and pixelated object mask of each object in the image. It has good development prospects and important significance. At present, this technology has been widely used in fields such as autonomous driving, urban monitoring and robot grasping.
[0003] In the prior art, image target segmentation methods mainly include the following three categories: a single-stage method based on deep learning, a two-stage method based on segmentation, and a two-stage method based on detection.
[0004] First, the single-stage method based on deep learning aims to complete the three tasks of positioning, classifying and segmenting objects in one stage. Among them, some technical solutions can use polar coordinates to model contours, and complete target segmentation by classifying target center points and dense distance regression. Some technical solutions can merge mask prediction into the fully convolutional network, and encode the target shape by using a set of contour coefficients to form a single-stage target segmentation framework. There are also technical solutions that can perform point-based segmentation prediction at adaptively selected locations through iterative subdivision methods.
[0005] However, this method needs to complete the three tasks of object localization, classification and segmentation in one stage, so the mask quality is usually low. This method still faces great challenges in object detection and pixel-level feature alignment in complex scenes.
[0006] Secondly, the two-stage segmentation-based method can perform pixel-level semantic segmentation on the image and combine the pixels of each object through clustering, metric learning and other means to distinguish different targets. Some technical solutions in this method can use deep metric learning to learn the embedding of the target, and then group the pixels to form target-level segmentation. Some technical solutions can also use a series of neural networks to solve the complex task of target segmentation. Since each neural network can be used to solve a subclass problem with increased semantic complexity, the method can gradually construct target instances using a simple structure.
[0007] However, this method often leads to inaccurate segmentation when processing the underlying features of the target, and the generalization ability of the algorithm is cross-cutting and cannot cope with complex scenarios with more categories.
[0008] Third, the two-stage detection-based method can first detect the target in the image, and after finding the area where the target is located, perform semantic segmentation within a specific detection box. Different targets can be output as separate segmentation results. In this method, some algorithms can slide on the image using windows of different sizes, and use classifiers to determine the probability of the existence of the target in the sliding box, thereby generating candidate sub-regions based on this probability, and ultimately achieving target recognition and segmentation. Other algorithms can support selective search to generate candidate sub-regions, that is, only the areas in the image that are most likely to contain the target are detected.
[0009] However, since the essence of this method is still to perform pixel-by-pixel segmentation in the predicted detection box, the task will be overly dependent on the accuracy of the detection box. During the sampling process, detailed feature information may be lost or the spatial feature resolution may be too low, making it difficult to classify pixels close to the boundary, and the mask only captures the general shape of the target. Since the boundary outline of the target is not clear enough, the target boundary detail information may often be missing during the target extraction process, which further affects the quality of target segmentation.
[0010] In view of the above problems, the present invention provides an improved target segmentation method of Mask R-CNN. Summary of the invention
[0011] In order to solve the shortcomings of the prior art, the purpose of the present invention is to provide an improved target segmentation method of Mask R-CNN, by fusing global features into the features of the region of interest, and generating a detection frame under the premise of mutual supervision between the regression branch and the classification branch through a coordinated loss function, and realizing the generation of boundaries and masks based on boundary reinforcement.
[0012] The present invention adopts the following technical solution.
[0013] An improved Mask R-CNN target segmentation method comprises the following steps: step 1, using the Mask R-CNN target segmentation method to obtain ROI region features in the original image, wherein after the global context information is subjected to feature conversion using a one-dimensional attention mechanism, the global features obtained by the conversion are fused into the ROI region features; step 2, a classification loss function and a regression loss function are calculated for a feature map containing the ROI region features, and a coordinated loss function is constructed using mutually supervised classification loss weights and regression loss weights to predict an optimal detection frame; step 3, target segmentation is performed on the local features extracted from the optimal detection frame, wherein the features of the target boundary are enhanced using a boundary enhancement method, thereby generating a target mask.
[0014] Preferably, the method for acquiring the ROI region features in the original image in step 1 is: step 1.1.1, using the residual network ResNetXt-101 and the feature pyramid network FPN to generate multiple feature maps of different scales; step 1.1.2, using the region candidate network RPN to generate a region of interest; step 1.2.3, using the feature of interest alignment method ROI Align to extract local region features in the region of interest, and using the channel multiplication method to realize the fusion of local region features and global features.
[0015] Preferably, the method for obtaining the global features in step 1 is as follows: step 1.2.1, using global average pooling GAP to reduce the dimension of each feature map of different scales generated in step 1.1.1, and then merging the reduced dimension information; step 1.2.2, using a lightweight attention mechanism to perform feature conversion on the merged reduced dimension information to obtain global features.
[0016] Preferably, the lightweight attention mechanism is implemented using a one-dimensional convolution with a convolution kernel of 5.
[0017] Preferably, a coordinated loss function is used to respectively train the classification branch and the regression branch of the feature map of the ROI region feature to predict the optimal detection frame.
[0018] Preferably, the coordination loss function is:
[0019]
[0020] Among them, i is the sequence number of the predicted candidate box in the small batch sample,
[0021] p i is the predicted probability of the i-th predicted candidate box,
[0022] y i is the positive and negative label of the original marked box of the i-th predicted candidate box,
[0023] d i is the coordinate vector of the i-th predicted candidate box,
[0024] is the coordinate vector of the original labeled box corresponding to the i-th predicted candidate box,
[0025] CE() is the cross entropy loss function, L() is the smooth L1 loss function,
[0026] γ r and γ r They are regression coordination factor and classification coordination factor respectively.
[0027] Preferably, a regression coordination factor is used to realize mutual supervision between classification loss and regression loss; and the regression coordination factor is
[0028]
[0029] The classification coordination factor is
[0030]
[0031] Preferably, in step 3, the method for generating the target mask is: step 3.1, using the boundary enhancement branch to generate the prediction features of the target boundary, and generating the predicted target boundary based on the prediction features; step 3.2, adding the prediction features of the target boundary to the mask branch to realize the generation of the target object mask.
[0032] Preferably, the boundary reinforcement branch includes a first sub-branch, a second sub-branch, a third sub-branch and a fourth sub-branch; wherein, the first sub-branch transforms and deforms the local features extracted from the feature map, and multiplies them with the position attention matrix jointly generated by the second sub-branch and the third sub-branch to obtain a result matrix; the second sub-branch reduces the dimension and transforms the local features extracted from the feature map, and after multiplying them with the transposed matrix generated by the third branch, generates the position attention matrix through a normalized exponential function; the third sub-branch reduces the dimension and transforms the local features extracted from the feature map, and multiplies them with the feature matrix of the second sub-branch after transposition; the fourth branch adds the local features extracted from the feature map to the result matrix using a residual connection, and realizes feature fusion through convolutional layers and upsampling.
[0033] Preferably, the loss function of the boundary strengthening branch is a binary cross entropy loss function L; and,
[0034]
[0035] Among them, y i is the positive or negative label of the original marked box of the i-th predicted candidate box, which takes a value of 1 or 0; N is the total number of training samples of the small batch samples of the predicted candidate box.
[0036] Preferably, the number of convolutional layers in the fourth branch is 4; two layers of convolution are performed on the predicted features of the target boundary generated after feature fusion to generate a predicted target boundary.
[0037] Preferably, the mask branch performs 4-layer convolution and upsampling on the local features extracted from the feature map, adds them to the predicted features of the target boundary, and then performs deconvolution to generate a mask.
[0038] The beneficial effect of the present invention is that, compared with the prior art, an improved target segmentation method of Mask R-CNN in the present invention can generate a detection frame by integrating global features into the features of the region of interest, and by coordinating the loss function under the premise of mutual supervision between the regression branch and the classification branch, and realize the generation of boundaries and masks based on boundary reinforcement. The method of the present invention can extract and enhance the target boundary features in multiple processes of image processing, predict a complete detection frame of the target, thereby avoiding the problems of incomplete mask segmentation and rough target boundaries, and greatly improving the accuracy of target segmentation.
[0039] The beneficial effects of the present invention also include:
[0040] 1. The method of the present invention is based on deep convolutional neural networks (CNN), so it has wider applicability in deep learning. It can be applied to various scenarios such as accurately calculating the layout of building scenes and helping to focus on the lesion area in the patient's body. It has good adaptability in different types of target segmentation and detection processes.
[0041] 2. The method of the present invention focuses on improving Mask R-CNN (Regions with CNN features, a regional method based on CNN features). While integrating the beneficial effects of Mask R-CNN, it also improves the technical problems of Mask R-CNN, and can effectively change the problems of target detail feature loss and rough segmentation edges in the image, making the generated mask more accurate, natural and complete. In addition, the method of the present invention can also face more complex scenes in real life and achieve more accurate target segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic diagram of an implementation flow of an improved Mask R-CNN target segmentation method in the present invention;
[0043] Figure 2 It is a schematic diagram of the implementation process of the global context information fusion module in an improved Mask R-CNN target segmentation method in the present invention;
[0044] Figure 3 A schematic diagram of a method for obtaining a detection box loss function in an improved Mask R-CNN target segmentation method in the present invention;
[0045] Figure 4 It is a schematic diagram of the implementation process of a boundary enhancement method in an improved Mask R-CNN target segmentation method in the present invention;
[0046] Figure 5 This is a comparison diagram of the target segmentation result achieved by an improved Mask R-CNN target segmentation method in the present invention and the target segmentation result achieved by Mask R-CNN in the prior art. DETAILED DESCRIPTION
[0047] The present application is further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present application.
[0048] Figure 1 FIG. 4 is a schematic diagram of an implementation flow of an improved Mask R-CNN target segmentation method in the present invention. Figure 1 As shown, an improved Mask R-CNN target segmentation method in the present invention specifically includes steps 1 to 3.
[0049] Among them, in step 1, the target segmentation method of Mask R-CNN is used to obtain the ROI area features in the original image, wherein the global context information is transformed by a one-dimensional attention mechanism, and the global features obtained by the transformation are fused into the ROI area features.
[0050] Specifically, the method used in step 1 can be implemented by referring to the Mask R-CNN method in the prior art. However, it is slightly different from the method in the prior art. In the present invention, global features are extracted at the same time and effectively integrated into ROI (Region of Interest) regional features. Through this method, some boundary information in the global features can be effectively obtained before mask generation and target segmentation of the image. Even if the target boundary information cannot be completely included in the region of interest, the target boundary information will not be missed by the global features.
[0051] Preferably, the method for acquiring the ROI region features in the original image in step 1 is: step 1.1.1, using the residual network ResNetXt-101 and the feature pyramid network FPN (Feature Pyramid Network) to generate multiple feature maps of different scales; step 1.1.2, using the region candidate network RPN to generate a region of interest; step 1.2.3, using the feature alignment method of interest (ROI Align) to extract local region features in the region of interest, and using the channel multiplication method to realize the fusion of local region features and global features.
[0052] It can be understood that, in the present invention, a Mask R-CNN method similar to that in the prior art can be used to extract ROI information in the backbone network.
[0053] Figure 2 FIG. 1 is a schematic diagram of the implementation process of the global context information fusion module in the target segmentation method of an improved Mask R-CNN in the present invention. Figure 2 As shown, preferably, the method for obtaining the global features in step 1 is: step 1.2.1, using global average pooling GAP to reduce the dimension of each feature map of different scales generated in step 1.1.1, and then merging the reduced dimension information; step 1.2.2, using a lightweight attention mechanism to perform feature conversion on the merged reduced dimension information to obtain the global features.
[0054] The difference between the present invention and the prior art is that a global context information fusion module (GCIFM) is added in the present invention, and the extraction of global features can be realized by adding this module.
[0055] Specifically, the feature maps of multiple different scales extracted in the above steps, such as the features {P2, P3, P4, P5} in each layer of the feature pyramid network in the backbone network, can be subjected to global average pooling (GAP, Global Average Pooling). The pooling process here is mainly to compress the global context features, while retaining the important features in the global information, it can also effectively reduce the scale of the features. Then, the multiple pooled features are merged and input into the lightweight attention module. The merging method in the present invention can be implemented by using the method of element-by-element addition to achieve the fusion of global context information at four different levels. The global features are refined by the lightweight attention module. Here, the merged pooled information can first be subjected to a feature transformation (i.e., Figure 1 The Transform described in , can be transformed by GLU (Gated Linear Units) or FC (Fully Connected Layers) which are often used in lightweight attention modules. After one transformation, the generated information can conform to the input format of the one-dimensional convolution layer, so the spatial relationship between the input channel features and the adjacent channel features can be realized by the one-dimensional convolution layer (Conv1d). After the output of the one-dimensional convolution layer, the output of the lightweight attention mechanism is realized by transformation.
[0056] Preferably, the lightweight attention mechanism is implemented using a one-dimensional convolution with a convolution kernel of 5.
[0057] In other words, the method of the present invention can realize local feature interaction between each channel and the five adjacent channels.
[0058] The fusion between the local area features and the global features included in the aforementioned step 1.2.3 of the present invention can be understood as the global context feature weight parameters output by the lightweight attention mechanism and the ROI features generated by the backbone network are multiplied according to the channel to obtain information fusion. This fusion process can effectively improve the accuracy of subsequent detection frame generation and the accuracy of target edge positioning, thereby improving the detection and segmentation performance of the network.
[0059] Step 2: Calculate the classification loss function and regression loss function for the feature map containing the ROI region features, and use the mutually supervised classification loss weight and regression loss weight to construct a coordinated loss function to predict the optimal detection box.
[0060] Figure 3 Schematic diagram of a method for obtaining a detection box loss function in an improved Mask R-CNN target segmentation method in the present invention. Figure 3 As shown, preferably, a coordinated loss function is used to respectively train the classification branch and the regression branch of the feature map of the ROI region feature to predict the optimal detection box.
[0061] It is understandable that in order to better ensure the consistency of detection box prediction, the present invention proposes the concept of coordination loss function. Figure 3 As shown, in the present invention, the feature map obtained after fusion is input into the detection module to realize the prediction of the detection frame. First of all, in the detection module, two branches, a classification branch and a regression branch, may generally be included. In the prior art, the above two branches usually complete the prediction of the detection frame separately. This makes the loss functions of the two branches unrelated in the calculation process, and the loss function obtained by simple summation is not accurate enough, nor can it fully reflect the mutual influence between the classification branch and the regression branch in the detection module. In this way, it is easy to obtain a detection result with a high classification score but a low IOU (Intersection-over-Union, the intersection-over-union ratio between the predicted candidate box and the original labeled box) during the calculation process, or a detection result with a low classification score and a high IOU score. The above two inconsistent detection results are both caused by the fact that the classification loss and the regression loss are not associated during the independent training process.
[0062] To address the above problems, the present invention adopts the following coordination loss function.
[0063] Preferably, the coordination loss function is:
[0064]
[0065] Among them, i is the sequence number of the predicted candidate box in the small batch sample,
[0066] pi is the predicted probability of the i-th predicted candidate box,
[0067] y i is the positive and negative label of the original marked box of the i-th predicted candidate box,
[0068] d i is the coordinate vector of the i-th predicted candidate box,
[0069] is the coordinate vector of the original labeled box corresponding to the i-th predicted candidate box,
[0070] CE() is the cross entropy loss function, L() is the Smooth L1 loss function,
[0071] γ r and γ c They are regression coordination factor and classification coordination factor respectively.
[0072] It can be understood that the CE(p i ,y i ) is the cross entropy loss function used by the classification branch, and is the smooth L1 loss function used by the smooth branch. In order to ensure the correlation between the two during separate training, the weights represented by the regression coordination factor and the classification coordination factor are also added (1+γ r ) and (1+γ c ).
[0073] Preferably, a regression coordination factor is used to realize mutual supervision between classification loss and regression loss; and the regression coordination factor is
[0074]
[0075] The classification coordination factor is
[0076]
[0077] In the present invention, a regression coordination factor is assigned to the classification branch, so that the optimization of the classification branch can be dynamically supervised. Similarly, a classification coordination factor can be assigned to the regression branch, so that the optimization of the regression branch can be supervised. Therefore, during the optimization process, the regression loss can be perceived by the classification branch, and the classification loss can also be perceived by the regression branch.
[0078] Therefore, the method of the present invention can very accurately predict the optimal detection frame with relatively consistent values of classification loss and regression loss.
[0079] Specifically, the present invention can determine the target category of the region of interest and regress the location information of the region of interest, and simultaneously use the coordinated loss function to complete the training of the classification branch and the regression branch respectively, thereby generating an optimal detection frame.
[0080] like Figure 3 As shown, in one embodiment of the present invention, a 7*7 region of interest is generated, and 256 target categories are used for each region of interest, and then a 1024 fully connected layer is used to implement regression branch and classification branch training respectively.
[0081] Step 3: segment the target using the local features extracted by the optimal detection frame, wherein the boundary enhancement method is used to enhance the features of the target boundary, thereby generating a target mask.
[0082] After achieving the optimal detection frame, local features can be extracted based on the optimal detection frame.
[0083] Preferably, in step 3, the method for generating the target mask is: step 3.1, using the boundary enhancement branch to generate the prediction features of the target boundary, and generating the predicted target boundary based on the prediction features; step 3.2, adding the prediction features of the target boundary to the mask branch to realize the generation of the target object mask.
[0084] It can be understood that the method for generating the target mask in the present invention can be implemented based on enhanced boundary information and an optimal detection frame.
[0085] First, the mask generation process can include two branches: the mask branch and the boundary enhancement branch. The content of the boundary enhancement module in the boundary enhancement branch is as follows: Figure 4 shown.
[0086] Figure 4 The figure is a schematic diagram of the implementation process of the boundary enhancement method in the target segmentation method of an improved Mask R-CNN in the present invention. Figure 4 In the method, preferably, the boundary reinforcement branch includes a first sub-branch, a second sub-branch, a third sub-branch and a fourth sub-branch; wherein, the first sub-branch transforms and deforms the local features extracted from the feature map, and multiplies them with the position attention matrix jointly generated by the second sub-branch and the third sub-branch to obtain a result matrix; the second sub-branch reduces the dimension and transforms the local features extracted from the feature map, and after multiplying them with the transposed matrix generated by the third branch, generates the position attention matrix through a normalized exponential function; the third sub-branch reduces the dimension and transforms the local features extracted from the feature map, and multiplies them with the feature matrix of the second sub-branch after transposition; the fourth branch adds the local features extracted from the feature map to the result matrix using a residual connection, and realizes feature fusion through convolutional layers and upsampling.
[0087] It is understandable that in Figure 4 In the embodiment shown, the input local feature A can be subjected to channel dimension reduction to obtain features B and C, that is, the contents of the second branch and the third branch. B and C are converted respectively, and the transposed matrices of the converted B and the converted C are matrix multiplied. The obtained product is calculated using Softmax (normalized exponential function) to generate a position attention matrix S. The position attention matrix S can model the position relationship between any two pixels in the feature.
[0088] In the present invention, the local feature A can also be directly transformed and matrix transposed, and the generated result needs to be multiplied with the aforementioned position attention matrix S to obtain a result matrix. The result matrix is then connected with the local feature A through a residual connection to achieve the addition of each corresponding element, and finally obtain the fusion feature at each position. In the present invention, in order to achieve further refinement of the fusion feature, four convolutional layers can be used to achieve it.
[0089] Preferably, the loss function of the boundary strengthening branch is a binary cross entropy loss function L; and,
[0090]
[0091] Among them, y i is the positive and negative labels of the original marked box of the i-th predicted candidate box, specifically, the positive and negative labels of the original marked box of the small batch sample, with a value of 1 or 0; N is the total number of training samples of the small batch sample of the predicted candidate box
[0092] When the original label box label of the predicted candidate box is positive, y i The value of is 1, and when it is a negative class, the value is 0. In addition, p i The value of y is i In the present invention, a label with a positive value is also called a target true category label.
[0093] Since the boundary reinforcement branch also introduces a certain degree of loss, the present invention adopts a corresponding loss function. The present invention has been verified many times and found that the binary cross entropy loss function has a better result.
[0094] Preferably, the number of convolutional layers in the fourth branch is 4; two layers of convolution are performed on the predicted features of the target boundary generated after feature fusion to generate a predicted target boundary.
[0095] In the present invention, in addition to sending the fused features to the mask branch through certain conversion, the boundary enhancement branch also realizes the generation of the predicted target boundary through further 2 layers of convolution.
[0096] In one embodiment of the present invention, since the feature size of the output of the boundary supervision module is 14*14, in order to obtain a feature map with the same feature size as that in the mask branch and effectively implement element addition, thereby strengthening the feature representation of the target position, especially the feature representation of the boundary, the output of the boundary enhancement branch can be upsampled after four convolutional layers. Similarly, in order to predict the generation of the target boundary, the upsampled output can be restored to a 14*14 feature, and then the target boundary map can be generated.
[0097] Preferably, the mask branch performs 4-layer convolution and upsampling on the local features extracted from the feature map, adds them to the predicted features of the target boundary, and then performs deconvolution to generate a mask.
[0098] The process of generating mask branches is relatively simple, which is achieved through four layers of convolution and upsampling. Since the present invention simultaneously upsamples both the boundary features and the mask features, the boundary-related information in the mask generation process is increased, and the accuracy of the mask boundary is also increased, preventing the problem of rough boundaries.
[0099] After synthesizing the boundary-enhanced features, the mask branch generates the final mask through deconvolution.
[0100] Figure 5 FIG. 4 is a comparison diagram of the target segmentation result achieved by an improved Mask R-CNN target segmentation method in the present invention and the target segmentation result achieved by Mask R-CNN in the prior art. Figure 5 As shown, the above method is applied to the COCO (Common Objects in Context) dataset. The four images in the upper row are target segmentation implemented by the Mask R-CNN method commonly used in the prior art. In the figure, different targets have boundary missing to varying degrees. The images in the lower row are targets extracted according to the method of the present invention. The detection frame can completely cover the entire target and solve the problems of incomplete mask and rough boundary segmentation. The target mask implemented by the present invention is more accurate, more natural and more complete.
[0101] The beneficial effect of the present invention is that, compared with the prior art, an improved target segmentation method of Mask R-CNN in the present invention can generate a detection frame by integrating global features into the features of the region of interest, and by coordinating the loss function under the premise of mutual supervision between the regression branch and the classification branch, and realize the generation of boundaries and masks based on boundary reinforcement. The method of the present invention can extract and enhance the target boundary features in multiple processes of image processing, predict a complete detection frame of the target, thereby avoiding the problems of incomplete mask segmentation and rough target boundaries, and greatly improving the accuracy of target segmentation.
[0102] The applicant of the present invention has made a detailed explanation and description of the implementation examples of the present invention in conjunction with the drawings in the specification. However, those skilled in the art should understand that the above implementation examples are only preferred implementation schemes of the present invention, and the detailed description is only to help readers better understand the spirit of the present invention, but not to limit the scope of protection of the present invention. On the contrary, any improvements or modifications based on the inventive spirit of the present invention should fall within the scope of protection of the present invention.
Claims
1. An improved Mask R-CNN object segmentation method, It is characterized in that The method comprises the following steps: Step 1, using the Mask R-CNN target segmentation method to obtain the ROI region features in the original image, wherein the global context information is subjected to feature conversion using a one-dimensional attention mechanism, and the global features obtained by the conversion are fused into the ROI region features; Step 2: Calculate the classification loss function and regression loss function for the feature map containing the ROI region features, and use the mutually supervised classification loss weight and regression loss weight to construct the coordinated loss function to predict the optimal detection box; Step 3, segmenting the target using the local features extracted by the optimal detection frame, wherein the boundary enhancement method is used to enhance the features of the target boundary, thereby generating a target mask; The boundary reinforcement branch includes a first sub-branch, a second sub-branch, a third sub-branch and a fourth sub-branch; wherein, The first sub-branch transforms and deforms the local features extracted from the feature map, and multiplies the local features with the position attention matrix jointly generated by the second sub-branch and the third sub-branch to obtain a result matrix; The second sub-branch performs dimension reduction and transformation on the local features extracted from the feature map, and generates a position attention matrix through a normalized exponential function after multiplying the local features with the transposed matrix generated by the third sub-branch; The third sub-branch performs dimension reduction and transformation on the local features extracted from the feature map, and multiplies the local features with the feature matrix of the second sub-branch after transposition; The fourth sub-branch adds the local features extracted from the feature map to the result matrix using a residual connection method, and realizes feature fusion through a convolution layer and upsampling.
2. According to the object segmentation method of improving Mask R-CNN described in claim 1, Features: The method for acquiring the ROI region features in the original image in step 1 is: Step 1.1.1, use the residual network ResNetXt-101 and the feature pyramid network FPN to generate multiple feature maps of different scales; Step 1.1.2, using the region candidate network RPN to generate the region of interest; Step 1.2.3, using the feature of interest alignment method ROIAlign to extract the local area features in the region of interest, and using the channel multiplication method to achieve the fusion of the local area features and the global features.
3. According to the object segmentation method of improving Mask R-CNN described in claim 2, Features: The method for obtaining the global features in step 1 is: Step 1.2.1, after reducing the dimension of each of the feature maps of different scales generated in step 1.1.1 by using global average pooling GAP, the reduced dimension information is merged; In step 1.2.2, a lightweight attention mechanism is used to perform feature transformation on the merged dimensionality reduction information to obtain global features.
4. According to the object segmentation method of improving Mask R-CNN described in claim 3, Features: The lightweight attention mechanism is implemented using a one-dimensional convolution with a convolution kernel of 5.
5. According to the object segmentation method of improving Mask R-CNN described in claim 1, Features: The classification branch and the regression branch of the feature map of the ROI region feature are trained separately using a coordinated loss function to predict the optimal detection box.
6. According to the object segmentation method of improving Mask R-CNN as described in claim 5, Features: The coordination loss function is: Among them, i is the sequence number of the predicted candidate box in the small batch sample, p i is the predicted probability of the i-th predicted candidate box, y i is the positive and negative label of the original marked box of the i-th predicted candidate box, d i is the coordinate vector of the i-th predicted candidate box, is the coordinate vector of the original labeled box corresponding to the i-th predicted candidate box, CE() is the cross entropy loss function, L() is the smooth L1 loss function, γ r and γ c They are regression coordination factor and classification coordination factor respectively.
7. According to the object segmentation method of improving Mask R-CNN as described in claim 6, Features: A regression coordination factor is used to achieve mutual supervision between classification loss and regression loss; and, The regression coordination factor is The classification coordination factor is 8. According to the object segmentation method of improving Mask R-CNN described in claim 1, Features: In step 3, the method for generating the target mask is: Step 3.1, using the boundary enhancement branch to generate prediction features of the target boundary, and generating a predicted target boundary based on the prediction features; Step 3.2, adding the predicted features of the target boundary to the mask branch to realize the generation of the target object mask.
9. According to the object segmentation method of improving Mask R-CNN as described in claim 8, Features: The loss function of the boundary strengthening branch is a binary cross entropy loss function L; and, Among them, y i is the positive or negative label of the original marked box of the i-th predicted candidate box, which takes a value of 1 or 0; N is the total number of training samples of the small batch samples of the predicted candidate box.
10. According to the object segmentation method of improving Mask R-CNN as described in claim 9, Features: The number of convolutional layers in the fourth sub-branch is 4; A two-layer convolution is performed on the predicted features of the target boundary generated after feature fusion to generate a predicted target boundary.
11. According to the object segmentation method of improving Mask R-CNN as described in claim 10, Features: The mask branch performs 4-layer convolution and upsampling on the local features extracted from the feature map, adds the local features to the predicted features of the target boundary, performs deconvolution and generates a mask.
Citation Information
Patent Citations
Two-stage hand target detection method
CN112183435A