Ship segmentation method based on unsupervised region saliency detection

Through the method based on regional unsupervised significance detection, the significance detector is trained and the vessel mask is generated, which solves the problems of high labeling costs and degraded cross-domain segmentation performance in the ship segmentation task in the prior art, and achieves efficient and accurate ship mask generation.

CN115170796BActive Publication Date: 2025-05-13SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210556492.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-20
Publication Date
2025-05-13
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

The prior art requires pixel-level manual labeling in ship segmentation tasks, which is expensive to label, and the segmentation performance is degraded in cross-domain and unknown ship categories.

Method used

Using a method based on regional unsupervised significance detection, by training the significance detector, using significance pseudo-label generation and detector training, the salient areas in each bounding box in the ship image are extracted to generate a ship mask.

Benefits of technology

It realizes the accurate and efficient generation of masks for all ships in the image without pixel-level manual annotation, which improves segmentation performance and has better generalization performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115170796B_ABST
    Figure CN115170796B_ABST
Patent Text Reader

Abstract

The present invention discloses a ship segmentation method based on regional unsupervised saliency detection, comprising: step 1, training to obtain a saliency detector; wherein the training to obtain a saliency detector comprises saliency pseudo-label generation and saliency detector training; step 2, ship image segmentation based on regional saliency detection, using the saliency detector obtained in step 1 to extract the ship mask in each bounding box in the ship image. The saliency detection model of the present invention can adaptively mine saliency information through the correlation between high-level features of the image, has strong robustness, and achieves the most advanced performance in the field. At the same time, the present invention uses the bounding box information of objects in the image to accurately and efficiently generate masks for all ships in the image. Compared with other alternative methods, the segmentation performance has been significantly improved, and has better generalization performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of image processing, and in particular to a ship segmentation method based on regional unsupervised saliency detection. Background Art

[0002] The saliency detection task refers to locating the most attractive objects, i.e., salient objects, from the input image information by analyzing it, and performing pixel-level segmentation on them. This task originates from the human visual attention behavior in cognitive research, that is, humans can quickly shift their attention to the area with the most information in the visual scene. In an image, the location of salient objects is not only affected by the apparent color, but also closely related to the complexity of the background. Existing unsupervised salient object detection algorithms based on deep learning use the results of multiple traditional unsupervised salient object detection algorithms as prior knowledge, and fuse multiple generated saliency maps into a higher quality saliency map by designing complex optimization methods. For example, four traditional methods are first used to obtain multiple saliency prediction results for the same image. Then multiple networks are designed to fit and optimize the results of each method respectively. Finally, the optimized results of the four methods are collected and used together to train a saliency detector.

[0003] It is worth noting that saliency detection cannot determine the category of the target object, so it is significantly different from general semantic segmentation. In some specific tasks, such as ship detection and pedestrian detection, usually only rough annotations are made for the position of each object, such as bounding boxes. Since the bounding box itself has category information, a relatively more accurate mask of the target object can be obtained by performing saliency detection on the image inside the bounding box, thereby guiding the algorithm to extract more discriminative features for the target object. Similar semantic segmentation models can also predict pixel-level labels for ships, but the model is fully supervised, that is, it requires manual pixel-level annotation for training, and the annotation cost is relatively high. Summary of the invention

[0004] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is to provide a ship segmentation method based on regional unsupervised saliency detection to solve the shortcomings of the prior art.

[0005] To achieve the above object, the present invention provides a ship segmentation method based on regional unsupervised saliency detection, comprising:

[0006] Step 1: training a saliency detector; wherein the training to obtain a saliency detector includes generating saliency pseudo labels and training the saliency detector;

[0007] Step 2: Ship image segmentation based on regional saliency detection. The saliency detector obtained in step 1 is used to extract the ship mask in each bounding box in the ship image.

[0008] Furthermore, the saliency pseudo-label generation is specifically as follows:

[0009] Use ResNet-50 pre-trained on the ImageNet dataset as an encoder to extract multi-level features from the original image;

[0010] Adding additional SE modules after the multi-level features to further enhance the multi-level features;

[0011] After fusing the multi-level features together, an SE module is used for enhancement, and finally a feature map is output, where the value of each pixel in the feature map is called an activation value;

[0012] Design an adaptive decision boundary to find potential salient areas from the feature map, simplify the adaptive decision boundary to a threshold, use the global mean of the original image as the threshold, subtract the threshold from the activation values ​​of all pixels in the feature map, and then use the Sigmoid function to transform to between 0 and 1. The transformed result is the generated saliency map;

[0013] The distance between all pixels of the feature map and the global mean is increased to mine more salient information, and the generated salient information results are processed by conditional random fields to obtain salient pseudo labels.

[0014] Furthermore, the saliency pseudo-label is generated using the following formula:

[0015]

[0016] in refers to the activation value of the network at the i-th pixel, and refers to the mean value of the entire image, p i is the final significant prediction, σ is the sigmoid function, according to the output The mean of defines an adaptive decision boundary, and defines the distance from each pixel to the decision boundary as |p i -σ(0)| 2 .

[0017] Furthermore, the saliency detector training is specifically as follows:

[0018] ResNet-50 unsupervised pre-trained on the ImageNet dataset is used as an encoder to extract multi-level features from the original image; the multi-level features are processed by a convolutional layer and a residual attention network to obtain a saliency detector.

[0019] Furthermore, the saliency detector obtained in step 1 is used to extract the ship mask in each bounding box in the image, specifically:

[0020] Using the bounding boxes of all ships in the ship image dataset, extract the salient areas in each bounding box and expand the bounding box to twice the area;

[0021] The enlarged ship image is cut out, the cut out ship image is put into the saliency model, and the saliency result in the original bounding box is cut out as the mask of the ship;

[0022] The masks of multiple ships are stitched onto the same image to obtain the ship mask of the entire image.

[0023] The present invention also provides a ship segmentation system based on regional unsupervised saliency detection, comprising:

[0024] A saliency detector training unit, used for training a saliency detector; wherein the training to obtain a saliency detector includes saliency pseudo-label generation and saliency detector training;

[0025] The saliency detection ship image segmentation unit uses the saliency detector to extract the ship mask in each boundary box in the ship image.

[0026] The beneficial effects of the present invention are:

[0027] The saliency detection model of the present invention can adaptively mine saliency information through the correlation between high-level features of the image, has strong robustness, and achieves the most advanced performance in the field. At the same time, the present invention uses the bounding box information of the objects in the image to accurately and efficiently generate masks for all ships in the image. Compared with other alternative methods, the segmentation performance has been greatly improved and has better generalization performance.

[0028] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the accompanying drawings to fully understand the purpose, characteristics and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is an overall step diagram of the present invention.

[0030] Figure 2 It is the overall framework diagram of the unsupervised saliency detection model of the present invention.

[0031] Figure 3 It is a flow chart of ship segmentation based on regional saliency of the present invention.

[0032] Figure 4 It is a comparison chart of the results of the present invention and the mainstream significance model.

[0033] Figure 5 This is a comparison diagram of the present invention and other methods.

[0034] Figure 6 It is a table diagram of the comparison results between the present invention and the mainstream method. DETAILED DESCRIPTION

[0035] The present invention provides a ship segmentation method based on regional unsupervised saliency detection, the overall steps are as follows: Figure 1 .

[0036] Step 1: Propose an unsupervised saliency detection framework and train a saliency detector.

[0037] Step 2: Propose a ship segmentation method based on region saliency detection. Using the saliency detector obtained in step 1, extract the ship mask in each bounding box in the image.

[0038] The specific process of each step is as follows:

[0039] Step 1: Unsupervised saliency detection framework.

[0040] The saliency model of the present invention is a two-stage framework from activation to saliency (A2S). Figure 2 The overall framework of the unsupervised saliency detection model is presented, where stage 1 is used for saliency pseudo-label generation and stage 2 is used for saliency detector training.

[0041] Phase 1: Use ResNet-50 pre-trained on the ImageNet dataset as an encoder to extract multi-level features from the image. In addition, add additional SE after each feature

[0042] The Squeeze-and-Excitation (Squeeze-and-Excitation) module is used to further enhance these features. Multiple features are fused together and then enhanced using an SE module. Finally, the network outputs a feature map, and the value of each pixel in the feature map is called an activation value.

[0043] After obtaining the feature map output by the network, an adaptive decision boundary is designed to find potential salient areas from the feature map. Since the network output is a single feature map, for the saliency detection task, the adaptive decision boundary can be simplified to a threshold, and the mean of the entire image is used as the threshold in the implementation. The activation value of all pixels is subtracted from this mean, and then the Sigmoid function is used to transform it to between 0 and 1. The transformed result is the generated saliency map. However, some images may not be able to extract accurate saliency information by themselves. Therefore, the supervisory signal increases the distance from all pixels to the global mean to mine more saliency information. The formula can be expressed as:

[0044]

[0045] in refers to the activation value of the network at the i-th pixel, and refers to the mean value of the entire image. i is the final significant prediction, and σ is the sigmoid function. The mean of defines an adaptive decision boundary, and defines the distance from each pixel to the decision boundary as |p i -σ(0)| 2 The objective function mines more salient information from the image by increasing the distance from all pixels to the decision boundary. Finally, the generated salient results are post-processed with the Conditional Random Field (CRF) to obtain salient pseudo labels.

[0046] Phase 2: A saliency detector is constructed and trained using the pseudo labels generated in Phase 1. The detector also uses the ResNet-50 pre-trained on the ImageNet dataset as an encoder. After extracting multi-level features, the encoder fuses the features of the lower levels in pairs and multiplies them to obtain the final output result.

[0047] Step 2: Ship segmentation method based on region saliency detection.

[0048] For ship images, the most direct way to obtain ship masks is to train a segmentation model. However, training such models usually requires a large amount of pixel-level manual annotation to achieve sufficiently generalized performance. In addition, since existing semantic segmentation datasets usually also have the ship category, semantic segmentation models trained on these datasets can also be used. However, due to the differences in ship categories in existing datasets and ship datasets, as well as differences in image styles such as image contrast and resolution, the results of directly using these models on ship datasets will be significantly reduced.

[0049] In view of this, the present invention extracts the salient areas in each box based on the bounding boxes of all ships in the ship dataset, and combines them into the segmentation result of the entire image. Specifically, first, in order to avoid the salient object area being too large, which causes the saliency model to only focus on the more detailed areas, the bounding box is expanded to twice the area, and then the expanded image is cut out. Subsequently, the cut image is placed in the saliency model, and the saliency result in the original bounding box is cut out as the mask of the ship. Finally, the masks of multiple ships are spliced ​​onto the same image to obtain the ship mask of the entire image. The specific process is shown in Figure 3 .

[0050] This embodiment illustrates the effect of the method through experiments. The training set in the experiment uses the training subset of MSRA-B, which contains a total of 3000 images. Six commonly used saliency detection datasets are used in the test, including DUTS-TE, PASCAL-S, DUT-OMRON, HKU-IS, ECSSD and the test subsets in MSRA-B, which contain 5019, 850, 5168, 4447, 1000 and 2000 images and pixel-level saliency annotations respectively. The experimental evaluation indicators use F-measure (F) and mean absolute error (MAE).

[0051] This example compares the method of the present invention with some existing mainstream deep learning-based methods. The comparison results on six test sets are as follows: Figure 6 As shown, (* indicates the use of supervised encoders). Due to the unsupervised setting, the weights of all networks are initialized using unsupervised pre-trained ResNet-50. Since the existing methods all use supervised pre-trained models, the results of the methods using supervised pre-trained ResNet-50 are also reported (with *). Figure 6 The results show that the average F-measure of the present invention on six commonly used test sets is 0.880, 0.886, 0.683, 0.775, 0.714, 0.856, and the MAE is 0.060, 0.043, 0.082, 0.106, 0.076, 0.048, respectively. Most indicators are ahead of the existing unsupervised salient object detection methods. At the same time, compared with the salient object detection algorithm using pixel-level manual annotation, the present method also achieves competitive results.

[0052] At the same time, in order to intuitively demonstrate the superiority of the present invention, the salient object graphs of the present invention and the mainstream method are compared, such as Figure 4 As shown. Figure 4From the results in , we can see that our method can locate salient objects more accurately and even achieve better results than the supervised algorithm. For example, in the results in row 7, we can see that the supervised algorithm will also consider the tree trunk as a salient object, while our algorithm can more accurately determine the salient object and segment its edges.

[0053] For the ship segmentation task, 100 images from the ship dataset mcship were randomly selected for manual annotation to obtain the ship mask of each image. The results of the latest fully supervised semantic segmentation algorithm Segformer and the present invention are compared as shown in the following table:

[0054]

[0055] Due to the influence of cross-domain and unknown ship categories, the results of existing fully supervised semantic segmentation algorithms are relatively poor. However, this paper only needs the bounding box of the ship to obtain a more accurate mask. Figure 5 shown.

[0056] It can be seen that due to the problem of image style, the pre-trained semantic segmentation model can only segment part of the ship area. Applying the saliency model directly to the image will lead to additional attention to other irrelevant areas. The regional saliency method proposed in the present invention can achieve better performance than these methods.

[0057] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.

Claims

1. A ship segmentation method based on unsupervised saliency detection of regions, characterized in that: include: Step 1: training a saliency detector; wherein the training to obtain a saliency detector includes generating saliency pseudo labels and training the saliency detector; Step 2: Ship image segmentation based on regional saliency detection, using the saliency detector obtained in step 1 to extract the ship mask in each bounding box in the ship image; The saliency pseudo-label generation is specifically as follows: Use ResNet-50 pre-trained on the ImageNet dataset as an encoder to extract multi-level features from the original image; Adding additional SE modules after the multi-level features to further enhance the multi-level features; After fusing the multi-level features together, an SE module is used for enhancement, and finally a feature map is output, where the value of each pixel in the feature map is called an activation value; Design an adaptive decision boundary to find potential salient areas from the feature map, simplify the adaptive decision boundary to a threshold, use the global mean of the original image as the threshold, subtract the threshold from the activation values ​​of all pixels in the feature map, and then use the Sigmoid function to transform to between 0 and 1. The transformed result is the generated saliency map; Increasing the distance between all pixels of the feature map and the global mean, mining more saliency information, and obtaining saliency pseudo labels by post-processing the generated saliency information results with conditional random fields; The saliency pseudo-label generation adopts the following formula: in refers to the activation value of the network at the i-th pixel, and refers to the mean value of the entire image, p i is the final significant prediction, σ is the sigmoid function, according to the output The mean of defines an adaptive decision boundary, and defines the distance from each pixel to the decision boundary as |p i -σ(0)| 2 ; The saliency detector training is specifically as follows: Use ResNet-50 pre-trained on the ImageNet dataset as an encoder to extract multi-level features from the original image; After the multi-level features are processed by a convolutional layer and a residual attention network respectively, a saliency detector is obtained.

2. A ship segmentation method based on region unsupervised saliency detection as claimed in claim 1, characterized in that: The saliency detector obtained in step 1 is used to extract the ship mask in each bounding box in the image, specifically: Based on the bounding boxes of all ships in the ship image dataset, the bounding boxes are expanded to twice the area, and the salient areas in each bounding box are extracted; The enlarged ship image is cut out, the cut out ship image is put into the saliency model, and the saliency result in the original bounding box is cut out as the mask of the ship; The masks of multiple ships are stitched onto the same image to obtain the ship mask of the entire image.

3. A ship segmentation system based on unsupervised saliency detection of regions, characterized in that: include: A saliency detector training unit, used for training a saliency detector; wherein the training to obtain a saliency detector includes saliency pseudo-label generation and saliency detector training; A saliency detection ship image segmentation unit, using the saliency detector, extracts the ship mask in each bounding box in the ship image; The saliency pseudo-label generation is specifically as follows: Use ResNet-50 pre-trained on the ImageNet dataset as an encoder to extract multi-level features from the original image; Adding additional SE modules after the multi-level features to further enhance the multi-level features; After fusing the multi-level features together, an SE module is used for enhancement, and finally a feature map is output, where the value of each pixel in the feature map is called an activation value; Design an adaptive decision boundary to find potential salient areas from the feature map, simplify the adaptive decision boundary to a threshold, use the global mean of the original image as the threshold, subtract the threshold from the activation values ​​of all pixels in the feature map, and then use the Sigmoid function to transform to between 0 and 1. The transformed result is the generated saliency map; Increasing the distance between all pixels of the feature map and the global mean, mining more saliency information, and obtaining saliency pseudo labels by post-processing the generated saliency information results with conditional random fields; The saliency pseudo-label generation adopts the following formula: in refers to the activation value of the network at the i-th pixel, and refers to the mean value of the entire image, p i is the final significant prediction, σ is the sigmoid function, according to the output The mean of defines an adaptive decision boundary, and defines the distance from each pixel to the decision boundary as |p i -σ(0)| 2 ; The saliency detector training is specifically as follows: Use ResNet-50 pre-trained on the ImageNet dataset as an encoder to extract multi-level features from the original image; After the multi-level features are processed by a convolutional layer and a residual attention network respectively, a saliency detector is obtained.

Citation Information

Patent Citations

  • Method for realizing weak supervision image saliency detection by using detection frame

    CN111680702A

  • Video salient object detection method based on confidence degree self-adaption and difference enhancement

    CN112784745A