A Video Salient Object Detection Method Based on Image Annotation
Through the progressive significance target detection network based on image annotation and the spatiotemporal positioning label generation method, the problem of high labeling cost and difficulty in adapting to diverse scenarios in video significance target detection is solved, and efficient video significance target detection is achieved.
Patent Information
- Application Number
- CN202411846050.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-12-16
AI Technical Summary
Existing video significance object detection methods have challenges in reducing labeling costs and adapting to rich scenarios, especially due to the high cost of video pixel-level labeling, which leads to poor performance in diverse scenarios.
A video significance object detection method based on image annotation is proposed. By constructing an incremental significance object detection network, the knowledge in the image significance object dataset is transferred to the video, and the rich scenes are adapted to the spatiotemporal positioning labels and dual-stream positioning networks are generated.
It effectively reduces the annotation cost of video significance object detection data set, improves the ability to adapt to rich scenes, and adds motion information through optical flow images, making the segmented significant object positioning map more accurate.
Smart Images

Figure CN119313885B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video salient object detection, and in particular, to a video salient object detection method based on image annotation. Background Art
[0002] Video salient object detection is an important task in intelligent video content understanding, and it has a wide range of applications in real visual tasks, such as video object segmentation, video classification, video summarization, video surveillance, video compression, autonomous driving, etc. Therefore, video salient object detection has always been one of the research hotspots in the field of computer vision. Different from isolated image salient object detection, it is more challenging to capture video in terms of context association and rich inter-frame dynamic information. Existing methods mainly use recurrent networks or introduce optical flow information to supplement dynamic information. Although great progress has been made in video salient object detection, limited training samples and scene diversity restrict its further development and application. The main reason is the high cost of video pixel-level annotation.
[0003] Although unsupervised methods can completely eliminate the annotation cost by using handcrafted low-level features, the performance of existing methods is poor and they only work well in a few considered scenarios. Some studies have proposed a semi-supervised saliency detection method using sparse labeled frames, which greatly reduces the manual annotation cost, but still requires 20% pixel-level annotation. In addition, some methods design a weakly supervised video salient object detection model by scribbling to label the foreground and background to reduce the burden of pixel-level labeling. Although scribbling annotation can effectively reduce the time and cost compared with pixel-level annotation, it still requires expensive and time-consuming annotation. The main reason is not only the large scale of the video dataset, but also the dependence on high-cost fixation annotation, which is achieved by tracking the eye trajectories of many people. Although these deep learning-based methods can reduce the cost of video annotation, they are difficult to handle large-scale datasets with richer scenarios in the future. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to reduce the annotation cost of the video salient object detection dataset and achieve adaptation to rich scenarios. To overcome the above defects of the prior art (or related technologies), the present invention provides a video salient object detection method based on image annotation.
[0005] The present invention provides a video salient object detection method based on image annotation, including the following steps:
[0006] Step S1: Obtain an image saliency target dataset with pixel-level annotations and an unannotated video saliency target detection dataset, and randomly divide the image saliency target dataset and the video saliency target detection dataset into a training set and a test set after preprocessing;
[0007] Step S2: Construct a progressive saliency target detection network. The progressive saliency target detection network includes a rough localization network, an attention sampler, and a fine segmentation network connected in sequence. Use the rough localization network to locate the position of the salient object in the input image to obtain a main attention map, use the attention sampler to magnify the details of the salient object in the main attention map to obtain an attention magnification map, and use the fine segmentation network to refine the boundary of the salient object and segment the image in the attention magnification map to obtain a salient object localization map;
[0008] Step S3: According to Step S2, use the input image provided by the training set to train the progressive saliency target detection network, and then use the trained rough localization network to identify the salient region of the salient object in the image corresponding to the video frame in the test set to generate a first spatio-temporal localization label;
[0009] Step S4: Obtain adjacent frames of the video frame, and locate the salient region in the image corresponding to the adjacent frame that has the same salient object as the video frame based on the inter-frame difference method to generate a second spatio-temporal localization label;
[0010] Step S5: Input the first spatio-temporal localization label and the second spatio-temporal localization label into a pre-trained two-stream localization network for the localization of video salient objects;
[0011] Step S6: Use the trained attention sampler and fine segmentation network to sequentially identify and segment the image corresponding to the video frame after the video salient object is localized to obtain the corresponding salient object localization map.
[0012] Compared with the prior art, a video salient object detection method based on image annotation in this application has the following advantages:
[0013] A progressive saliency object detection network is proposed in this application. This network adopts a strategy of first localizing and then segmenting, and is composed of a coarse localization network, an attention sampler, and a fine segmentation network. By decoupling the image saliency object dataset with pixel-level annotations, it can efficiently transfer the knowledge learned from images to videos. Utilizing the differences in localization and similarities in segmentation between image saliency object detection and video saliency object detection, it can reduce the annotation cost without expensive video annotations. And a novel method for generating spatio-temporal localization labels is proposed, including generating high-saliency frames and tracking salient objects in adjacent frames, thereby indirectly transferring the rich scene knowledge learned from images to locate salient objects in videos and achieving adaptation to rich scenes. The two-stream localization network can bridge static and dynamic information to predict the saliency object localization map in videos. This network can adapt to the data structure of sparse spatio-temporal localization labels and add important motion information through optical flow images, making the segmented saliency object localization map more accurate.
[0014] In a possible implementation manner, in the step S1, the process of preprocessing the image saliency object dataset and the video saliency object detection dataset is as follows:
[0015] Normalize the pixel values of all pixel points in the R channel of the images in the image saliency object dataset and the video saliency object detection dataset to a mean of 0.485 and a variance of 0.229, normalize the pixel values of all pixel points in the G channel to a mean of 0.456 and a variance of 0.224, and normalize the pixel values of all pixel points in the B channel to a mean of 0.406 and a variance of 0.225. Subsequently, perform data augmentation on the images through horizontal flipping or random rotation processing.
[0016] In a possible implementation manner, the binary dilation operation with a kernel size of is used in the coarse localization network constructed in the step S2 to magnify the input image to obtain the position distribution of the salient object and retain the background information around the salient object. Subsequently, apply Gaussian blur to generate the body attention map using the same kernel size as the binary dilation operation.
[0017] In a possible implementation manner, the attention sampler constructed in the step S2 decomposes the body attention map into two dimensions through the function: and , where w represents the width of the body attention map, h represents the height of the body attention map, represents the pixel value. Subsequently, sample the salient object through the sampling function to obtain the attention magnification map: , where denotes the attention amplification map, denote the inverse function of.
[0018] In a possible implementation, the refined segmentation network constructed in step S2 performs scale uniformization processing and resolution improvement processing on the attention amplification map to refine the boundary of the significant object, and then places the high-resolution significant object in the central area of the image and compresses the background information, and obtains the significant object localization map containing the significant object through segmentation processing.
[0019] In a possible implementation, in step S3, the generation of the first spatio-temporal localization label includes a static saliency prediction branch, a dynamic saliency prediction branch, and a high-saliency frame discriminator. For the static saliency prediction branch, the coarse localization network is used to identify the corresponding image of the video frame to obtain the static saliency region of the significant object. For the dynamic saliency prediction branch, the RAFT model is used to estimate the optical flow of the video frame and perform rendering to obtain an optical flow image, and the coarse localization network is used to identify the optical flow image to obtain the dynamic saliency region of the significant object. Subsequently, the high-saliency frame discriminator is used to determine the similarity degree between the static saliency region and the dynamic saliency region to generate a high-saliency frame and the corresponding video saliency region as the first spatio-temporal localization label.
[0020] In a possible implementation, in step S3, during the process of using the trained coarse localization network to identify the static saliency region of the corresponding image of the video frame and the dynamic saliency region of the optical flow image, when the static saliency region and the dynamic saliency region contain the same object in the same frame, the object is regarded as the significant object.
[0021] In a possible implementation, in step S4, based on the high-saliency frame, multiple adjacent frames of the high-saliency frame are obtained, the spatio-temporal localization maps of the adjacent frames are calculated using the inter-frame difference method, the displacement of the significant object is calculated by obtaining the optical flow information between the adjacent frames through the RAFT optical flow algorithm, and then based on the spatio-temporal localization map and the displacement, remapping processing is performed through the OpenCV software library to obtain the significant object localization maps of the adjacent frames as the second spatio-temporal localization label. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is the flowchart of the steps of the present invention;
[0023] Figure 2 is the schematic diagram of the framework of the progressive saliency object detection network of the present invention;
[0024] Figure 3Schematic diagram of the framework for generating spatio-temporal localization tags of the present invention;
[0025] Figure 4 Schematic diagram of the framework for expanding spatio-temporal localization tags based on the adjacent frame tracking method of the present invention;
[0026] Figure 5 Schematic diagram of the framework of the dual-stream localization network of the present invention;
[0027] Figure 6 Schematic diagram of the comparison result between the method of the present invention and the existing method on the air switch dataset. Detailed implementation manners
[0028] First of all, those skilled in the art should understand that these implementation manners are only used to explain the technical principles of the embodiments of the present application, and are not intended to limit the protection scope of the embodiments of the present application. Those skilled in the art can adjust them according to needs to adapt to specific application scenarios.
[0029] The following further describes the present application in detail with reference to the drawings and specific embodiments.
[0030] See Figure 1 , an embodiment of the present application discloses a video salient object detection method based on image annotation, including:
[0031] Step S1: Obtain an image salient object dataset and an unannotated video salient object detection dataset, and preprocess the images;
[0032] Step S2: Construct a progressive salient object detection network (PAKRN) based on the knowledge review network, and train it with the preprocessed image input;
[0033] Step S3: Use the coarsely localized network trained in step S2 to identify the static salient regions in the video frames and the dynamic salient regions in their optical flow images. When the static salient regions and the dynamic salient regions contain the same object in the same frame, these objects are regarded as video salient objects; by analyzing the similarity between the static and dynamic salient regions, identify the highly salient frames and localize their video salient regions, thereby generating spatio-temporal localization tags;
[0034] Step S4: According to the continuity of the salient objects in adjacent frames, the salient regions containing the same salient objects can be localized through the adjacent frame tracking method, so as to generate richer spatio-temporal localization tags;
[0035] Step S5: Input the spatio-temporal localization tags generated in step S3 and step S4 into the dual-stream localization network to train the localization of video salient objects;
[0036] Step S6: After localizing the salient object in step S5, use the attention sampler to process the video frames, directly perform segmentation using the trained fine segmentation network, and map the segmentation result back to the original format.
[0037] Continue to refer to Figure 1 , the specific steps of step S1 are as follows: Normalize the image saliency object dataset and video saliency object detection dataset of the input image. Normalize the pixel values of all pixel points in the R channel of the image to a mean of 0.485 and a variance of 0.229, the pixel values of all pixel points in the G channel to a mean of 0.456 and a variance of 0.224, and the pixel values of all pixel points in the B channel to a mean of 0.406 and a variance of 0.225. Then, perform data augmentation on the image by horizontally flipping and randomly rotating with a probability of 0.5.
[0038] Refer to Figure 2 , the progressive saliency object detection network based on knowledge review includes a coarse localization network (CLM), an attention sampler, and a fine segmentation network (FSM), which are specifically as follows:
[0039] (1) Coarse localization network
[0040] The goal of the coarse localization network CLM is to find the exact location of the salient object, laying a solid foundation for subsequent accurate segmentation. Its training label is the object attention map, which mainly focuses on position information to guide the network. In order to fully include the salient object and smooth the edges, use a binary dilation operation with a kernel size of to magnify the label. In addition, in order to help the model learn the object position distribution and retain the surrounding background information, apply Gaussian blur. is 8, and use the same kernel size as the binary dilation operation to generate the object attention map. The main idea is to sample the pixels of the original image according to the attention values of the object attention map;
[0041] (2) Attention sampler
[0042] After obtaining the object attention maps, use them to improve the resolution of the regions related to the salient object in the image, thereby magnifying the details of the salient object. The salient object is magnified to be closer to the image size, which will narrow the scale gap between the salient objects in the dataset. This sampler takes the original image and the object attention map as inputs and generates an image that retains the salient object. First, decompose the object attention map into two dimensions through the function: , , where w and h are the width and height of the original image , is the pixel values, and then calculate the inverse function of the distribution function to achieve sampling. The sampling function is defined as follows: , where is 's inverse function. Specifically, the regions with higher attention values are sampled more densely;
[0043] (3) The fine segmentation network is similar to other saliency object detection (SOD) networks in tasks and functions. The FSM also needs to complete the tasks of saliency object localization and segmentation. The biggest difference is that the input images of the FSM have been preprocessed, and the saliency objects in these images have more uniform scales and higher resolutions, so it can effectively help refine the boundaries of saliency objects. In addition, in the processed images, the high-resolution saliency objects are located in the central region, while the background is compressed, thus greatly reducing the difficulty of locating saliency objects. In the last step, the output will be restored to the saliency map corresponding to the original image.
[0044] Continue to refer to Figure 2 , the coarse localization network and the fine segmentation network structures are based on the Knowledge Review Network (KRN). The KRN is based on the U-shaped Feature Pyramid Network (FPNs), uses the pre-trained ResNet-50 as the backbone network, and is a bottom-up and top-down encoder-decoder that can fully combine multi-scale features to obtain rich semantic information. Among them, the Knowledge Review (KR) module is introduced to review the unlearned and diluted information by recombining the finest feature map with the features of each layer. After fusion, there are five groups of feature maps, named and , arranged in ascending order of resolution. Different groups of feature maps retain different degrees of details and semantic information, but their sizes and quantities are different. First, through convolution, compress the channels of and to be the same as . Then, through the upsampling operation, adjust these feature maps to the same size as . Then, through pixel-by-pixel addition and convolution, fuse them with to supplement the diluted and undiscovered important information. To avoid introducing interference information due to the huge difference between the finest feature map and the rough top-level feature map, intermediate supervision is added to guide all feature maps to retain only the useful information related to the saliency object. Next, these four groups of fused feature maps are integrated through the concatenation operation and convolution. The final saliency map will be obtained through Generated by convolution and upsampling operations, the re - fusion of features in the KR module can compensate for missing or un - fused useful information. However, obtaining useful information as efficiently as possible during the feature integration process can further improve the usability of features. Therefore, a simple SA module is added during the feature integration process to improve the efficiency of feature fusion; similar to FPNs, through multiple downsamplings, average pooling, and convolutions to obtain multi - scale feature maps. Then, all feature maps are simply merged and a single convolutional filter is applied. By combining multi - scale features, more comprehensive information can be extracted from different scale spaces to avoid missing important information; in addition, the SA module can further enhance the receptive field of the entire network. It should be noted that KRN is applied in both CLM and FSM, but there are differences between them. Like general saliency object detection (SOD) networks, FSM needs to accurately distinguish salient objects at the pixel level. Therefore, intermediate edge supervision is added to guide the features provided during the encoding process to have clear boundaries.
[0045] Continue to refer to Figure 2 The loss function of body - attention supervision is calculated as follows: , is a variant of the normalized saccade path saliency, is a variant of the Pearson correlation coefficient. Denote the predicted body - attention map as P and the annotated body - attention map as Q. Then, extract the pixels with a value of 255 from Q to construct a new fixation ground - truth map F. This region represents the high - probability region of the salient object. NSS is used to measure the average normalized value of P at the fixation prediction points in the fixation ground - truth map F, emphasizing the importance of these fixation points for saliency detection. The NSS loss function is denoted as . In the task of locating salient objects, the high - value pixels in F need to be focused on. The modified loss function is: , where represents the th pixel, is the number of pixels with a value of 255 in F. The functions and represent the mean and standard deviation of the image respectively. The Pearson correlation coefficient CC is often used to measure the linear correlation degree between two variables: is the covariance of P and Q. The modified is denoted as: . KLD is used to evaluate the similarity between two distributions. A lower KLD value indicates a higher similarity between P and S, thus reflecting better performance of the model. Its calculation process is: , where is the regularization term, and the loss function of the coarse localization network is as follows: , where is the loss function of the -th middle body attention supervision, and are set to 2 and 1; the loss function of the saliency supervision is as follows: , where and are the binary cross-entropy (BCE) loss and the intersection over union (IoU) loss respectively. The loss function of the sampled saliency supervision is the same, except that the corresponding ground truth map is sampled. is the loss function of the edge supervision, which is constructed using the binary cross-entropy loss. The loss function of the FSM is as follows: , where is the -th middle sampled saliency loss, is the -th middle sampled edge supervision loss, , and are set to 2, 1, and 1 respectively. In the joint training of the coarse localization network and the fine segmentation module, the total loss L is as follows: .
[0046] See Figure 3, the generation of spatio-temporal localization tags includes a static saliency prediction branch, a dynamic saliency prediction branch, and a high-saliency frame discriminator. For the static saliency prediction branch, a CLM trained on an image dataset is used to find static salient regions. In the dynamic saliency prediction branch, RAFT is used to estimate the optical flow of video frames and perform rendering. When the motion of an object is different from the background, its region in the rendered optical flow image has distinct colors and clear boundaries. Therefore, the CLM can easily detect dynamic salient regions. After obtaining the predicted body attention position maps of the static frame and the optical flow image, the high-saliency frame discriminator is used to determine whether their salient regions are similar. The present invention designs strict discrimination conditions to ensure the high quality of the generated high-saliency frames. The intersection over union (IoU) is used to measure the similarity between the body attention maps generated from the static frame and the optical flow image. When the IoU value of a certain frame is higher than the threshold of 0.7, this frame is considered a high-saliency frame. To avoid situations with poor optical flow quality, five optical flow images are generated at different frame intervals, thus obtaining five dynamic salient maps. Only one of them needs to meet the above conditions. If the quality of all optical flow images is poor, this frame will not be selected as a high-saliency frame. Therefore, the occurrence of optical flow model failures has little impact on this process. For these high-saliency frames, the static image salient regions predicted by the CLM are used as spatio-temporal localization tags. Since the dynamic salient regions may not contain the complete object due to the local motion of the salient object, while the static salient regions are obtained through the trained CLM and can contain the complete object with a higher probability. Finally, these frames and their corresponding salient object localization maps are used as part of the training data for the subsequent two-stream localization network learning.
[0047] See Figure 4 , given a high-saliency frame and its corresponding spatio-temporal localization map , the spatio-temporal localization map of adjacent frames is calculated using the adjacent frame tracking method, where denotes the first six frames and the last six frames of . Based on the RAFT optical flow algorithm, the between the Optical flow information between them is obtained to acquire the displacement of the salient object. Based on the spatio-temporal localization map and the optical flow image of the high-salience frame, the localization map of the salient object in adjacent frames is calculated through the remapping function of OpenCV. However, this may result in the remapped salient object localization map not containing the complete object, and such spatio-temporal localization labels lack the correct shape of the salient object, which may reduce the performance of subsequent training. To avoid the above problems, a wider region containing the remapped salient object is cropped, and the CLM is reused to calculate its salient object localization map. Specifically, first, the remapped salient object localization map is dilated and eroded, then the bounding box of the connected domain is found, and the size of the bounding box is doubled to include the complete object. According to the coordinates of these bounding boxes, local images containing the salient object are generated by cropping adjacent frames. Finally, the CLM is used to predict the localization maps of these local images and map them back to the original size. In addition, the adjacent frame may be a high-salience frame, or it may happen that different spatio-temporal localization maps are generated from different high-salience frames. For this case, the spatio-temporal localization map is selected according to the proximity principle.
[0048] See Figure 5 , the dual-stream localization network mainly consists of a static feature extraction branch, a dynamic feature extraction branch, and a feature fusion module. The static feature extraction branch and the dynamic feature extraction branch are based on the CLM pre-trained on the image saliency object detection dataset as the backbone network, which can provide multi-scale features with rich semantic information. The input of the static feature extraction branch is the RGB image, while the input of the dynamic feature extraction branch is five optical flow images. To obtain these optical flow images, five images are randomly selected from the first five frames and the last five frames to calculate the optical flow images with the current frame. If only one optical flow frame is input to the optical flow branch, the result of the network is easily affected by the quality of the optical flow image. However, the five optical flow images contain more or less different motion information and can provide richer information. In addition, random selection is similar to data augmentation, which can improve the generalization ability of the model; the feature fusion module not only needs to combine static features and dynamic features to obtain more comprehensive semantic information, but also needs to restore the spatial size of the feature map and predict a high-resolution saliency map by fusing high-level and low-level features. The basic structure of this module is the same as the decoder of FPN, and the features are integrated layer by layer from top to bottom to obtain comprehensive information. The main difference is that a side aggregation (SA) module is introduced after feature fusion at each layer. In the feature fusion module, the output feature maps of the static feature extraction branch and the dynamic feature extraction branch are used as inputs, denoted as and ; since the quality of each input optical flow image in the dynamic feature extraction branch is different, weights are added to each feature map; specifically, the feature map of the last output layer passes through The convolution is compressed into one dimension, and the weights of a certain optical flow image are obtained through average pooling operations. These weights are then multiplied by the feature maps of the corresponding optical flow images. Finally, the feature maps of the five optical flow images are added together as the input of the feature fusion module. In addition, intermediate supervision is added after each fusion to guide the network to retain only the useful information related to significant objects. Feature integration includes four steps. In the first fusion, , , the processed by the SA module and the processed by the SA module are combined through pixel-level summation and convolution. In subsequent series of fusions, the corresponding static features, the corresponding motion features, and the features processed by the SA module from the previous layer are combined through pixel-level summation and convolution. In addition, intermediate supervision is added after each fusion to guide the network to retain only the useful information related to significant objects.
[0049] Continue to refer to Figure 5 . In this embodiment, a combined loss function is used in the loss function supervised by the main body attention map. This loss function is composed of the linear correlation coefficient (CC), the Kullback-Leibler divergence (KLD), and the normalized scanpath saliency (NSS), denoted as: . The loss function of the overall two-stream localization network is: , where represents the loss of the -th intermediate attention supervision.
[0050] Continue to refer to Figure 1 . In order to further verify the feasibility and effectiveness of the method of the present invention, experiments are conducted on the method of the present invention;
[0051] During the experiment, publicly available video saliency object detection datasets are directly selected for training and testing, including ViSal, VOS, FBMS, DAVIS, DAVSOD. Three metrics are used to evaluate and compare different video saliency object detection methods, including MAE, max F-measure, S-measure. The schematic diagram of the comparison results is as shown in Figure 6 . The following Table 1 and Table 2 give the quantitative comparison tables of the evaluation metrics of the method of the present invention and common image saliency object detection methods and video saliency object detection methods on the ViSal, VOS, FBMS, DAVIS, DAVSOD datasets:
[0052] Table 1 Quantitative comparison table of the evaluation metrics of the method of the present invention and common image / video saliency object detection methods on the ViSal, VOS, and DAVSOD datasets
[0053] ;
[0054] Table 2 Quantitative comparison table of the method of the present invention and common image / video saliency object detection methods in terms of evaluation metrics on the FBMS and DAVIS datasets
[0055] ;
[0056] As can be seen from Table 1 and Table 2 above, the method of the present invention has better evaluation metrics.
[0057] In the description of the present application, the descriptions with reference to terms such as "one embodiment", "some embodiments", "in this embodiment", "specific example", or "some examples" mean that the specific features, mechanisms, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, mechanisms, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0058] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A video salient object detection method based on image annotation, characterized in that: The following steps are involved: Step S1, obtaining a pixel-level annotated image salient object dataset and an unannotated video salient object detection dataset, and preprocessing the image salient object dataset and the video salient object detection dataset and randomly dividing them into a training set and a test set; Step S2, constructing a progressive salient object detection network, the progressive salient object detection network comprising a coarse positioning network, an attention sampler and a fine segmentation network connected in sequence, the coarse positioning network performs position positioning of salient objects in the input image to obtain a main attention map, the attention sampler performs detail magnification of salient objects in the main attention map to obtain an attention magnification map, and the fine segmentation network performs boundary refinement and image segmentation of salient objects in the attention magnification map to obtain a salient object positioning map; Step S3, according to step S2, the input image is provided by the training set to train the progressive salient object detection network, and then the trained coarse positioning network is used to identify the static salient area in the image corresponding to the video frame in the test set and the dynamic salient area in its optical flow image, and by analyzing the similarity between the static salient area and the dynamic salient area, the high salient frame is identified and located to the corresponding video salient area, so as to generate a first spatiotemporal positioning label; Step S4, obtaining adjacent frames of the video frame, and locating a salient area having the same salient object as the video frame in the image corresponding to the adjacent frame based on an adjacent frame difference method to generate a second spatiotemporal positioning label; Step S5, inputting the first spatiotemporal positioning tag and the second spatiotemporal positioning tag into a pre-trained dual-stream positioning network to locate salient objects in the video; Step S6, using the trained attention sampler and the fine segmentation network to sequentially identify and segment the images corresponding to the video frames after the salient objects in the video are located to obtain the corresponding salient object location map.
2. The video salient object detection method based on image annotation according to claim 1 is characterized in that: In step S1, the process of preprocessing the image salient object dataset and the video salient object detection dataset is as follows: The pixel values of all pixels in the R channel of the images in the image salient target dataset and the video salient target detection dataset are normalized to a mean of 0.485 and a variance of 0.229, the pixel values of all pixels in the G channel are normalized to a mean of 0.456 and a variance of 0.224, and the pixel values of all pixels in the B channel are normalized to a mean of 0.406 and a variance of 0.225, and then the images are data enhanced by horizontal flipping or random rotation.
3. The video salient object detection method based on image annotation according to claim 1 is characterized in that: The coarse positioning network constructed in step S2 uses a kernel size of The binary dilation operation enlarges the input image to obtain the location distribution of salient objects and preserve the background information around the salient objects, and then applies Gaussian blur using the same kernel size as the binary dilation operation to generate the subject attention map.
4. The video salient object detection method based on image annotation according to claim 1, characterized in that: The attention sampler constructed in step S2 is The function decomposes the subject attention map into two dimensions: and , where w represents the width of the subject attention map, and h represents the height of the subject attention map. Represents the pixel value, and then the sampling function is used to sample the salient object to obtain the attention magnification map: ,in, represents the attention magnification map, express The inverse function of .
5. The video salient object detection method based on image annotation according to claim 1, characterized in that: The fine segmentation network constructed in step S2 performs scale homogenization and resolution improvement processing on the attention magnification map to refine the boundaries of the salient objects, then places the high-resolution salient objects in the central area of the image and compresses the background information, and obtains the salient object positioning map containing the salient objects through segmentation processing.
6. The video salient object detection method based on image annotation according to claim 1, characterized in that: In step S3, the generation of the first spatiotemporal positioning label includes a static saliency prediction branch, a dynamic saliency prediction branch and a high saliency frame discriminator. For the static saliency prediction branch, the coarse positioning network is used to identify the corresponding image of the video frame to obtain a static salient area of the salient object. For the dynamic saliency prediction branch, the RAFT model is used to estimate the optical flow of the video frame and render it to obtain an optical flow image. The coarse positioning network is used to identify the optical flow image to obtain a dynamic salient area of the salient object. Subsequently, the high saliency frame discriminator is used to determine the similarity between the static salient area and the dynamic salient area to generate a high saliency frame and the corresponding video salient area as the first spatiotemporal positioning label.
7. The method for detecting salient objects in a video based on image annotation according to claim 6, characterized in that: In the step S3, in the process of using the trained coarse positioning network to identify the static salient area of the image corresponding to the video frame and the dynamic salient area of the optical flow image, when the static salient area and the dynamic salient area contain the same object in the same frame, the object is regarded as a salient object.
8. The method for detecting salient objects in a video based on image annotation according to claim 6, characterized in that: In the step S4, a plurality of adjacent frames of the high-saliency frame are obtained based on the high-saliency frame, a spatiotemporal positioning map of each of the adjacent frames is calculated using an adjacent frame difference method, optical flow information between each of the adjacent frames is obtained using a RAFT optical flow algorithm to calculate the displacement of the salient object, and then based on the spatiotemporal positioning map and the displacement, a remapping process is performed using the OpenCV software library to obtain the salient object positioning map of each of the adjacent frames as the second spatiotemporal positioning label.
Citation Information
Patent Citations
A video saliency detection method based on candidate area merging
CN106611427A
Infrared video salient target detection method based on deep learning and differential clustering
CN116385752A