Memory-edge guided weakly supervised video salient object detection method and system
By constructing a multi-scale feature encoder and a memory pool module, combined with an edge-guided dual-branch decoder, the problem of target edge capture under complex spatiotemporal variations in video salient target detection is solved, achieving efficient salient target detection and improved boundary accuracy.
Patent Information
- Application Number
- CN202511053022.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-07-30
AI Technical Summary
Existing video salient object detection technologies struggle to accurately capture object edges when faced with complex, spatiotemporally dynamic video data. In particular, deep learning-based methods rely on densely labeled data, resulting in high costs, while sparse labeling in weakly supervised methods leads to unsatisfactory detection results.
A multi-scale feature encoder and a multi-scale memory pool module are constructed. Combined with an edge-guided dual-branch decoder, spatiotemporal features are extracted through Video Swin Transformer. Multi-scale self-attention mechanism and dynamic storage of historical frame features are utilized, along with edge fitting loss that minimizes distance, to optimize salient object detection.
It improves the accuracy and boundary precision of salient target detection, effectively suppresses background noise interference, reduces dependence on dense annotation, and enhances detection performance.
Smart Images

Figure CN120564108B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a weakly supervised video saliency target detection method and system based on memory-edge guidance, belonging to the field of target detection technology. Background Technology
[0002] In recent years, with the booming development of artificial intelligence technology and the multimedia industry, research on video content analysis and understanding has attracted much attention. The demand for these technologies in fields such as intelligent surveillance, autonomous driving, video editing, and virtual reality has experienced explosive growth.
[0003] Video salient object detection (VSOD) is a classic computer vision task. Its core objective is to accurately locate salient objects from consecutive video frames with complex dynamic backgrounds by fusing spatiotemporal features to fully utilize spatial cues from static images, fine edge information of objects, temporal correlations, and inherent motion patterns in video sequences. These salient objects can be dynamic elements such as moving objects, unusual behaviors, or key events. Therefore, VSOD technology, as a fundamental preprocessing step, is widely used in advanced video analytics applications such as object tracking, video object segmentation, video compression, and video summarization.
[0004] Existing video salient object detection technologies mainly fall into two categories: 1) Detection techniques based on traditional handcrafted features, which utilize static spatial features and optical flow motion features, designing specific rules to extract motion and appearance features for object detection. However, these methods rely on prior knowledge from human hands, making it difficult to adapt to complex spatiotemporal dynamic changes in videos, and are sensitive to background edge noise interference; 2) Detection techniques based on deep learning, which utilize end-to-end network models such as Convolutional Neural Networks (CNNs) to directly learn spatiotemporal features from large amounts of data, eliminating the need for complex handcrafted features. However, these methods improve detection accuracy through data-driven approaches, and their performance is highly dependent on a large amount of densely labeled ground truth data (i.e., pixel-by-pixel labeled fully supervised data). The cost of dense labeling of video data is extremely high, severely limiting its practical application. The resulting weakly supervised detection methods, on the other hand, only require simple weakly annotated data (such as keypoint annotations, image-level annotations, and graffiti annotations), significantly reducing annotation costs while improving applicability. However, because videos often feature complex spatiotemporal dynamics and encompass a wide range of visual information features, while the temporal supervision provided by weakly supervised labels is sparse and incomplete, it is difficult to accurately capture the edges of targets. This hinders deep models from fully utilizing spatiotemporal context information and fine-grained edge details of salient objects, making it difficult for current technologies to achieve ideal detection results. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a weakly supervised video salient target detection method based on memory-edge guidance; it predicts salient targets in a given series of video frames.
[0006] This invention also provides a weakly supervised video salient object detection system based on memory-edge guidance.
[0007] This invention first constructs a multi-scale feature encoder, primarily utilizing a multi-stage Video Swin Transformer to extract spatiotemporal features from several consecutive video frames in a long video sequence at different scales. Secondly, this invention proposes a multi-scale memory pool (MSMP) module for dynamically storing historical frame features. It uses a parallel multi-scale self-attention mechanism to generate multi-scale historical memory features and guides deep fusion with current frame features through parallel multi-scale cross-attention. Finally, this invention designs a dual-branch decoder. The SaliencyPrediction Branch (SPB) progressively upsamples and concatenates features to generate an initial saliency map, while the Edge Enhancement Branch (EEB) fuses three shallow visual features to learn fine-grained edge features. In both branches, a distance-minimized edge fitting loss is proposed to jointly optimize the geometric consistency between the predicted saliency edges and the true edges of the object, while suppressing cluttered background edge noise, thereby improving object integrity and boundary accuracy.
[0008] The technical solution of this invention is as follows:
[0009] A memory-edge guided weakly supervised video salient object detection method includes:
[0010] The given long video sequence is divided into non-overlapping sliding windows to extract several consecutive video frames. These video frames are then input into the trained memory edge-guided network model to achieve weakly supervised video salient object detection.
[0011] Inputting the data into a trained memory edge-guided network model enables salient object detection in weakly supervised videos; specifically including:
[0012] Feature extraction is performed on consecutive video frames at the current time step to obtain spatiotemporal features extracted at different scales;
[0013] The semantic representation of relevant objects in the current frame is enhanced by using salient cues mined from historical frames. The cues are then input into the decoder for decoding to obtain the final salient object detection map.
[0014] During the training of the memory edge-guided network model, the edge fitting loss is used to align with the true boundaries of salient objects while suppressing the interference of background edge noise.
[0015] As a further preferred option, the memory edge-guided network model includes a multi-scale encoder, a multi-scale memory pool module, and an edge-guided dual-branch decoder;
[0016] A multi-scale encoder is used to extract features from consecutive video frames at the current time step to obtain spatiotemporal features extracted at different scales.
[0017] The spatiotemporal features of historical frames are dynamically stored using a multi-scale memory pool module, and the semantic representation of relevant objects in the current frame is enhanced by using salient cues mined from historical frames.
[0018] The final salient object detection map is obtained through an edge-guided dual-branch decoder.
[0019] As a further preferred approach, the specific steps for extracting spatiotemporal features from a given series of video frames using a multi-scale encoder include:
[0020] The given long video sequence is divided into four consecutive non-overlapping sliding windows of length 4, and four consecutive video frames are extracted as the encoding input for each time step.
[0021] A four-stage Video Swin Transformer network model was used to capture information at different semantic levels in video frames, obtaining spatiotemporal features at four different scales.
[0022] As a further preferred approach, a multi-scale memory pool module is used to effectively mine salient cues in historical frames to enhance the semantic representation of salient objects in the current video frame. Specific steps include:
[0023] The memory pool is designed to store the spatiotemporal features of four scales extracted from several historical frames in the past.
[0024] After the video frames at each time step have their features extracted by the multi-scale encoder, they enter the memory pool. If the memory pool is full, the oldest historical frame is replaced to achieve dynamic storage.
[0025] A self-attention mechanism is applied to the features of historical frames at each scale to obtain parallel multi-scale spatiotemporal context information.
[0026] By leveraging multi-scale spatiotemporal context information across four scales, and through parallel multi-scale cross-attention guidance and deep fusion with current frame features, the semantic representation of salient objects in the current video frame is enhanced.
[0027] As a further preferred approach, a multi-scale memory pool module is used to effectively mine salient cues in historical frames to enhance the semantic representation of salient objects in the current video frame. Specific steps include:
[0028] S21 The multi-scale features extracted at this time step are dynamically represented in a highly generalized manner:
[0029] ;
[0030] in, This represents a four-stage Video Swin Transformer network model. This represents the four consecutive video frames input at this time step. This represents the spatiotemporal features obtained at four different scales; The subscript indicates the scale, and the superscript indicates the video frame number. X i Represents the spatiotemporal characteristics of a single historical video frame;
[0031] S22 : The spatiotemporal features extracted at this time step Updated into the memory pool; the features of each historical video frame in the memory pool are spatiotemporal features, denoted as... If the memory pool is full, the oldest historical frame features are removed and the new video frame features are updated.
[0032] S23 The salient cues of historical frame features are extracted using a parallel self-attention mechanism at four scales, as shown below:
[0033] ;
[0034] ;
[0035] ;
[0036] in, For the spliced first Features at each scale The feature dimension of each video frame at this scale. For splicing operations, For the first in the memory pool The first historical video frame Features at each scale for Features after dimensional reshaping For the tensor dimension reshaping function, To integrate all historical information and emphasize the new features obtained from key spatiotemporal patterns, express function, These are three independent linear transformations;
[0037] S24 The spatiotemporal feature representation of the current time step is guided by salient cues from historical frames, as shown below: ;
[0038] ;
[0039] ;
[0040] in, This indicates that in the current time step, there are 4 consecutive video frames at the 1st... Feature representation at each scale express Features after dimensional reshaping The dimension is ,Right now ,express Features after dimensional reshaping For use The first one obtained after augmentation representation New features at each scale.
[0041] As a further preferred embodiment, an edge-guided dual-branch decoder is used to accurately capture the boundaries of salient objects and generate the final salient object detection map; the edge-guided dual-branch decoder includes a saliency prediction branch and an edge enhancement branch; including:
[0042] The output features, after being enhanced by the spatiotemporal context information of historical frames, are input into two branches of the decoder for decoding.
[0043] We utilize an edge enhancement branch to fuse three shallow visual features and align them with the true boundaries of prominent objects, while suppressing cluttered background edge noise.
[0044] The features are progressively upsampled and concatenated using a saliency prediction branch, and an initial saliency feature map is obtained through a residual channel attention mechanism.
[0045] The saliency features obtained from the saliency prediction branch and the edge enhancement branch are fused together, and the final saliency map is obtained through the residual channel attention mechanism.
[0046] As a further preferred solution, an edge-guided dual-branch decoder is used to accurately capture the boundaries of salient objects and generate the final salient object detection map.
[0047] The edge-guided dual-branch decoder includes a saliency prediction branch and an edge enhancement branch. The saliency prediction branch is responsible for generating the initial saliency map, while the edge enhancement branch is responsible for learning the edge features of salient objects. The edge fitting loss based on distance minimization is introduced into both the saliency prediction branch and the edge enhancement branch to supervise the learning of the edge features of salient objects. At the same time, the saliency loss and the gated structure perception loss are combined to jointly supervise the learning of the memory edge-guided network model.
[0048] The output features are enhanced by historical frame spatiotemporal context information. Decoding is performed using a dual-branch decoder; the specific implementation steps include:
[0049] S31 First, the edge enhancement branch applies to three shallow visual features. , , Feature transformation is performed using convolutional modules respectively. express The part in the middle represents the feature of the first video frame, the subscript number represents the scale, and the superscript number represents the video frame number;
[0050] Then, an upsampling transformation is performed to the same feature dimension. Next, the three transformed features are concatenated, and then channel enhancement is performed through the residual channel attention module to obtain the edge feature map of the salient object. As shown below:
[0051] ;
[0052] ;
[0053] ;
[0054] in, This represents the output feature at the j-th scale after enhancement based on the spatiotemporal context information of historical frames. For convolutional modules, For upsampling operation, The features are those after convolution and upsampling. This indicates that features are concatenated along the channel dimension. This represents the feature obtained by concatenating features from the three scales along the channel dimension. For residual channel attention module, This represents the generated edge feature map;
[0055] S32 In the saliency prediction branch, from the deepest features... The process begins with sequential upsampling and concatenation, followed by enhancement of channel representations using a residual channel attention module to generate initial saliency feature maps. , This refers to the features at the fourth scale of each video frame; as shown below:
[0056] ;
[0057] ;
[0058] ;
[0059] ;
[0060] in, This indicates that features are concatenated along the channel dimension. This represents the intermediate features after the first convolution calculation and concatenation. This represents the intermediate features after the second convolution calculation and concatenation. This represents the intermediate features after the third convolution calculation and concatenation. For residual channel attention module, This represents the generated initial salient feature map;
[0061] S33 The feature maps generated by fusing the saliency prediction branch and the edge enhancement branch are used to accurately predict the final saliency map. :
[0062] .
[0063] As a further preferred option, a distance-minimization-based edge fitting method is used for supervised, accurate boundary learning of salient objects; including:
[0064] First, two feature maps are generated through dual-branch decoding. and The predicted edge map is extracted using the Sobel operator, denoted as . and ;
[0065] Next, the RCF edge detection network is used to generate a precise edge map from the original image, denoted as . The accurate edge map includes accurate object boundaries while introducing background edge noise;
[0066] Calculate separately and Each edge pixel in the middle arrive The middle corresponds to the nearest edge pixel The mean Euclidean distance;
[0067] Pre-processing edge maps Calculate the distance map and Each pixel value represents the distance from that location to the nearest edge point; including:
[0068] Through two predicted edge maps and Corresponding distance map and The dot product operation is used to efficiently solve for the loss;
[0069] Edge fitting loss of saliency prediction branch and edge enhancement branch and The definition is as follows:
[0070] ;
[0071] ;
[0072] in, This indicates the calculation of the Euclidean distance between two pixels. For edge map The number of pixels at the middle edge. For edge map The number of pixels at the middle edge. and To predict the edge graph of the first Each edge pixel to The distance to the nearest edge point; the subscript EEB indicates the edge enhancement branch, SPB indicates the significance prediction branch, and DEL indicates the edge fitting loss. This represents the edge fitting loss of the edge enhancement branch. The marginal fit loss represents the significance prediction branch. This represents the predicted edge map extracted from the feature map output from the EEB branch using the Sobel operator. This represents the predicted edge map extracted from the initial feature map output from the SPB branch using the Sobel operator. Represents the first edge in the predicted edge graph One edge pixel, Represents distance in the RCF edge graph The nearest edge pixel;
[0073] During the training of the memory edge-guided network model, a multi-supervision mechanism is adopted, and edge loss is introduced. and The loss includes saliency loss and gated structure perception loss; among which, the saliency loss is used to calculate the predicted saliency feature map. Weak labeling of truth values Pixel-level cross-entropy loss , The final salient feature map of the gated structure-aware loss supervises the structural consistency between the output and the input image within the salient region; saliency loss, gated structure-aware loss, and the overall loss of the memory edge-guided network model. The design is as follows:
[0074] ;
[0075] ;
[0076] ;
[0077] in, Indicating the first true value weak annotation 1 pixel, The first saliency in the final salient feature map 1 pixel, Indicates the total number of pixels. Represents pixel coordinates. Represents the gradient direction. The coordinates in the final salient feature map are: pixels, The final salient feature map is in Gradient of direction, Represents image intensity. Represents the gradient of image intensity in the direction. Represents edge-aware weights. This represents a smoothing function.
[0078] A memory-edge guided weakly supervised video salient object detection system includes:
[0079] The non-overlapping sliding window partitioning module is configured to partition a given long video sequence into a non-overlapping sliding window and extract several consecutive video frames.
[0080] The salient object detection module is configured to: input several video frames into a trained memory edge-guided network model to achieve weakly supervised video salient object detection; specifically including:
[0081] Feature extraction is performed on consecutive video frames at the current time step to obtain spatiotemporal features extracted at different scales;
[0082] The semantic representation of relevant objects in the current frame is enhanced by using salient cues mined from historical frames. The cues are then input into the decoder for decoding to obtain the final salient object detection map.
[0083] During the training of the memory edge-guided network model, the edge fitting loss is used to align with the true boundaries of salient objects while suppressing the interference of background edge noise.
[0084] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0085] 1. The multi-scale encoder constructed in this invention captures local detail information and global information simultaneously by deploying a Video Swin Transformer network, and captures feature information at different scales to cover information at different semantic levels, thereby achieving accurate localization and fine segmentation of salient targets.
[0086] 2. This invention proposes a multi-scale memory pool module, which utilizes an innovative dynamic storage strategy and a parallel multi-scale self-attention mechanism to mine useful salient clues from historical frames, thereby effectively mitigating the performance degradation caused by the sparsity and incompleteness of weak annotations.
[0087] 3. The edge-guided dual-branch decoder based on distance minimization proposed in this invention innovatively designs an edge enhancement branch and proposes a new edge fitting loss supervision method based on distance minimization, which effectively aligns with the true boundary of salient objects while suppressing the interference of background edge noise. Attached Figure Description
[0088] Figure 1 This is a diagram of the overall framework of the model;
[0089] Figure 2 This is a schematic diagram of a multi-scale memory pool module;
[0090] Figure 3 This is a schematic diagram of a dual-branch video decoder;
[0091] Figure 4 This is a schematic diagram of the edge fitting loss based on distance minimization proposed in this invention;
[0092] Figure 5 This is a schematic diagram comparing the renderings of the present invention with those of other existing models;
[0093] Figure 6 This is an architecture diagram of the four-stage Video Swin Transformer network model.
[0094] Figure 7 This is an architecture diagram of the residual channel attention module;
[0095] Figure 8 This is a diagram of the RCF edge detection network architecture. Detailed Implementation
[0096] The present invention will be further defined below with reference to the accompanying drawings and embodiments, but is not limited thereto.
[0097] Example 1
[0098] Terminology Explanation:
[0099] 1. Four-Stage Video Swin Transformer Network Architecture: The four-stage Video Swin Transformer network is an existing model used for spatiotemporal feature extraction from input video. This model employs a hierarchical feature extraction strategy, processing several consecutive video frames in four stages, progressively reducing the feature map resolution while increasing the number of channels, thereby capturing spatial semantic information and temporal dependencies at different levels. For example... Figure 6 As shown, Figure 6 In this model, T represents the number of video frames, H×W represents the resolution, and C represents the number of channels. The specific model structure is shown below: after the input video frames are divided into 3D blocks, feature extraction is performed sequentially through four stages. Within each stage, the feature resolution of the input video frames is first reduced, while the number of channels is increased (in the first stage, each block is linearly embedded after 3D segmentation to reduce the resolution of the original input video frames and increase the number of channels; the subsequent three stages achieve this through block merging). Then, spatiotemporal feature extraction is performed through several layers of sliding Transformer blocks. Each video sliding Transformer block sequentially performs layer normalization, multi-head attention computation within a 3D window (dividing the input feature map into regular windows, each containing several blocks, and calculating multi-head attention within each window to capture local spatiotemporal feature dependencies within the window), layer normalization, multi-layer perceptron, layer normalization, multi-head self-attention computation within a 3D sliding window (sliding the regular window from the second step, calculating multi-head self-attention within each new window to capture spatiotemporal dependencies between windows and introduce global feature interactions), layer normalization, and multi-layer perceptron, and adds multiple residual connections (each in the figure below) to the input features. Each represents a residual connection, which involves adding the transformed features to the original features (without the transformation) to obtain the spatiotemporal features output by this layer. The method for obtaining features at four different scales is as follows: the features at each scale correspond to the features output at each stage.
[0100] 2. Edge detection network RCF, such as Figure 8The image shown is an existing edge detection network (from the paper "Richerconvolutional features for edge detection," published in IEEE Transactions on Pattern Analysis and Machine Intelligence). This model can perform feature recognition on input RGB images and then output an edge map. The overall model structure is modified from the VGG model. By fusing the features output from each convolutional layer, it fully utilizes semantic feature information and detail information at different levels to detect edge information, specifically for extracting features from the input image and drawing an edge map. The model structure is shown in the figure below, consisting of five stages of convolution (one color represents one stage, and each stage contains several convolution calculations). In the model, a form like "3*3-128 convolution" indicates that the convolution kernel size is 3*3, and the number of feature channels output after calculation is 128. All loss calculations are performed only during training, and vector concatenation indicates concatenation along the channel dimension. Accurate edge map generation: Using a pre-trained RCF model, the original image is input into the network, and after a series of convolution and pooling calculations, the edge map is obtained.
[0101] A memory-edge guided weakly supervised video salient object detection method includes:
[0102] The given long video sequence is divided into non-overlapping sliding windows to extract several consecutive video frames. These video frames are then input into the trained Memory-Edge Guided Network (MEGNet) model to achieve weakly supervised video salient object detection.
[0103] Inputting the data into a trained memory edge-guided network model enables salient object detection in weakly supervised videos; specifically including:
[0104] Feature extraction is performed on consecutive video frames at the current time step to obtain spatiotemporal features extracted at different scales;
[0105] The semantic representation of relevant objects in the current frame is enhanced by using salient cues mined from historical frames. The cues are then input into the decoder for decoding to obtain the final salient object detection map.
[0106] During the training of the memory edge-guided network model, the edge fitting loss is used to align with the true boundaries of salient objects while suppressing the interference of background edge noise.
[0107] Example 2
[0108] The difference between this and the memory-edge guided weakly supervised video salient object detection method described in Example 1 is as follows:
[0109] like Figure 1 As shown, the memory edge-guided network model includes a multi-scale encoder, a multi-scale memory pool module, and an edge-guided dual-branch decoder;
[0110] A multi-scale encoder is used to extract features from consecutive video frames at the current time step to obtain spatiotemporal features extracted at different scales.
[0111] The spatiotemporal features of historical frames are dynamically stored using a multi-scale memory pool module, and the semantic representation of relevant objects in the current frame is enhanced by using salient cues mined from historical frames.
[0112] The final salient object detection map is obtained through an edge-guided dual-branch decoder.
[0113] The specific steps for extracting spatiotemporal features from a given series of video frames using a multi-scale encoder include:
[0114] The given long video sequence is divided into four consecutive non-overlapping sliding windows of length 4, and four consecutive video frames are extracted as the encoding input for each time step.
[0115] A four-stage Video Swin Transformer network model was used to capture information at different semantic levels in video frames, obtaining spatiotemporal features at four different scales.
[0116] A specially designed multi-scale memory pool module is used to effectively mine salient cues from historical frames to enhance the semantic representation of salient objects in the current video frame. Specific steps include:
[0117] The design uses a memory pool of length 10 to store the spatiotemporal features of four scales extracted from several (10) historical frames; that is, each slot stores the features of only one video frame, rather than storing the features of all video frames in a time step.
[0118] After the video frames at each time step have their features extracted by the multi-scale encoder, they enter the memory pool. If the memory pool is full, the oldest historical frame is replaced to achieve dynamic storage.
[0119] A self-attention mechanism is applied to the features of 10 historical frames at each scale to obtain parallel multi-scale spatiotemporal context information.
[0120] By leveraging multi-scale spatiotemporal context information across four scales, and through parallel multi-scale cross-attention guidance and deep fusion with current frame features, the semantic representation of salient objects in the current video frame is enhanced.
[0121] By utilizing a specially designed multi-scale memory pool module, salient cues in historical frames are effectively mined to enhance the semantic representation of salient objects in the current video frame.
[0122] like Figure 2 As shown, the multi-scale memory pool module mainly consists of 10 temporal slots, each corresponding to the spatiotemporal features of a historical video frame, used to store its features at four scales (consistent with the four-stage feature scales output by Video SwinTransformer, such as...). Figure 2 In The subscripts represent scale numbers, and the superscripts represent video frame numbers at each time step, with resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the input, respectively. Each time, all video frames at the current time step are first updated into the memory pool (each slot corresponds to one video frame, and four video frames are updated per time step), and then subsequent calculations are performed. When a new video frame is input, the memory pool uses a first-in, first-out (FIFO) strategy; if the memory pool is full, the oldest feature is discarded and the newest feature is added. Next, parallel multi-scale self-attention is calculated. Within each scale, the features of the 10 most recent historical frames are concatenated, and the spatial format is adjusted to a sequence format suitable for attention calculation. Then, self-attention calculation is performed (the sequence-form features are transformed through three independent linear transformations to obtain three vectors: Query, Key, and Value, which are then calculated according to the formula). Finally, the spatial format is reconstructed as the output of the multi-scale memory pool module. Specific steps include:
[0123] S21 : like Figure 2 As shown, the multi-scale features extracted at this time step are dynamically represented in a highly generalized manner:
[0124] ;
[0125] in, This represents a four-stage Video Swin Transformer network model. This represents the four consecutive video frames input at this time step. This represents the spatiotemporal features obtained at four different scales; The subscript indicates the scale, and the superscript indicates the video frame number. X i Represents the spatiotemporal characteristics of a single historical video frame;
[0126] S22 : The spatiotemporal features extracted at this time step Updated into the memory pool; in this embodiment, the time step capacity of the memory pool is 10, and the feature of each historical video frame in the memory pool is a spatiotemporal feature, denoted as . If the memory pool is full, the oldest historical frame features are removed and new video frame features are added. Video frame features are spatiotemporal features, that is, spatiotemporal features extracted from video frames. Historical frame features are the same, that is, spatiotemporal features extracted from video frames at historical time steps.
[0127] S23 The historical frame features in the memory pool contain many effective salient cues that can be used for long-term memory modeling, compensate for the lack of spatiotemporal background in the current frame, and provide detailed edge supervision. Therefore, a parallel self-attention mechanism is used to mine salient cues from historical frame features at four scales, as shown below:
[0128] ;
[0129] ;
[0130] ;
[0131] in, For the spliced first Features at each scale The feature dimension of each video frame at this scale. For splicing operations, For the first in the memory pool The first historical video frame Features at each scale for Features after dimensional reshaping For the tensor dimension reshaping function, To integrate all historical information and emphasize the new features obtained from key spatiotemporal patterns, express function, These are three independent linear transformations;
[0132] S24 The spatiotemporal feature representation of the current time step is guided by salient cues from historical frames, as shown below: ;
[0133] ;
[0134] ;
[0135] in, This indicates that in the current time step, there are 4 consecutive video frames at the 1st... Feature representation at each scale express Features after dimensional reshaping The dimension is ,Right now ,express Features after dimensional reshaping For use The first one obtained after augmentation representation New features at each scale.
[0136] An edge-guided dual-branch decoder accurately captures the boundaries of salient objects and generates the final salient object detection map. The edge-guided dual-branch decoder includes a saliency prediction branch and an edge enhancement branch; it includes:
[0137] The output features, after being enhanced by the spatiotemporal context information of historical frames, are input into two branches of the decoder for decoding.
[0138] We utilize an edge enhancement branch to fuse three shallow visual features and align them with the true boundaries of prominent objects, while suppressing cluttered background edge noise.
[0139] The features are progressively upsampled and concatenated using a saliency prediction branch, and an initial saliency feature map is obtained through a residual channel attention mechanism.
[0140] The saliency features obtained from the saliency prediction branch and the edge enhancement branch are fused together, and the final saliency map is obtained through the residual channel attention mechanism.
[0141] The edge-guided dual-branch decoder accurately captures the boundaries of salient objects and generates the final salient object detection map.
[0142] The edge-guided dual-branch decoder includes a saliency prediction branch (SPB) and an edge enhancement branch (EEB). The saliency prediction branch is responsible for generating the initial saliency map, while the edge enhancement branch is responsible for learning the edge features of salient objects. Innovatively, a distance-based edgefitting loss (DEL) is introduced into both the saliency prediction branch and the edge enhancement branch to supervise the learning of the edge features of salient objects. At the same time, the saliency loss (SL) and the gated structure-aware loss (GSAL) are combined to jointly supervise the learning of the memory edge-guided network model.
[0143] The output features are enhanced by historical frame spatiotemporal context information. Decoding is performed using a dual-branch decoder; the specific implementation steps include:
[0144] S31 :like Figure 3 As shown, firstly, the edge enhancement branch applies to three shallow visual features. , , Feature transformation is performed using convolutional modules respectively. express The first part of the video frame features is represented by the subscript number indicating the scale and the superscript number indicating the video frame number. The convolution module is a common concept, representing the name of a computational process. Based on a designed convolution kernel (small matrix), it slides across the input feature matrix with a set stride. After each slide, it performs a dot product with the corresponding local region and sums the results, which is used as the element value at the corresponding position in the output feature map. Here, a 3x3 convolution kernel is used. The feature matrix obtained from the convolution calculation is then sequentially processed by Rectified Linear Unit (ReLU) and Batch Normalization (BN) to accelerate network training and improve stability.
[0145] Then, an upsampling transformation is performed to the same feature dimension. Next, the three transformed features are concatenated, and then channel enhancement is performed through the residual channel attention module to obtain the edge feature map of the salient object. The residual channel attention module, such as Figure 7 As shown, the calculation process of the features input to this module is as follows: convolution calculation, channel attention calculation, and residual connection (adding the values that have not undergone convolution and channel attention calculations to the values that have undergone both calculations), as shown below:
[0146] ;
[0147] ;
[0148] ;
[0149] in, This represents the output feature at the j-th scale after the enhancement expression is guided by the spatiotemporal context information of historical frames (the superscript CA is an abbreviation for Cross Attention). For convolutional modules, For upsampling operation, The features are those after convolution and upsampling. This indicates that features are concatenated along the channel dimension. This represents the feature obtained by concatenating features from three scales along the channel dimension (cat is short for concatenate). For residual channel attention module, This represents the generated edge feature map (EFM is an abbreviation for edge feature map).
[0150] S32 :like Figure 3 As shown, in the saliency prediction branch, from the deepest features... The process begins with sequential upsampling and concatenation, followed by enhancement of channel representations using a residual channel attention module to generate initial saliency feature maps. , This refers to the features at the fourth scale of each video frame; as shown below:
[0151] ;
[0152] ;
[0153] ;
[0154] ;
[0155] in, This indicates that features are concatenated along the channel dimension. This represents the intermediate features after the first convolution calculation and concatenation. This represents the intermediate features after the second convolution calculation and concatenation. This represents the intermediate features after the third convolution calculation and concatenation. For residual channel attention module, This represents the generated initial salient feature map (SFM is an abbreviation for saliency feature map).
[0156] S33 The feature maps generated by fusing the saliency prediction branch and the edge enhancement branch are used to accurately predict the final saliency map. :
[0157] .
[0158] Compared to fully supervised methods that provide accurate and reliable edge supervision for model training through dense pixel-level annotations, weakly supervised methods rely on sparse and incomplete labels. These labels can only indicate the rough location of salient objects and are difficult to effectively supervise the boundaries of salient objects. If the weakly supervised edge information of the entire image is directly used for supervised learning, the presence of background edge information may interfere with the model's learning of salient object boundaries, leading to suboptimal performance. Therefore, this invention innovatively proposes a distance-based edge fitting loss (DEL) method for supervising accurate boundary learning of salient objects; including:
[0159] First, two feature maps are generated through dual-branch decoding. and The Sobel operator is used to extract the predicted edge map, denoted as . and ;include:
[0160] The Sobel operator is a gradient-based edge detection method that uses a specially designed convolutional kernel to calculate the gradient intensity in the horizontal and vertical directions of an image to locate edges. Its core idea is that edge regions exhibit significant pixel intensity changes and thus larger gradient values. The algorithm process is as follows:
[0161] (1) Define the horizontal and vertical gradient kernels:
[0162] Horizontal gradient kernel: ;
[0163] Vertical gradient kernel: ;
[0164] (2) For the input feature map Convolution is performed using horizontal and vertical gradient kernels respectively to obtain the horizontal and vertical gradients. These two gradients are then combined to obtain the edge feature map. The calculation process is as follows:
[0165] ;
[0166] ;
[0167] ;
[0168] Next, the classic edge detection network RCF is used to generate an accurate edge map from the original image, denoted as . The accurate edge map includes accurate object boundaries while introducing background edge noise;
[0169] To suppress the interference of edge background noise on model learning, calculations were performed separately. and Each edge pixel in the middle arrive The middle corresponds to the nearest edge pixel The mean Euclidean distance;
[0170] Furthermore, to improve computational efficiency, the edge map is pre-processed. Calculate the distance map and Each pixel value represents the distance from that location to the nearest edge point; including:
[0171] The RCF edge map (a two-dimensional matrix where each element's value ranges from [0,1], representing the probability that the location is an edge) is binarized (edge pixels are converted to 255, and other pixels are converted to 0). Then, the RCF edge map is inverted (edge pixels are converted to 0, representing the background; other pixels are converted to 255, representing the foreground). Next, the `distanceTransform()` function from the OpenCV library in Python is called, which calculates the Euclidean distance from each pixel to the nearest background pixel. In this embodiment, using the inverted RCF edge map as input, the Euclidean distance from each pixel in the RCF edge map to the nearest edge (0-value pixels) can be calculated. The output two-dimensional matrix, where each element represents the Euclidean distance from that location to the nearest edge pixel, is the distance map. When calculating the mean Euclidean distance, the predicted edge map is binarized (edge pixels are converted to 1, and the rest are 0), and the edge fitting loss is obtained by averaging the dot product with the distance map (the RCF edge map and the predicted edge map have the same dimension, and the Euclidean distance from the edge pixel in the predicted edge map to the nearest edge pixel in the RCF edge map is equal to the Euclidean distance from the corresponding pixel in the RCF edge map to the nearest edge pixel, which is the element value at that position in the function output matrix).
[0172] Finally, through two predicted edge maps and Corresponding distance map and The dot product operation efficiently solves the loss; it significantly accelerates the training process while effectively reducing memory overhead.
[0173] The edge fitting loss of the saliency prediction branch and the edge enhancement branch in this process. and The definition is as follows:
[0174] ;
[0175] ;
[0176] in, This indicates the calculation of the Euclidean distance between two pixels. For edge map The number of pixels at the middle edge. For edge map The number of pixels at the middle edge. and To predict the edge graph of the first Each edge pixel to The distance to the nearest edge point; the subscript EEB indicates the edge enhancement branch, SPB indicates the significance prediction branch, and DEL indicates the edge fitting loss. This represents the edge fitting loss of the edge enhancement branch. The marginal fit loss represents the significance prediction branch. This represents the predicted edge map extracted from the feature map output from the EEB branch using the Sobel operator. This represents the predicted edge map extracted from the initial feature map output from the SPB branch using the Sobel operator. Represents the first edge in the predicted edge graph One edge pixel, Represents distance in the RCF edge graph The nearest edge pixel;
[0177] During the training of the memory edge-guided network model, in order to ensure the effective learning and optimization of the model, a multi-supervision mechanism was adopted, and detailed monitoring was carried out from multiple perspectives; edge loss was introduced. and The loss includes saliency loss (SL) and gated structure-aware loss (GSAL); among which, the saliency loss calculates the predicted saliency feature map. Weak labeling of truth values Pixel-level cross-entropy loss , The gated structure-aware loss supervises the output salient feature map to ensure structural consistency with the input image within the salient region. This joint supervision strategy enhances the model's ability to refine object edges while mitigating the challenges of weak supervision, thereby improving the performance of video salient object detection tasks. The loss consists of saliency loss, gated structure-aware loss, and the overall loss of the edge-memory guided network model. The design is as follows:
[0178] ;
[0179] ;
[0180] ;
[0181] in, Indicating the first true value weak annotation 1 pixel, The first saliency in the final salient feature map 1 pixel, Indicates the total number of pixels. Represents pixel coordinates. Represents the gradient direction. The coordinates in the final salient feature map are: pixels, The final salient feature map is in Gradient of direction, Represents image intensity. Represents the gradient of image intensity in the direction. Represents edge-aware weights. This represents a smoothing function.
[0182] Table 1 introduces the international standard dataset for evaluating the performance of this invention;
[0183]
[0184] Table 2 introduces the evaluation indicators for the performance of this invention;
[0185]
[0186] Table 3 shows a comparison of the detection accuracy of the model of the present invention;
[0187]
[0188] In this embodiment, quantitative performance is compared with two existing categories of video salient object detection models (fully supervised models and weakly supervised / unsupervised models) on three international standard datasets. For each row of data, "-" indicates no data available, "*" indicates training with enhanced pseudo-labels, and "‡" indicates training with only pure doodle weak labels. The following conclusions can be drawn from the experimental data analysis:
[0189] Fully supervised models (such as EGNet and MMN) on partial datasets (such as Visal and Davis) and Superior performance metrics (e.g., EGNet on ViSal) The indicator reached 0.946. The metric reached 0.941, but after training with augmented pseudo-labels, the weakly supervised model performed close to or even better than some fully supervised methods on datasets such as DAVSOD and Visal, such as MEGNet* (in this example) on the Visal dataset. The indicator reached 0.935. The index reached 0.923, which is better than the fully supervised RCRN model.
[0190] The model of this invention demonstrates outstanding performance across multiple metrics on various datasets. After training with enhanced pseudo-labels, the model achieved best performance in 8 out of 9 comparisons, and second-best performance in 1. Furthermore, it still achieved two top-three results when trained using only graffiti labels. This indicates that the method of this invention can effectively improve the accuracy and efficiency of salient object detection in weakly supervised scenarios, especially demonstrating excellent performance in structural similarity and error control.
[0191] Example 3
[0192] A memory-edge guided weakly supervised video salient object detection system includes:
[0193] The non-overlapping sliding window partitioning module is configured to partition a given long video sequence into a non-overlapping sliding window and extract several consecutive video frames.
[0194] The salient object detection module is configured to: input several video frames into a trained memory edge-guided network model to achieve weakly supervised video salient object detection; specifically including:
[0195] Feature extraction is performed on consecutive video frames at the current time step to obtain spatiotemporal features extracted at different scales;
[0196] The semantic representation of relevant objects in the current frame is enhanced by using salient cues mined from historical frames. The cues are then input into the decoder for decoding to obtain the final salient object detection map.
[0197] During the training of the memory edge-guided network model, the edge fitting loss is used to align with the true boundaries of salient objects while suppressing the interference of background edge noise.
Claims
1. A weakly supervised video salient object detection method based on memory-edge guidance, characterized in that, include: The given long video sequence is divided into non-overlapping sliding windows to extract several consecutive video frames. These video frames are then input into the trained memory edge-guided network model to achieve weakly supervised video salient object detection. Inputting the data into a trained memory edge-guided network model enables salient object detection in weakly supervised videos; specifically including: Feature extraction is performed on consecutive video frames at the current time step to obtain spatiotemporal features extracted at different scales; The semantic representation of relevant objects in the current frame is enhanced by using salient cues mined from historical frames. The cues are then input into the decoder for decoding to obtain the final salient object detection map. During the training of the memory edge-guided network model, the edge fitting loss is used to align with the true boundaries of salient objects while suppressing the interference of background edge noise. The memory edge-guided network model includes a multi-scale encoder, a multi-scale memory pool module, and an edge-guided dual-branch decoder; A multi-scale encoder is used to extract features from consecutive video frames at the current time step to obtain spatiotemporal features extracted at different scales; the specific steps include: The given long video sequence is divided into four consecutive non-overlapping sliding windows of length 4, and four consecutive video frames are extracted as the encoding input for each time step. A four-stage Video Swin Transformer network model was used to capture information at different semantic levels in video frames, obtaining spatiotemporal features at four different scales. A multi-scale memory pool module is used to dynamically store the spatiotemporal features of historical frames, and salient cues mined from historical frames are used to enhance the semantic representation of relevant objects in the current frame; the specific steps include: The memory pool is designed to store the spatiotemporal features of four scales extracted from several historical frames in the past. After the video frames at each time step have their features extracted by the multi-scale encoder, they enter the memory pool. If the memory pool is full, the oldest historical frame is replaced to achieve dynamic storage. A self-attention mechanism is applied to the features of historical frames at each scale to obtain parallel multi-scale spatiotemporal context information. By leveraging multi-scale spatiotemporal context information across four scales, and through parallel multi-scale cross-attention guidance and deep fusion with current frame features, the semantic representation of salient objects in the current video frame is enhanced. The final salient object detection map is obtained through an edge-guided dual-branch decoder; the edge-guided dual-branch decoder includes a salient prediction branch and an edge enhancement branch; including: The output features, after being enhanced by the spatiotemporal context information of historical frames, are input into two branches of the decoder for decoding. We utilize an edge enhancement branch to fuse three shallow visual features and align them with the true boundaries of prominent objects, while suppressing cluttered background edge noise. The features are progressively upsampled and concatenated using a saliency prediction branch, and an initial saliency feature map is obtained through a residual channel attention mechanism. The saliency features obtained from the saliency prediction branch and the edge enhancement branch are fused together, and the final saliency map is obtained through the residual channel attention mechanism.
2. The weakly supervised video salient object detection method based on memory-edge guidance according to claim 1, characterized in that, A multi-scale memory pool module is used to effectively mine salient cues from historical frames to enhance the semantic representation of salient objects in the current video frame. Specific steps include: S21: The multi-scale features extracted at this time step are dynamically represented in a highly generalized manner: X=Encoder vst (f); Among them, Encoder vst (·) represents a four-stage Video Swin Transformer network model, f = {f 1 ,f 2 ,f 3 ,f 4 } represents the four consecutive video frames input at this time step, X = {X 1 ,X 2 ,X 3 ,X 4 } represents the four spatiotemporal features obtained at different scales; The subscript indicates the scale, and the superscript indicates the video frame number. X i Represents the spatiotemporal characteristics of a single historical video frame; S22: Update the spatiotemporal feature X extracted in this time step into the memory pool; the feature of each historical video frame in the memory pool is the spatiotemporal feature, denoted as M. i If the memory pool is full, the oldest historical frame features are removed and the new video frame features are updated. S23: The salient cues of historical frame features are mined using a parallel self-attention mechanism at four scales, as shown below: MP j '=reshape(MP j ); in, For the features of the j-th scale after splicing, H j W j C j For each video frame, Cat(·) represents the feature dimension at this scale, and Cat(·) is the stitching operation. For the i-th historical video frame in the memory pool, the feature at the j-th scale is... For MP j Features after dimensional reshaping, where reshape(·) is the tensor dimension reshaping function. To integrate all historical information and highlight the new features obtained from key spatiotemporal patterns, softmax(·) represents the softmax function, and Q(·), K(·), and V(·) are three independent linear transformations; S24: Use salient cues from historical frames to guide the spatiotemporal feature representation of the current time step, as shown below: in, This represents the feature representation of four consecutive video frames at the j-th scale in the current time step. X represents j Features after dimensional reshaping The dimension is 10H j W j ×C j ,Right now express Features after dimensional reshaping For use The new feature at the j-th scale obtained after augmentation representation.
3. The weakly supervised video salient object detection method based on memory-edge guidance according to claim 1, characterized in that, The edge-guided dual-branch decoder accurately captures the boundaries of salient objects and generates the final salient object detection map. The edge-guided dual-branch decoder includes a saliency prediction branch and an edge enhancement branch. The saliency prediction branch is responsible for generating the initial saliency map, while the edge enhancement branch is responsible for learning the edge features of salient objects. The edge fitting loss based on distance minimization is introduced into both the saliency prediction branch and the edge enhancement branch to supervise the learning of the edge features of salient objects. At the same time, the saliency loss and the gated structure perception loss are combined to jointly supervise the learning of the memory edge-guided network model. The output features are enhanced by historical frame spatiotemporal context information. Decoding is performed using a dual-branch decoder; the specific implementation steps include: S31: First, the edge enhancement branch applies to three shallow visual features. Feature transformation is performed using convolutional modules respectively. express The part in the middle represents the feature of the first video frame, the subscript number represents the scale, and the superscript number represents the video frame number; Then, an upsampling transformation is performed to the same feature dimension. Next, the three transformed features are concatenated, and then channel enhancement is performed through the residual channel attention module to obtain the edge feature map X of the salient object. efm As shown below: X efm =RCAB(X cat ); in, This represents the output feature at the j-th scale after enhancement based on the spatiotemporal context information of historical frames. Conv(·) is the convolutional module, and Upsampling(·) is the upsampling operation. The features are the result of convolution and upsampling operations. `cat(·)` indicates that the features are concatenated along the channel dimension. X cat This represents the feature obtained by concatenating features from three scales along the channel dimension. RCAB(·) is the residual channel attention module. efm This represents the generated edge feature map; S32: In the saliency prediction branch, from the deepest features... The process begins with sequential upsampling and concatenation, followed by enhancement of channel representations using a residual channel attention module to generate an initial saliency feature map X. sfm , This refers to the features at the fourth scale of each video frame; as shown below: X sfm =RCAB(F3); Where Cat(·) represents the concatenation of features along the channel dimension, F1 represents the intermediate features after the first convolution calculation and concatenation, F2 represents the intermediate features after the second convolution calculation and concatenation, F3 represents the intermediate features after the third convolution calculation and concatenation, RCAB(·) is the residual channel attention module, and X sfm This represents the generated initial salient feature map; S33: The feature map generated by fusing the saliency prediction branch and the edge enhancement branch accurately predicts the final saliency map X. sm : X sm =RCAB(Cat(X efm ,X sfm ))。 4. A weakly supervised video saliency target detection method based on memory-edge guidance according to any one of claims 1-3, characterized in that, Edge fitting methods based on distance minimization are used for supervised and accurate boundary learning of salient objects; including: First, two feature maps X are generated through dual-branch decoding. efm and X sfm The predicted edge map is extracted using the Sobel operator, denoted as edge. EEB and edge SPB ; Next, the RCF edge detection network is used to generate an accurate edge map from the original image, denoted as edge. RCF The accurate edge map includes accurate object boundaries while introducing background edge noise; Calculate edge separately EEB and edge SPB Each edge pixel p i to the edge RCF The middle corresponds to the nearest edge pixel The mean Euclidean distance; Pre-processing edge graph RCF Calculate the distance map D RCF_EEB and D RCF_SPB Each pixel value represents the distance from that location to the nearest edge point; including: Using two predicted edge maps EEB and edge SPB Corresponding distance map D RCF_EEB and D RCF_SPB The dot product operation is used to efficiently solve for the loss; The edge fitting loss L of the saliency prediction branch and the edge enhancement branch DEL_EEB and L DEL_SPB The definition is as follows: Where dist(·) represents calculating the Euclidean distance between two pixels, and M is the edge map. EEB The number of pixels in the edge map, N is the edge number. SPB The number of pixels at the middle edge, D RCF_EEB (i) and D RCF_SPB (i) is the prediction of the i-th edge pixel in the edge map to the edge. RCF The distance to the nearest edge point; the subscript EEB indicates the edge enhancement branch, SPB indicates the saliency prediction branch, DEL indicates the edge fitting loss, and L... DEL_EEB L represents the edge fitting loss of the edge enhancement branch. DEL_SPB The edge fitting loss represents the significant prediction branch. EEB This represents the predicted edge map extracted from the feature map output from the EEB branch using the Sobel operator. SPB p represents the predicted edge map extracted from the initial feature map output from the SPB branch using the Sobel operator. i This represents the i-th edge pixel in the predicted edge map. Indicates the distance p in the RCF edge graph i The nearest edge pixel; During the training of the memory edge-guided network model, a multi-supervision mechanism is adopted, and an edge loss L is introduced. DEL_EEB and L DEL_SPB The loss includes saliency loss and gated structure perception loss; among which, the saliency loss is used to calculate the predicted saliency feature map X. sfm The pixel-level cross-entropy loss between the ground truth weak label G and the gated structure-aware loss supervise the structural consistency of the final salient feature map output with the input image within the salient region; the saliency loss, the gated structure-aware loss, and the overall loss L of the memory edge-guided network model are also considered. ALL The design is as follows: L ALL =L DEL_EEB +L DEL_SPB +L SL +0.3·L GSAL ; Among them, G i S represents the i-th pixel in the truth weak annotation. i Let (u,v) represent the i-th pixel in the final salient feature map, M represent the total number of pixels, (u,v) represent the pixel coordinates, d represent the gradient direction, and S represent the gradient direction. u,v This represents the pixel with coordinates (u,v) in the final salient feature map. I represents the gradient of the final salient feature map in the d direction. u,v Represents image intensity. Represents the gradient of image intensity in the direction. Represents edge-aware weights. This represents a smoothing function.
5. A memory-edge guided weakly supervised video salient object detection system, used to execute the memory-edge guided weakly supervised video salient object detection method according to any one of claims 1-4, characterized in that, include: The non-overlapping sliding window partitioning module is configured to partition a given long video sequence into a non-overlapping sliding window and extract several consecutive video frames. The salient object detection module is configured to: input several video frames into a trained memory edge-guided network model to achieve weakly supervised video salient object detection; specifically including: Feature extraction is performed on consecutive video frames at the current time step to obtain spatiotemporal features extracted at different scales; The semantic representation of relevant objects in the current frame is enhanced by using salient cues mined from historical frames. The cues are then input into the decoder for decoding to obtain the final salient object detection map. During the training of the memory edge-guided network model, the edge fitting loss is used to align with the true boundaries of salient objects while suppressing the interference of background edge noise.
Citation Information
Patent Citations
Weak supervision video target segmentation method based on multi-source saliency and time-space model adaptation
CN113283438A
An edge-guided RGBD underwater salient object detection method with multi-attention
JP7605548B1