Video retargeting method and system based on video saliency ranking
Patent Information
- Application Number
- CN202410935102.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-12
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-07-12
AI Technical Summary
[0005]本发明正是针对现有重定向过程不能保证视显著性区域的质量问题,出现了扭曲、伪影等时序不一致,显著性区域提取错误等问题,提供一种基于视频显著性排序的视频重定向方法及系统,首先获取至少包含每个实例的边界框,掩码结果和显著性排名得视频显著物体排序数据集;构建基于Mask R-CNN和图神经网络的显著物体排序模型,通过MaskR-CNN区分图像中的显著物体和非显著物体,通过注意力机制和位置保护注意力(PPA)提取实例的特征后,再将特征输入到图神经网络中,比较每个实例的显著性值并进行排序,得到每个实例的显著性程度;根据每个实例的显著性程度为其分配权重,根据权重以及实例位置加权得到裁剪中心;根据目标长宽比得到裁剪框的尺寸,最后利用LOESS方法平滑裁剪中心的时间序列,根据平滑后的裁剪中心和裁剪框的尺寸对视频帧进行裁剪得到重定向后的视频
[0027](1)本发明基于显著物体排序实现了视频重定向,相比于之前基于接缝雕刻,扭曲的方法,本发明方法利用裁剪的方式杜绝了伪影的发生;和其他基于裁剪的方式相比,本发明方法利用显著物体排序而不是显著性检测,防止了显著性较弱的前景和背景对裁剪区域的影响,能够最大限度的保留视觉显著性区域。
Smart Images

Figure CN119048948B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision technology, and particularly relates to video salient object ranking tasks and video redirection technology. It mainly relates to a video redirection method and system based on video salientity ranking. Background Technology
[0002] With the development of multimedia technology, people can watch videos on various devices such as smartphones and tablets. Video retargeting adjusts videos to different aspect ratios to suit different devices. The purpose of image retargeting is to preserve visually salient areas for a better visual experience. Based on this, image retargeting has achieved great success in the past through methods such as seam carving, cropping, and distortion. However, simply applying these methods to video retargeting will lead to serious temporal inconsistencies, greatly affecting the human visual experience.
[0003] The smartVidCrop method uses saliency detection for cropping to achieve video retargeting. However, this method only distinguishes between foreground and background, ignoring the saliency differences between different instances in the foreground. Therefore, this method often leads to bias towards background or less saliency objects during retargeting. Specifically, in a foreground scene, different objects may have different importance and attention levels, but smartVidCrop fails to accurately capture these differences. This makes it difficult to accurately retain truly important saliency objects when handling complex scenes. As a result, some less important background or less saliency foreground objects may be incorrectly identified as the focus of the retargeting, thus affecting the overall quality of the video content and the user's viewing experience.
[0004] The recent surge in salient object ranking (SOR) tasks has provided an important method for addressing saliency differences. SOR aims to rank objects in an image based on their saliency. Salientity typically refers to the degree to which an object attracts attention in a visual scene. The main objective is to identify and prioritize visually most prominent objects to better understand the scene's structure and focus. In the image domain, relative saliency is predicted by capturing multi-scale spatial contrast, including inter-instance contrast, local contrast between instances, and global contrast. In the video domain, temporal features of instances, including motion, should also be considered. This presents new challenges for saliency ranking. However, current video salient object ranking methods still cannot accurately capture motion information of instances, thus failing to achieve accurate saliency ranking. Therefore, how to better utilize the results of saliency ranking for video redirection has become a focus of research for those skilled in the art. Summary of the Invention
[0005] This invention addresses the problem that existing retargeting processes cannot guarantee the quality of visually salient regions, resulting in issues such as distortion, artifacts, temporal inconsistencies, and errors in salient region extraction. It provides a video retargeting method and system based on video saliency ranking. First, a salient object ranking dataset is obtained, containing bounding boxes, masking results, and saliency rankings for at least each instance. A salient object ranking model based on Mask R-CNN and a graph neural network is constructed. Mask R-CNN distinguishes between salient and non-salient objects in the image. Features of instances are extracted using attention mechanisms and position-preserving attention (PPA), and then input into the graph neural network. The saliency values of each instance are compared and ranked to obtain the saliency level of each instance. Weights are assigned to each instance based on its saliency level, and a cropping center is obtained by weighting the weights and instance positions. The size of the cropping box is obtained based on the target aspect ratio. Finally, the LOESS method is used to smooth the temporal series of the cropping center. The video frames are cropped based on the smoothed cropping center and the size of the cropping box to obtain the retargeted video. This invention's method can preserve salient regions to the greatest extent and ensure scene continuity.
[0006] To achieve the above objectives, the technical solution adopted by this invention is: a video redirection method based on video saliency ranking, specifically including the following steps:
[0007] S1, Obtain the dataset: Obtain the video salient object ranking dataset, which contains at least the bounding box, masking results and salientity ranking for each instance;
[0008] S2, Construct a salient object ranking model: The model is built based on Mask R-CNN and graph neural network. Mask R-CNN distinguishes salient and non-salient objects in the image. After extracting the features of the instances through the attention mechanism and position-preserving attention (PPA), the features are input into the graph neural network to compare the salience values of each instance and rank them to obtain the salience degree of each instance.
[0009] S3, Calculate the saliency center: Assign weights to each instance based on its saliency, and obtain the clipping center by weighting the weights and the instance positions;
[0010] S4, Determine the cropping frame: Obtain the size of the cropping frame based on the target aspect ratio, specifically: make the long side fill the original image and scale the short side to the target ratio;
[0011] S5, Smoothing: The LOESS method is used to smooth the time series of the crop center. The video frames are then cropped according to the smoothed crop center and the size of the crop frame to obtain the redirected video.
[0012] As an improvement of the present invention, the graph neural network in step S2 is used to mine the spatiotemporal relationships of instances;
[0013] At the spatial level, for local feature extraction, the bounding box area of each instance is proportionally doubled to obtain the bounding box of the local region of the instance, and the features within the bounding box are extracted using ROI Align.
[0014] For global feature extraction, the entire video frame is divided into 3×3 parts, and features of each part are extracted through an average pooling layer.
[0015] At the temporal level, for trajectory feature extraction, the bounding box area of the current frame instance is proportionally doubled, and the features within the bounding box at the corresponding position in the adjacent frame are extracted as the trajectory features of the instance using ROI Align.
[0016] As another improvement of the present invention, the loss function of the salient object ranking model in step S2 includes object detection loss L1 and salient ranking loss L2, wherein the object detection loss L1 is the classification loss L... cls , regression box loss L box and mask loss L mask The sum; the significance ranking loss L2 is specifically:
[0017]
[0018] Where {q1,q2} is any pair of instances; The significance score; β q These are the weighting coefficients; The formula for the number of combinations is given by , where N is the number of instances.
[0019] As another improvement of the present invention, the weighting coefficient β q Specifically:
[0020]
[0021] in, This represents the rank prediction for instances {q1, q2}. This represents the actual ranking result of instance {o1,o2}. The formula for the number of combinations is given by , where N is the number of instances.
[0022] As another improvement of the present invention, the calculation method of the cutting center in step S3 is as follows:
[0023]
[0024] Among them, (x i,y i The center coordinates of each instance, r i For significance ranking, r i =1,2,…,r i The larger the value, the stronger the significance.
[0025] To achieve the above objectives, the present invention also adopts the following technical solution: a video redirection system based on video saliency ranking, comprising a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the methods described above.
[0026] Compared with the prior art, the present invention has the following beneficial effects:
[0027] (1) The present invention achieves video redirection based on salient object sorting. Compared with the previous methods based on seam carving and distortion, the present invention eliminates artifacts by cropping. Compared with other cropping-based methods, the present invention uses salient object sorting instead of salientity detection to prevent the influence of weak foreground and background on the cropping area, and can preserve the visually salient area to the maximum extent.
[0028] (2) The video saliency ranking model framework proposed in this invention first segments instances based on Mask R-CNN, then enhances instance features through an attention mechanism, and introduces position and scale as saliency priors using a position-preserving attention module; subsequently, a spatiotemporal graph neural network is constructed to jointly optimize multi-scale spatial and temporal contrast cues to generate instance-level saliency features; finally, a fully connected neural network is used to infer saliency scores, and combined with the results of instance segmentation, the final saliency ranking result is obtained, which is more accurate.
[0029] (3) The graph neural network proposed in this paper is used to mine the spatiotemporal relationships of instances. The salience of instances in spatiotemporal relationships is considered at both the spatial and temporal levels. At the spatial level, the interaction relationships between instances, the contrast relationships between instances and their local parts, and the contrast relationships between instances and the global context are considered. At the temporal level, the interrelationships between instances at different times and the motion features of each instance are considered. The spatiotemporal relationships are jointly optimized to capture spatiotemporal salience. Attached Figure Description
[0030] Figure 1 This is a flowchart of the video redirection method based on video saliency ranking according to the present invention;
[0031] Figure 2 This is a schematic diagram of the video redirection framework of the present invention;
[0032] Figure 3 This is a schematic diagram of the structure of the video saliency ranking model of the present invention;
[0033] Figure 4 This is a schematic diagram illustrating the effect of video saliency ranking in Embodiment 1 of the present invention;
[0034] Figure 5 This is a comparison chart of the effects of different redirection methods in the test examples of this invention. Detailed Implementation
[0035] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.
[0036] Example 1
[0037] Video retargeting methods based on ranking salient objects in a video, such as Figure 1 and Figure 2 As shown, a video salient object ranking model is trained using a video salient object ranking dataset and a video salient detection dataset. The trained model obtains salient instance information for each frame, including the bounding box, mask, and salientity ranking of each instance. Then, a weight is assigned to each instance based on its salientity ranking, with higher weights for more salient instances. The global salient center is then used as the cropping center, and cropping boxes are generated based on the target aspect ratio. Finally, the positions of the obtained cropping centers are smoothed over time to ensure scene coherence. The specific steps include:
[0038] Step S1: Obtain the video salient object ranking dataset for training the video saliency model. To meet the requirements of the video salient object ranking task, for each frame in the video, the information that needs to be annotated includes: the bounding boxes, masks, and saliency levels of all salient instances.
[0039] The video salient object ranking dataset used was extracted from the publicly available datasets RVSOD and DAVSOD. In this embodiment, the data processing in RVSOD and DAVSOD is as follows:
[0040] 1) For RVSOD, the bounding box and mask information for each instance is missing. The instance-level mask is extracted from the existing image-level mask and the bounding box is generated. For the saliency level of the instance, the saliency level is assigned to each instance based on the density of the gaze points provided by RVSOD. The division of the training set and the test set is the same as that of RVSOD.
[0041] 2) For DAVSOD, instance-level annotations are obtained in the same manner as for RVSOD. However, some videos in the DAVSOD dataset contain only a single salient instance, making them unsuitable for saliency ranking. Therefore, this portion of the data is discarded, and the remaining videos are divided into two categories: those with some frames containing multiple salient instances and those with all frames containing multiple salient instances. Finally, the training and test sets for each category are split in a 4:1 ratio.
[0042] Step S2: Construct a salient object ranking model using Mask R-CNN and graph neural networks.
[0043] This invention employs Mask R-CNN to segment instances and mines salient features of instances using graph neural networks. The video salient object ranking model comprises an object detection module (Mask R-CNN) and a salient ranking module based on graph convolutional neural networks. These two modules are trained sequentially to ensure end-to-end training and prediction capabilities for the entire model. In the detection module, the model only distinguishes between salient and non-salient objects in the image without performing identification. In the ranking module, salient values are obtained based on the features of the detected instances combined with spatiotemporal relationships to rank different instances. The entire salient ranking model is as follows: Figure 3 As shown: After extracting instance features using Mask R-CNN, the attention mechanism and Position Preserving Attention (PPA) are used to fully extract the instance features. These features are then input into a graph neural network, and finally, the saliency values of each instance are compared and ranked. Detailed steps include:
[0044] S21: Use MaskR-CNN for instance segmentation. In instance segmentation, only salient objects and non-salient objects are distinguished.
[0045] S22: In instance segmentation, MaskR-CNN uses ROI Align to extract the features of the Region Proposal. Considering the influence of shallow features such as location information and size on saliency, the extracted instance features are encoded with location and attention mechanism to extract the saliency information of the instance.
[0046] S23: To fully extract the saliency information of instances, a graph convolutional neural network is constructed to mine the spatiotemporal relationships of instances. At the spatial level, the interactions between instances, between instances and their local features, and between instances and the global scene are mainly considered. For local feature extraction, the bounding box area of each instance is proportionally doubled to obtain the bounding box of the instance's local region, and then ROI Align is used to extract features within the bounding box. For feature extraction, the entire frame is divided into 3×3 parts, and features are extracted from each part using an average pooling layer. At the temporal level, the interactions between different instances and the influence of instance trajectory information on instances are mainly considered. For trajectory feature extraction, the bounding box area of the instance in the current frame is proportionally doubled, and ROI Align is used to extract features within the bounding box at the corresponding positions in adjacent frames as the instance's trajectory features.
[0047] S24: The loss function for constructing a video salient object ranking model based on graph convolutional neural networks mainly consists of two parts: first, the loss for object detection, including classification loss, bounding box loss, and mask loss, with the specific calculation formula as follows:
[0048] L1 = L cls +L box +L mask
[0049] For significance ranking, considering all instance pairs, for any pair of instances {q1, q2}, where the significance of q1 is less than that of q2, the predicted significance scores are respectively The goal is to reduce The value and increase The value of . Therefore, the specific calculation formula is as follows:
[0050]
[0051] Where β q It is a weighting coefficient whose purpose is to assign greater weight to pairs with large ranking differences and less weight to pairs with similar rankings, thereby clearly optimizing instances with extreme levels. The specific calculation formula is as follows:
[0052]
[0053] In this embodiment, the parameters of the video saliency ranking model based on graph neural networks are set as follows: learning rate is 5e-6, Adam is used as the optimizer, batch size is set to 1, maximum number of iterations is set to 200,000, and the learning rate is decayed to 1 / 10 of the current learning rate at the 80,000th and 150,000th iterations.
[0054] Figure 4This demonstrates the effectiveness of the present invention on the task of ranking salient objects in videos. The first two lines are the input video frames and salient instances. The third, fourth, fifth, and sixth lines show the results of methods with only spatial salient signals, methods with only global temporal and spatial signals, methods with only instance-level temporal and spatial signals, and the method of the present invention, respectively. Figure 4 As can be seen above, before the introduction of the temporal feature extraction method in this invention, the model could not accurately mine the saliency of instances in time. This invention introduces the trajectory features of instances, enabling the model to accurately extract the saliency features of instances, achieving the best results to date.
[0055] Step S3: Obtain the clipping center based on the predicted saliency results: The predicted saliency results include the location and saliency level of each salient instance. A weight is assigned to each instance based on its saliency level, and the clipping center is obtained by weighting the instances based on their weights and locations. The center coordinates of each instance are (x...). i ,y i The significance ranking is r. i (where r) i =1,2…,and r The larger the value of i, the stronger the significance. Therefore, the clipping center is:
[0056]
[0057] Step S4: Obtain the dimensions of the cropping frame based on the target aspect ratio. For the target length and width f... w / f h We make the longer side fill the original image and scale the shorter side to the target proportions. For example, if the target dimensions are f... w / f h =4 / 5, we keep the image height unchanged and only crop it on the x-axis.
[0058] Step S5: Use the LOESS method to smooth the time series of the crop center, and crop the video frames according to the size of the smoothed crop center and the crop frame to obtain the redirected video.
[0059] In summary, the method of this invention utilizes the publicly available datasets RVSOD and DAVSOD to train a graph neural network-based model for ranking salient objects in videos. Based on the extracted salient instances and their salience strength, an appropriate cropping center is selected, combined with the required aspect ratio, to choose a suitable cropping region. Finally, a smoothing method is used to ensure the consistency of the final result. This method overcomes the problem of poor visual experience within visually salient regions during retargeting. Through the salient object ranking module, all salient instances can be obtained while isolating non-salient foreground or background. Furthermore, by ranking, a weight is assigned to salient instances, so that even under extreme retargeting aspect ratios, this invention can retain the most salient visual regions. Moreover, through cropping, artifacts, distortions, and other phenomena that affect human visual experience are eliminated.
[0060] Test case
[0061] Performance testing of video salient object ranking based on graph neural network: To verify the effectiveness of the salient ranking module in this invention, this test case performs performance testing on the video salient object ranking task and compares it with several state-of-the-art models.
[0062] The evaluation metrics for the video salient object ranking task are as follows:
[0063]
[0064] Where S(x,y) represents the predicted pixel grayscale value, and GT(x,y) represents the actual pixel grayscale value. The MAE metric is used to measure the absolute error between the predicted image and the label image; the lower the value, the better the model.
[0065] SSA-SOR=Pearson-Correlation(Ranks pre ,Ranks gt )
[0066] Where Pearson-Correlation represents the calculation of the Pearson correlation coefficient, and Ranks pre ,Ranks gt These represent the predicted and actual saliency ranking sequences, respectively. SA-SOR is currently the most advanced metric for ranking salient objects. Compared to previous metrics like SOR and SSOR, SA-SOR can simultaneously reflect both segmentation and ranking performance.
[0067] The comparison results of the video salient object ranking method based on graph neural networks in this invention with two other image salient object ranking methods and another advanced video salient object ranking model are shown in the table below:
[0068]
[0069] As can be seen from the table above, the video salient object ranking method based on graph neural networks in this invention outperforms the state-of-the-art methods on both datasets and two metrics, demonstrating strong competitiveness.
[0070] Performance testing of the video redirection method based on video saliency ranking: To verify the effectiveness of the video redirection method of the present invention, this test case performs performance testing on a video redirection task and compares it with several advanced models.
[0071] Since there is no universal theorem evaluation metric for video redirection, we chose to use visual results for qualitative comparison.
[0072] The video redirection method based on salient object ranking in this invention is compared with two other video redirection methods as follows: Figure 5 As shown: the first column is the input video frame, and the next six columns are divided into two parts: results with aspect ratios of 2:3 and 4:5, respectively. The first, second, and third columns represent the seam engraving method, the saliency-based cropping method, and the method of this invention, respectively. From Figure 5 As can be seen, the seam-based engraving method often causes instance artifacts, which affects the visual experience (yellow box in the figure); while the saliency-based cropping method is easily interfered with by weakly saliency objects and background information, resulting in the generation of incorrect cropping boxes. In contrast, the method of the present invention uses instance-level saliency for redirection and achieves the best visual effect.
[0073] First, this invention employs a video salient object ranking method to extract salient regions from video frames. Leveraging the advantages of graph neural networks in feature interaction modeling, it fully mines the spatiotemporal saliency cues of instances using spatial and temporal contrast maps, achieving the best video salient object ranking results to date. Second, this invention employs a video retargeting method based on video salient object ranking. This method uses cropping to eliminate artifacts and other results that affect human visual experience. Furthermore, compared to previous methods based on video salientity, this invention's method effectively isolates the background and foreground and eliminates the influence of weakly salient foregrounds, resulting in more accurate cropping. Extensive comparative experiments demonstrate and verify the rationality of the proposed video retargeting method based on video salientity ranking.
[0074] The video retargeting method based on salient object ranking provided by this invention can transform input videos into target aspect ratios while preserving salient regions to improve the visual experience. Specifically, in the industrial sector, with the development of multimedia technology and the widespread adoption of devices with different aspect ratios such as computers, mobile phones, watches, and tablets, the advantages of this retargeting method can be maximized, resulting in models that meet the needs of enterprise domains and are close to supervised pre-trained models.
[0075] It should be noted that the above content merely illustrates the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. For those skilled in the art, various improvements and modifications can be made without departing from the principle of the present invention, and all such improvements and modifications fall within the scope of protection of the claims of the present invention.
Claims
1. A video redirection method based on video saliency ranking, characterized in that, Specifically, the steps include the following: S1, Obtain the dataset: Obtain the video salient object ranking dataset, which contains at least the bounding box, masking results and salientity ranking for each instance; S2, Constructing a salient object ranking model: This model is built based on Mask R-CNN and a graph neural network. Mask R-CNN distinguishes salient and non-salient objects in an image. After extracting instance features through an attention mechanism and position-protected attention (PPA), the features are input into the graph neural network. The salience values of each instance are compared and ranked to obtain the salience level of each instance. The graph neural network is used to mine the spatiotemporal relationships of instances. At the spatial level, for local feature extraction, the bounding box area of each instance is proportionally doubled to obtain the bounding box of the local region of the instance, and the features within the bounding box are extracted using ROI Align. For global feature extraction, the entire video frame is divided into... Each part is divided into portions, and features are extracted from each portion using an average pooling layer; At the temporal level, for trajectory feature extraction, the bounding box area of the current frame instance is proportionally doubled, and the features within the bounding box are extracted at the corresponding positions in adjacent frames using ROI Align as the trajectory features of the instance. S3, Calculate the saliency center: Assign weights to each instance based on its saliency, and obtain the clipping center by weighting the weights and the instance positions; S4, Determine the cropping frame: Obtain the size of the cropping frame based on the target aspect ratio, specifically: make the long side fill the original image and scale the short side to the target ratio; S5, Smoothing: The LOESS method is used to smooth the time series of the crop center. The video frames are then cropped according to the smoothed crop center and the size of the crop frame to obtain the redirected video.
2. The video redirection method based on video saliency ranking as described in claim 1, characterized in that: The loss function of the salient object ranking model in step S2 includes object detection loss. and significance ranking loss Among them, object detection loss Classification loss Regression box loss and mask loss The significance ranking loss Specifically: ; in, For any pair of instances; These are the weighting coefficients; The formula for the number of combinations represents the number of all instance pairs. The number of instances.
3. The video redirection method based on video saliency ranking as described in claim 2, characterized in that: The Specifically: ; in, Indicates for instance Ranking predictions Representation of instances The actual ranking results The formula for the number of combinations represents the number of all instance pairs. The number of instances.
4. The video redirection method based on video saliency ranking as described in claim 1, characterized in that: In step S3, the calculation method for the cutting center is as follows: ; in, The center coordinates of each instance For significance ranking, , The larger the value, the stronger the significance.
5. A video redirection system based on video saliency ranking, comprising a computer program, characterized in that: When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-4 above.