Rapid panoramic video saliency prediction method
Through the lightweight panoramic video significance prediction architecture, the significant areas of panoramic video are quickly extracted, solving the problems of high computing complexity and insufficient real-time performance, and achieving efficient and accurate significant prediction of panoramic videos.
Patent Information
- Application Number
- CN202510708582.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-29
AI Technical Summary
The existing panoramic video significance prediction methods have high computational complexity and insufficient real-time performance, which is difficult to meet the needs of real-time applications. In addition, the traditional central prior method is difficult to adapt to the unique spatial distribution and human visual attention characteristics of panoramic video.
A lightweight panoramic video significance prediction architecture is designed, and the equatorial prior and spatial semantic integration prior are generated through the basic visual feature extraction module, the space-time decoupling module and the dual-branch attention prior. Combined with the lightweight timing relationship modeling module, the significant area is quickly extracted.
It greatly reduces computing complexity, improves processing speed and accuracy, adapts to the visual bias of panoramic videos and human attention, and generates more time-consistent significant graphs.
Smart Images

Figure CN120298955A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and video processing, and particularly to a fast panoramic video saliency prediction method. Background Art
[0002] With the rapid development of the Internet and digital technologies, panoramic video, as a new multimedia form, is gradually changing our visual experience. Panoramic video can provide a 360-degree omnidirectional view, and users can adjust the viewing angle to comprehensively view the scenes in the video.
[0003] However, while panoramic video brings a unique viewing experience, its huge data volume and high resolution also pose challenges to transmission efficiency. To balance transmission efficiency and user experience, researchers have begun to explore panoramic video saliency prediction technology. This technology aims to simulate the human attention mechanism through algorithms to quickly and accurately locate and predict the regions in the current video with high user attention. Through the precise identification of salient regions, video coding and transmission strategies can be optimized accordingly, giving priority to ensuring the high-quality transmission of key regions and achieving high-quality video compression.
[0004] Existing methods for panoramic video saliency prediction have some significant drawbacks, which limit their performance and efficiency in practical applications. On the one hand, the high resolution and large data volume of panoramic video pose extremely high requirements for computing resources. Some existing saliency prediction algorithms have too high computational complexity when processing panoramic video, resulting in slow processing speed and difficulty in meeting the needs of real-time applications. This not only increases system costs but also limits the popularization of panoramic video saliency prediction technology in various real-time application scenarios. On the other hand, considering the unique spatial distribution pattern of panoramic video, people tend to focus more on the regions near the equator and less on the polar regions when watching, and human attention often pays more attention to regions with semantic information. The center prior method adopted by traditional video saliency prediction is difficult to adapt to this unique attention characteristic. Therefore, developing an efficient, accurate, and more adaptable saliency prediction method for panoramic video has become a key technical requirement at present. Summary of the Invention
[0005] Object of the Invention: The technical problem to be solved by the present invention is to provide a fast panoramic video saliency prediction method in view of the deficiencies of the prior art. Aiming at the problems of high computational complexity and insufficient real-time performance of existing panoramic video saliency prediction methods, the present invention innovatively designs a lightweight panoramic video saliency prediction architecture. Considering the unique spatial distribution of panoramic video and the characteristics of human visual attention, an adaptive prior model is constructed in the architecture.
[0006] The method of the present invention includes the following steps: Step 1: Process the panoramic video into video frames; Step 2: Take consecutive video frames as input, and extract multi-scale semantic features of each video frame through the basic visual feature extraction module; Step 3: Construct a spatio-temporal decoupling module to extract temporal features and spatial features of consecutive frames respectively; Step 4: Construct a two-branch attention prior module to generate an equator prior and a spatial semantic integration prior to simulate the visual bias phenomenon in panoramic video viewing behavior; Step 5: Input the spatio-temporal features and the prior map into a lightweight temporal relationship modeling module to generate the final saliency map.
[0007] Step 1 includes: First, parse and process the panoramic video in the equirectangular projection (ERP) format, convert the panoramic video into a series of equirectangular projection ERP video frames. The equirectangular projection ERP format is a projection method that maps the spherical surface to a rectangular plane. It assumes that the spherical surface and the cylindrical surface are tangent to the equator, projects the longitude and latitude lines on the spherical surface onto the cylindrical surface, and then unfolds along a generatrix of the cylindrical surface into a plane to obtain an equidistant cylindrical projection map. Considering that the original panoramic video has a high resolution, if the high-resolution equirectangular projection ERP video frames are directly input into the model, it will increase the computational burden of the model. The bilinear interpolation method is used to downsample the equirectangular projection ERP video frames, and the resolution of the equirectangular projection ERP video frames is adjusted to W×H, where W and H represent the width and height of the projection respectively.
[0008] Step 2 includes: The basic visual feature extraction module includes a Residual Network 18 (ResNet18), a dilated convolutional layer, a feature concatenation operation, an upsampling layer, and a convolutional layer.
[0009] In Step 2, the basic visual feature extraction module performs the following operations: Take T consecutive video frames { , ,... } as the input of the Residual Network 18 (ResNet18), where represents the T-th video frame. First, obtain the outputs of the 3rd, 4th, and 5th convolutional blocks of the Residual Network 18, which are respectively denoted as and . Process the output of the 5th convolutional block of the Residual Network 18, where represents the combined batch and time dimensions, The number of channels of is 512, and respectively represent the width and height of which represents the real number space; using m = 4 dilated convolutional layers to process to obtain multi-scale semantic information, the sampling rates r corresponding to the m dilated convolutional layers are respectively set as r = { , , , }, where represents the sampling rate corresponding to the 4th dilated convolutional layer; subsequently, the multi-scale semantic information is concatenated in the channel dimension and processed using 1×1 convolution for channel dimension adjustment to obtain the processed feature , where , , respectively represent the number of channels, width, and height of the feature : , , where represents the output of the 2nd convolutional block of the residual network ResNet18, , and respectively represent the 3rd, 4th, and 5th convolutional blocks of the residual network ResNet18, represents the dilated convolutional kernel with a sampling rate of , represents the feature concatenation operation, dim represents the dimension of the operation, is the convolutional kernel (in this invention, different letters (such as , , and etc.) are used to separately identify the convolutional kernels to distinguish convolutional kernels with different settings, such as the input and output channel numbers of the convolutional kernel and the convolutional kernel size are set independently); Two 1×1 convolutions are respectively used to adjust the channel dimensions of the output of the 3rd convolutional block and the output of the 4th convolutional block of the residual network ResNet18, and then the bilinear fusion operation is used to upsample the processed and to the same dimension size, and finally , and are concatenated in the channel dimension and processed using 1×1 convolution to obtain the basic feature of the video frame: , Among them 、 and represent the convolutional kernel, represents the upsampling operation of the upsampling layer.
[0010] Step 3 includes: The spatio-temporal decoupling module includes a dimension separation operation, a spatial convolutional layer, a batch normalization layer, an activation layer, a temporal convolutional layer, and a residual connection; The spatio-temporal decoupling module performs the following operations: First, perform a dimension transformation on the basic features of the video frame and separate the batch dimension and the temporal dimension into two dimensions to obtain the transformed features , 、 、 respectively represent the number of channels, width, and height of the basic feature ; Then, use a structure with serial 2D convolution and 1D convolution to extract spatio-temporal features: First, use a 3D convolution of 1×3×3 to represent the 2D convolution operation to construct a spatial convolutional layer for extracting spatial dimension features; Subsequently, perform standardization processing on the distribution of the spatial dimension features through a batch normalization layer, and further enhance its non-linear expression ability with the help of the ReLU activation function; Use a 3D convolution of 3×1×1 to represent the 1D convolution operation to construct a temporal convolutional layer to efficiently aggregate the temporal features of consecutive frames along the temporal dimension; Finally, process through the batch normalization layer and the ReLU activation layer again, and use the residual connection to obtain the spatio-temporal features , and the specific formula is: , where BN represents the batch normalization operation, represents the 1×3×3 convolutional kernel, represents the 3×1×1 convolutional kernel.
[0011] Step 4 includes: The dual-branch attention prior module includes a panoramic equator prior module and a spatial semantic integration prior module. The panoramic equator prior module is used to generate an equator prior for panoramic videos or panoramic images. The generation method of the panoramic equator prior map is: , where y represents the coordinate position in the vertical axis direction, represents the prior value at y, and respectively represent the mean and variance along the vertical axis direction; exp represents the natural exponential function; The calculation formula of is: The calculation formula is: , where ; ; Use N different variances to generate corresponding panoramic equatorial prior maps, and then use two layers of convolution to adaptively combine the N prior maps to generate the final panoramic equatorial prior features ; The spatial semantic integration prior module is used to generate a spatial semantic integration prior map for the spatio-temporal features extracted by the spatio-temporal decoupling module , where , , respectively represent the number of channels, width, and height of the spatio-temporal feature ; The spatial semantic integration prior module is used to obtain spatial semantic integration prior knowledge from the continuous T frames of the entire batch, and integrate the information in the time dimension to obtain prior features. The time dimension addition is used for temporal information integration. For the integrated spatio-temporal feature , two convolutional layers are used to extract semantic information, and bilinear interpolation is used for upsampling to obtain the final spatial semantic integration prior map. The specific formula is: , where represents the upsampling operation, and represent the convolutional kernels, and sum represents the addition operation in the time dimension; Combine and through feature concatenation and convolution operations to generate the final prior map : , where represents the convolutional kernel.
[0012] Step 5 includes: Concatenate the prior map with the spatio-temporal feature in the channel dimension and obtain the spatio-temporal prior feature through a convolutional layer : , where represents the convolutional kernel.
[0013] Step 5 also includes: For the spatio-temporal prior features generated from T consecutive video frames { , ,... }, The final predicted saliency map is obtained through moving weighted average, and the formula is: , , where represents the sigmod function, represents the convolution operation, represents the convolution kernel, represents the fusion weight, represents the input at the t-th time step, corresponding to the t-th feature in the time dimension; represents the output at the t-th time step; the outputs of each time step are concatenated in the time dimension, and then through convolution and sigmod function processing, the final predicted saliency map S is obtained: , where represents the convolution kernel.
[0014] The present invention also provides an electronic device, including a processor and a memory, where the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the above method.
[0015] The present invention also provides a storage medium, storing a computer program or instruction, and when the computer program or instruction runs on a computer, it executes the steps of the above method.
[0016] This method innovatively constructs a cascaded architecture: First, the panoramic video is processed into video frames, and the consecutive video frames are used as inputs. Aiming at the problems of large field of view and complex content in panoramic videos, a two-stage cascaded method is designed. First, the basic visual feature extraction module performs multi-scale semantic feature extraction on the input frame sequence, quickly extracts useful information, and filters background information. Subsequently, the spatio-temporal decoupling module extracts the temporal features and spatial features of consecutive frames in a serial manner. Then, the dual-branch attention prior module generates an equator prior and a spatial semantic integration prior to simulate the visual bias phenomenon in the panoramic video viewing behavior. Finally, the spatio-temporal features and prior features are fused, and the temporal relationship modeling module processes the temporal relationship between consecutive video frames and generates the final saliency map.
[0017] Beneficial effects: By constructing a lightweight cascaded architecture, the present invention significantly reduces the computational complexity and the number of parameters compared with traditional methods, solves the problem of excessive computational resource consumption of existing methods, and meets the processing speed and real-time requirements of panoramic video saliency prediction; Aiming at the visual bias in the equatorial region and the characteristics of the human attention mechanism unique to panoramic videos, a dual-branch attention prior module is used to generate an equatorial prior and a spatial semantic integration prior, simulating the distribution of human visual attention, avoiding the deficiencies of traditional central priors, making the algorithm design more reasonable and improving the accuracy of salient region prediction; A more lightweight temporal relationship modeling module is used to model the temporal relationship of consecutive frames, effectively capturing the temporal dependencies and dynamic changes in the video sequence, solving the problem of temporal incoherence of salient regions caused by perspective switching or scene movement, and generating a more temporally consistent saliency map, thus ensuring that the prediction results are more reliable. Description of the Drawings
[0018] Figure 1 is the overall flowchart of the method of the present invention.
[0019] Figure 2 is the flowchart of the basic visual feature extraction module.
[0020] Figure 3 is the flowchart of the spatio-temporal decoupling module.
[0021] Figure 4 is the panoramic equatorial prior map.
[0022] Figure 5 is the flowchart of the temporal relationship modeling module.
[0023] Figure 6 is the visualized prediction result. Detailed Embodiments
[0024] The following further detailed description of the present invention is made in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.
[0025] As Figure 1 shown, the embodiment of the present invention provides a fast panoramic video saliency prediction method, including the following steps: Step 1, process the panoramic video into video frames.
[0026] First, parse and process the panoramic video in the equidistant cylindrical projection ERP format, and convert it into a series of equidistant cylindrical projection ERP video frames. The equidistant cylindrical projection ERP format is a projection method that maps the spherical surface to a rectangular plane. It assumes that the spherical surface and the cylindrical surface are tangent to the equator, projects the longitude and latitude lines on the spherical surface onto the cylindrical surface, and then unfolds along a generatrix of the cylindrical surface into a plane to obtain an equidistant cylindrical projection map. Considering that the original panoramic video has a high resolution, if the high-resolution equidistant cylindrical projection ERP video frames are directly input into the model, it will increase the computational burden of the model. Therefore, the bilinear interpolation method is used to downsample the equidistant cylindrical projection ERP video frames, and the resolution of the equidistant cylindrical projection ERP video frames is adjusted to W×H, where W and H represent the width and height of the projection respectively.
[0027] Step 2: Construct a basic visual feature extraction module. Based on the principle of first extracting basic visual features in the human attention mechanism, this module extracts the basic information related to saliency in each video frame.
[0028] As Figure 2 shown, the basic visual feature extraction module includes: a residual network, an atrous convolution layer, a feature concatenation operation, an upsampling layer, and a convolution layer. The specific operation steps are as follows: Use T (in this embodiment, T = 5 is used for processing) consecutive video frames { , ,... } as the input of the model, where represents the T-th video frame. First, obtain the outputs of the 3rd, 4th, and 5th convolutional blocks of the residual network ResNet18, which are respectively denoted as and . Process the output of the 5th convolutional block of the residual network ResNet18 model, where represents merging the batch and time dimensions, and the number of channels is 512, and respectively represent 's width and height, represents the real number space; specifically, use m = 4 atrous convolution layers to process it, and the corresponding sampling rates are set as r = { , , , } (in this embodiment, , = 6, = 12, Taking 18 as an example, the number of output channels of each dilated convolutional layer is 256. In this way, multi-scale semantic information is obtained. Subsequently, these multi-scale semantic information are feature concatenated in the channel dimension, and processed by {1024, 256}×1×1 convolution for adjusting the channel dimension to obtain the processed features : , , where represents the output of the second convolutional block of the Residual Network ResNet18, 、 and respectively represent the third, fourth, and fifth convolutional blocks of the Residual Network ResNet18; represents a dilated convolutional kernel with a sampling rate of r, represents the feature concatenation operation, and dim represents the dimension of the operation, is the convolutional kernel (the setting of the present invention is {number of input channels, number of output channels}×height of convolutional kernel×width of convolutional kernel, taking {1024, 256}×1×1 as an example).
[0029] Use {2048,64}×1×1 convolution and {2048,32}×1×1 convolution to process the output of the third convolutional block and the output of the fourth convolutional block respectively, and then use the bilinear fusion operation to upsample the processed and to the same dimension size, and finally 、 and are feature concatenated in the channel dimension and processed by {352,512}×1×1 convolution to obtain the basic features of the video frame : , where 、 and represent convolutional kernels (the settings of the present invention 、 and are taken as {2048,64}×1×1, {2048,32}×1×1, and {352,512}×1×1 respectively), represents the upsampling operation.
[0030] Step 3: Construct a spatio-temporal decoupling module, which improves the C3D model based on the cascaded architecture of 2D convolution and 1D convolution for fast extraction of spatio-temporal features.
[0031] As Figure 3 shown, the spatio-temporal decoupling module includes a dimension separation operation, a spatial convolution layer, a batch normalization layer, an activation layer, a temporal convolution layer, and a residual connection. First, perform a dimension transformation on the basic features of the video frames and separate the batch dimension B and the temporal dimension T into two dimensions to obtain the transformed features , where , and respectively represent the number of channels, width, and height, and then use the structure of serial 2D convolution and 1D convolution to extract spatio-temporal features. This structure first uses a 3D convolution of 1×3×3 to represent the 2D convolution operation to construct a spatial convolution layer for extracting spatial dimension features. Subsequently, perform standardization processing on the obtained feature distribution through the batch normalization layer, and further enhance its non-linear expression ability with the help of the linear rectifier unit ReLU activation function. Use a 3D convolution of 3×1×1 to represent the 1D convolution operation to construct a temporal convolution layer to efficiently aggregate the temporal features of consecutive frames along the temporal dimension. Finally, process it again through the batch normalization layer and the linear rectifier unit ReLU activation layer, and use the residual connection to obtain the spatio-temporal features , and the specific formula is: , where BN represents the batch normalization operation, ReLU represents the linear rectifier unit ReLU activation function, represents the convolution kernel (the setting of the present invention takes {256,256}×1×3×3 as an example), represents the convolution kernel (the setting of the present invention takes {256,256}×3×1×1 as an example).
[0032] Step 4: Construct a dual-branch attention prior module to generate an equator prior and a spatial semantic integration prior to simulate the visual bias phenomenon in the panoramic video viewing behavior.
[0033] The dual-branch attention prior module consists of two sub-modules for constructing priors. Among them, the main function of the panoramic equator prior module is to generate an equator prior specifically for panoramic videos or panoramic images. In panoramic videos, based on their unique distribution characteristics, significant information is usually mainly concentrated in the regions near the equator, while the attention received by the polar regions is relatively less.
[0034] Given this characteristic, in order to obtain a central prior that matches the features of panoramic videos, this module adopts a strategy of using a Gaussian distribution on the y-axis and full padding on the x-axis. Specifically, the panoramic equatorial prior map is generated as follows: , where and represent the mean and variance along the y-direction respectively. N = 8 different variances are used to generate the corresponding panoramic equatorial prior maps. Statistically, each Gaussian map covers different distribution regions from small to large. and take the following specific values: , , Then, two 3×3 convolutional layers are used to adaptively combine the N prior maps to generate the final panoramic equatorial prior feature. , and the generated prior map is as shown in Figure 4 .
[0035] For the spatio-temporal features extracted by the spatio-temporal decoupling module , where , , represent the number of channels, width, and height of the spatio-temporal feature respectively. The spatial semantic integration prior module is used to obtain spatial semantic integration prior knowledge from the continuous T frames of the entire batch and integrate the information in the time dimension to obtain prior features. Here, addition in the time dimension is used for temporal information integration. For the integrated spatio-temporal feature , two ({64, 64}×3×3, stride = 2) convolutional layers are used to extract semantic information, where stride represents the convolutional stride, and bilinear interpolation is used for upsampling to obtain the final spatial semantic integration prior map , and the specific formula is: , where and represent the convolutional kernels (the settings of and in this invention take {64, 64}×3×3, stride = 2 as an example), and sum represents the addition operation in the time dimension.
[0036] The prior maps and are concatenated in the channel dimension, and then a 3×3 convolutional layer is used for fusion to obtain the final prior map : , wherein represents the convolutional kernel (the setting of the present invention takes {128, 64}×3×3 as an example).
[0037] Step 5, construct a temporal relationship modeling module for efficiently modeling the inter-frame temporal relationship.
[0038] Concatenate the prior map and the spatio-temporal features in the channel dimension and obtain spatio-temporal prior features through a convolutional layer : , wherein represents the convolutional kernel (the setting of the present invention takes {320, 256}×3×3 as an example).
[0039] As Figure 5 shown, for T consecutive video frames { , ,... } the spatio-temporal prior features obtain the finally predicted saliency map through a temporal post-processing operation, and the generation method of the saliency map is: , , wherein represents the sigmod function, represents the convolutional operation, represents the convolutional kernel (the setting of the present invention takes {512,256}×3×3 as an example), represents the fusion weight, represents the input at the t-th time step, corresponding to the t-th feature in the time dimension, represents the output at the t-th time step.
[0040] Concatenate the outputs of each time step in the time dimension, and then obtain the finally predicted saliency map S through convolutional and sigmod function processing: , wherein represents the convolutional kernel (the setting of the present invention takes {256,1}×3×3 as an example).
[0041] As Figure 6 shown, superimpose the predicted saliency map on the original video frame to obtain the visual prediction result, so as to visually display the saliency map.
[0042] This method is trained and tested based on the publicly available datasets Sports-360 and SVGC-AVA, which are widely used in the field of panoramic video saliency prediction, and five evaluation metrics are used: AUC-J (Area Under the Curve - Judd), NSS (Normalized Scanpath Saliency), KLD (Kullback-Leibler Divergence), SIM (Similarity), and CC (Pearsons Linear Correlation Coefficient) to evaluate the prediction results. As shown in Table 1, compared with the existing state-of-the-art saliency prediction methods (including ordinary video saliency prediction methods and panoramic video saliency prediction methods), this method achieves the optimal performance while maintaining a relatively low number of parameters (25.75M), as shown in Table 1.
[0043] Table 1
[0044]
[0045] The present invention provides a fast panoramic video saliency prediction method. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by existing technologies.
Claims
1. A fast panoramic video saliency prediction method, characterized by It includes the following steps: Step 1, process the panoramic video into video frames; Step 2, take consecutive video frames as input, and extract multi-scale semantic features of each video frame through a basic visual feature extraction module; Step 3, construct a spatio-temporal decoupling module to extract temporal features and spatial features of consecutive frames respectively; Step 4, construct a two-branch attention prior module to generate an equator prior and a spatial semantic integration prior to simulate the visual bias phenomenon in panoramic video viewing behavior; Step 5, input the spatio-temporal features and the prior map into a lightweight temporal relationship modeling module to generate a final saliency map.
2. The method according to claim 1, characterized in that, Step 1 includes: First, parse and process the panoramic video in the equirectangular projection (ERP) format, convert the panoramic video into equirectangular projection ERP video frames, perform downsampling on the equirectangular projection ERP video frames using bilinear interpolation, and adjust the resolution of the equirectangular projection ERP video frames to W×H, where W and H represent the width and height of the projection respectively.
3. The method according to claim 2, wherein Step 2 includes: The basic visual feature extraction module includes a ResNet18 residual network, a dilated convolutional layer, a feature concatenation operation, an upsampling layer, and a convolutional layer.
4. The method according to claim 3, wherein In step 2, the basic visual feature extraction module performs the following operations: taking T consecutive video frames { , ,... } as the input of the Residual Network ResNet18, where represents the T-th video frame. First, obtain the outputs of the 3rd, 4th, and 5th convolutional blocks of the Residual Network ResNet18, denoted as and respectively. Process the output of the 5th convolutional block of the Residual Network ResNet18, where represents the combined batch and temporal dimensions, the number of channels of is 512, and represent the width and height of respectively, and represents the real number space; Process using m = 4 dilated convolutional layers to obtain multi-scale semantic information. The sampling rates r corresponding to the m dilated convolutional layers are respectively set as r = { , , , }, where represents the sampling rate corresponding to the 4th dilated convolutional layer; Subsequently, the multi-scale semantic information is feature concatenated in the channel dimension and processed using 1×1 convolution for channel dimension adjustment to obtain the processed feature , where , , respectively represent the number of channels, width, and height of the feature : , , Among them represents the output of the second convolutional block of the Residual Network ResNet18, , and respectively represent the third, fourth, and fifth convolutional blocks of the Residual Network ResNet18, represents a dilated convolutional kernel with a sampling rate of , represents a feature concatenation operation, and dim represents the dimension of the operation, is the convolutional kernel; Adjust the outputs of the 3rd convolutional block and the 4th convolutional block of the Residual Network ResNet18 using two 1×1 convolutions respectively and the output of the 4th convolutional block in terms of the channel dimension, and then use the bilinear fusion operation to upsample the processed and to the same dimension size. Finally, concatenate , and along the channel dimension, and process them using a 1×1 convolution to obtain the basic features of the video frame : , Among them , and represent convolution kernels, represents the upsampling operation of the upsampling layer.
5. The method according to claim 4, characterized in that Step 3 includes: The spatio-temporal decoupling module includes a dimension separation operation, a spatial convolutional layer, a batch normalization layer, an activation layer, a temporal convolutional layer, and a residual connection; The spatio-temporal decoupling module performs the following operations: First, it performs dimensionality transformation on the basic features of the video frame and separates the batch dimension and the time dimension into two dimensions to obtain the transformed features , , , represent the number of channels, width, and height of the basic features respectively; Then, a structure with serial 2D convolution and 1D convolution is used to extract spatio-temporal features: First, a 3D convolution of 1×3×3 is used to represent the 2D convolution operation to construct a spatial convolution layer for extracting spatial dimension features; Subsequently, the distribution of the spatial dimension features is normalized through a batch normalization layer, and its non-linear expression ability is further enhanced with the help of a ReLU activation function; A 3D convolution of 3×1×1 is used to represent the 1D convolution operation to construct a temporal convolution layer to aggregate the temporal features of consecutive frames along the temporal dimension; Finally, it is processed again through a batch normalization layer and a ReLU activation layer, and a spatio-temporal feature is obtained using a residual connection , and the specific formula is as follows: , where BN represents the batch normalization operation, represents a 1×3×3 convolutional kernel, represents a 3×1×1 convolutional kernel.
6. The method according to claim 5, wherein Step 4 includes: The two-branch attention prior module includes a panoramic equator prior module and a spatial semantic integration prior module. The panoramic equator prior module is used to generate an equator prior for a panoramic video or panoramic image. The generation method of the panoramic equator prior map is: , where y represents the coordinate position in the vertical axis direction, represents the prior value at y, and represent the mean and variance along the vertical axis direction respectively; exp represents the natural exponential function; The calculation formula is as follows: , The calculation formula is as follows: , Among them ; ; Generate corresponding panoramic equatorial prior maps using N different variances, and then adaptively combine the N prior maps through two-layer convolution to generate the final panoramic equatorial prior features ; The spatial semantic integration prior module is used to generate a spatial semantic integration prior map for the spatio-temporal features extracted by the spatio-temporal decoupling module , where , , represent the number of channels, width, and height of the spatio-temporal feature respectively; the spatial semantic integration prior module is used to obtain spatial semantic integration prior knowledge from the continuous T frames of the entire batch, integrate the information in the time dimension to obtain prior features, and use addition in the time dimension for temporal information integration. For the integrated spatio-temporal feature , two convolutional layers are used to extract semantic information, and bilinear interpolation is used for upsampling to obtain the final spatial semantic integration prior map. The specific formula is as follows: , Among them represents an upsampling operation, and represents a convolution kernel, and sum represents an addition operation in the time dimension; Combine and to generate the final prior map through feature splicing and convolution operations : , Among them represents the convolution kernel.
7. The method according to claim 6, wherein Step 5 includes: The prior map and the spatio-temporal features are feature concatenated in the channel dimension and passed through a convolutional layer to obtain spatio-temporal prior features : , Among them represents the convolutional kernel.
8. The method according to claim 7, wherein Step 5 further includes: for the spatio-temporal prior features generated from T consecutive video frames { , ,... }, the final predicted saliency map is obtained through moving weighted average, and the formula is: , , Among them represents the sigmod function represents the convolution operation represents the convolution kernel represents the fusion weight represents the input at the t-th time step, corresponding to the t-th feature in the time dimension; represents the output at the t-th time step; the outputs of each time step are concatenated in the time dimension, and then processed through convolution and the sigmod function to obtain the final predicted saliency map S: , Among them represents the convolution kernel.
9. An electronic device, characterized in that, It includes a processor and a memory. The memory stores program code. When the program code is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 8.
10. A storage medium, characterized in that, Stores a computer program or instruction. When the computer program or instruction runs on a computer, it executes the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Video saliency prediction method and system based on audio and video features
CN116403135A
First-person-view-angle unmanned aerial vehicle video saliency prediction method and system
CN117671538A
Panoramic video saliency prediction method and system
CN117998093A
Panoramic video navigation method driven by subjective preference of user
CN119172634A
Panoramic video processing method and device, electronic equipment and storage medium
CN119893158A
Cited By
Panoramic video saliency prediction method and video compression method based on spherical geometric perception
CN121691682A