A panoramic video enhancement method with stable viewpoint

By acquiring local video sequences and perspective projections from the same viewpoint, and combining them with deformable convolutional neural networks, the problems of low efficiency and insufficient quality in panoramic video encoding are solved, achieving more efficient video quality enhancement.

CN116542889BActive Publication Date: 2025-12-19UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310500426.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-06
Publication Date
2025-12-19
Estimated Expiration
2043-05-06

AI Technical Summary

Technical Problem

Traditional coding techniques cannot effectively process panoramic video, resulting in low coding efficiency and quality degradation. Furthermore, existing neural network methods are not designed specifically for the characteristics of panoramic video, and their quality enhancement effects are insufficient.

Method used

By acquiring local video sequences from the same viewpoint, and utilizing perspective projection and deformable convolutional neural network models, moving targets are captured and their quality enhanced.

Benefits of technology

It improves the encoding efficiency and quality of panoramic video, reduces blockiness and distortion, and enhances video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116542889B_ABST
    Figure CN116542889B_ABST
Patent Text Reader

Abstract

The application discloses a panoramic video enhancement method with a stable viewpoint, comprising the following steps: S1, perspective projection is performed on an ERP format panoramic video to obtain a view window with the same field of view angle as a user; S2, a plurality of video sequences with the same viewpoint are collected as a training set; S3, 2T+1 perspective projection frames with the same viewpoint are taken as the input of an enhancement network model, a group of features are obtained after the input, the features are divided into an offset and the importance of the offset, fusion is performed through deformable convolution, then a mask Mask is obtained through L-layer convolution with a step of 1, and finally, an addition operation in pixel units is performed on the mask Mask and a target frame A t to obtain a final quality enhancement result. The application collects local video sequences with the same viewpoint to obtain a plurality of frames of information with the same background, thereby avoiding the problem of poor prediction accuracy of the viewpoint sequence; meanwhile, the neural network model using deformable convolution can more easily capture and obtain a moving target.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of image processing, and particularly relates to a panoramic video enhancement method with a stable viewpoint. BACKGROUND

[0002] Traditional video coding techniques are mainly based on planar images, that is, images are arranged in a certain order to form a planar image sequence, and then transmitted through coding compression. 360° video is a panoramic video, which presents a spherical image, which leads to the fact that the coding algorithm used in traditional coding techniques cannot be directly applied to 360° video. By projecting the spherical information to a plane and converting it into ERP or CMP format, H.264, H.265 and other technologies can be used to code it. However, due to the geometric distortion of the ERP format in the area far from the equator and the discontinuous boundary of the CMP format, their coding effect is inferior to that of ordinary 2D video.

[0003] The mismatch between traditional coding methods and 360° video can be compensated by perspective projection, which clips the panoramic image according to the specified angle size and position of the user, so that the encoder only needs to process the image information under a specific angle to improve the coding efficiency. However, the block-based coding strategy used in traditional coding leads to the fact that when there is a large area of texture or noise in the image, this strategy may divide these textures or noises into multiple blocks for coding, thereby producing block effects and causing obvious flaws and distortion in the image. Moreover, in high motion scenes, due to the large change in the position of objects between video frames, the difference between adjacent frames is also larger, so it is difficult for the encoder to accurately predict the content of the next frame, resulting in a significant increase in code rate and a decrease in video quality.

[0004] The patent application with the application number 201710878189.8 discloses an image or video quality enhancement method based on convolutional neural network. First, two convolutional neural networks for video quality enhancement are designed, and the two networks have different computational complexities. Then, a number of training images or videos are selected to train the parameters in the two convolutional neural networks. According to actual needs, a convolutional neural network with a relatively appropriate computational complexity is selected, and an image or video whose quality needs to be enhanced is input into the selected network. Finally, the network outputs the image or video after quality enhancement. This scheme can effectively enhance the video quality. Users can specify the convolutional neural network with a relatively appropriate computational complexity to enhance the quality of images or videos according to the computing power or remaining power of the device. This scheme designs two convolutional neural networks with different complexities, and users select the network according to the device. The difference between the two networks is only the depth of the convolutional neural network. It is not feasible to improve the quality enhancement effect by deepening the network depth. Moreover, the network is not designed for the characteristics of panoramic videos, and the quality enhancement effect needs to be improved.

[0005] The patent application with the application number 201910554229.2 discloses a fuzzy video super-resolution method and system based on deep learning. On the basis of a single-frame depth back-projection super-resolution model, a multi-frame fuzzy video super-resolution model is designed to improve the fuzzy video super-resolution reconstruction quality and support high-multiple (x8) reconstruction. In view of the problem that the edge contour and other detail information of a video after motion blur video super-resolution reconstruction is not clear and the video quality is low, the invention introduces recursive learning and multi-frame fusion strategy into the depth back-projection super-resolution model to construct a fuzzy video super-resolution model. The model can reconstruct an edge contour clear super-resolution video by learning the non-linear mapping from a fuzzy low-resolution video frame to a clear high-resolution video frame, improve the quality of motion blur video super-resolution reconstruction, and enable people to better obtain video information. This scheme proposes a multi-frame fuzzy video super-resolution model on the basis of single-frame super-resolution, uses an adversarial network and an optical flow scheme to remove the blur of a low-resolution video. It is difficult to obtain accurate inter-frame motion information by using the optical flow estimation scheme on a low-resolution image, so the motion compensation obtained is not accurate enough, which may cause distortion of the high-quality frame recovered finally. Moreover, the resolution of a panoramic video is greater than that of an ordinary video, and a small resolution image will cause serious distortion.

[0006] The patent application with the application number 201810603510.6 discloses a video quality enhancement method based on adaptive separable convolution. Adaptive separable convolution is applied as the first module in the network model, each two-dimensional convolution is converted into a pair of one-dimensional convolution kernels in the horizontal direction and the vertical direction, the parameter quantity is reduced from n 2The second is to use the adaptive change of the convolution kernel learned by the network for different inputs to realize the estimation of the motion vector, and by selecting two consecutive frames as the network input, a pair of separable two-dimensional convolution kernels can be obtained for each two consecutive inputs, and then the two-dimensional convolution kernel is unfolded into four one-dimensional convolution kernels, and the obtained one-dimensional convolution kernel changes with the change of the input, thereby improving the adaptability of the network. The application uses one-dimensional convolution kernel to replace two-dimensional convolution kernel, so that the network training model parameters are reduced, and the execution efficiency is high. The scheme uses five encoding modules and four decoding modules, a separation convolution module and an image prediction module, and the structure is that the last decoding module is replaced by a separation convolution module on the basis of the traditional symmetrical coding and decoding module network, although the parameters of the model are effectively reduced, but the quality enhancement effect still needs to be further improved.

[0007] With the development of the metaverse industry, more and more panoramic videos and images are produced. Panoramic videos usually need to be projected to obtain a small area with a normal field of view for users to watch, which results in that the panoramic video usually needs a larger field of view to ensure the user experience. At this time, video compression can effectively reduce the data amount, but inevitably leads to the decrease of video quality. SUMMARY

[0008] The present application aims to overcome the shortcomings of the prior art and provide a panoramic video enhancement method with a stable viewpoint. The present application fully utilizes the shooting characteristics of panoramic videos, acquires multiple frames of information with the same background by collecting local video sequences with the same viewpoint, and avoids the problem of poor prediction accuracy of the viewpoint sequence. At the same time, the use of a deformable convolution neural network model can more easily capture and acquire moving targets, so that the information of adjacent frames can better enhance the quality of the target frame.

[0009] The present application aims to overcome the shortcomings of the prior art and provide a panoramic video enhancement method with a stable viewpoint. The present application fully utilizes the shooting characteristics of panoramic videos, acquires multiple frames of information with the same background by collecting local video sequences with the same viewpoint, and avoids the problem of poor prediction accuracy of the viewpoint sequence. At the same time, the use of a deformable convolution neural network model can more easily capture and acquire moving targets, so that the information of adjacent frames can better enhance the quality of the target frame.

[0010] S1, perspective projection is performed on the panoramic video in ERP format to obtain a view window with the same field of view angle as that of the user; it is assumed that the video segment is V T , wherein T represents the number of panoramic video frames possessed by the panoramic video segment, and the t-th frame is represented as V t ;

[0011] The perspective projection operation is performed on V t , including the following sub-steps:

[0012] S11, each pixel point in the plane to be projected is converted into a three-dimensional coordinate: the horizontal and vertical sizes of the user's field of view angle are set as fov x and fov yThe size of the window resolution obtained by the projection is set as port w and port h The focal length f of the plane to be projected in the X and Y directions is calculated by formula (1) and (2) x and f y :

[0013]

[0014]

[0015] According to the above information, the intrinsic matrix of the camera is obtained:

[0016]

[0017] Among them, the two data of (1, 3) and (2, 3) in the matrix represent the coordinate position of the principal point on the image plane, that is, the position of the camera in the intrinsic matrix of the camera;

[0018] A grid with the same size as the window obtained by the projection is created, and two one-dimensional indexes u mesh and v mesh corresponding to the grid are obtained; by formula (4), 1 and u mesh , v mesh are spliced to obtain the matrix e:

[0019]

[0020] Among them, 1 represents that the z axis in the three-dimensional coordinates is 1; then the homogeneous projection coordinates q are obtained:

[0021] q=K -1 e (5)

[0022] After the projection coordinates q are normalized, the following is obtained:

[0023]

[0024] And multiply a diagonal matrix to flip the z axis coordinates to obtain the non-homogeneous projection coordinate grid:

[0025]

[0026] S12, according to the set window latitude and longitude The pixel position required by the projection plane in the spherical surface is calculated: first, the rotation matrix is calculated according to the given latitude and longitude point:

[0027]

[0028] Then the projection coordinate grid is rotated and translated by equation (9) to make the grid rotate to the corresponding position:

[0029] E = RP T (9)

[0030] where P T represents the transpose of matrix P;

[0031] Convert E to the latitude and longitude format in radian form:

[0032]

[0033]

[0034] where E1, E2 and E3 represent the 1st, 2nd and 3rd row data in matrix E respectively;

[0035] Correct the obtained pixel position:

[0036]

[0037]

[0038] Finally, a grid of size (port w , port h ) corresponding to the viewpoint position is obtained:

[0039]

[0040] Assign the pixels in the ERP format video frame to the grid, and use bilinear interpolation to fill the pixels with offset positions;

[0041] S2, collect multiple frames of video sequences with the same viewpoint, and process the video sequences according to step S1 as a training set;

[0042] S3, take 2T+1 perspective projection frames with the same viewpoint as the input of the enhancement network model, respectively perform downsampling and upsampling, then perform stitching operation on the features with the same dimension to obtain a group of features; divide the group of features into two parts, one part is offset, and the other part is the importance of the offset; the offset consists of two parts, which are the offset horizontal and vertical coordinates; then the offset, the importance of the offset and the multiple frame information of the original input are fused through deformable convolution; then after L layers of convolution with a step size of 1, the mask Mask is obtained, and finally the mask Mask is added to the target frame A t to obtain the final quality enhancement result.

[0043] The beneficial effects of the present application are: the present application makes full use of the shooting characteristics of panoramic video, acquires multiple frames of information with the same background by collecting local video sequences with the same viewpoint, avoids the problem of poor prediction viewpoint sequence accuracy; at the same time, the neural network model using deformable convolution can more easily capture and acquire moving targets, so that the information of adjacent frames can better enhance the quality of the target frame, thereby achieving better overall effect. BRIEF DESCRIPTION OF DRAWINGS

[0044] Figure 1 Video frames collected according to real viewpoint movement rules;

[0045] Figure 2 Structure diagram of video enhancement network;

[0046] Figure 3 Comparison diagram of average ΔPSNR performance on QP27, 32, 37 and 42. DETAILED DESCRIPTION

[0047] Abbreviations and key terms definition:

[0048] 1. H.264 / MPEG-4 AVC: It is a new video coding standard formulated by ITU-T VCEG after H.264. The H.265 standard is based on the existing video coding standard H.264, retains some of the original technologies, and improves some related technologies. The new technology uses advanced techniques to improve the relationship between bitstream, encoding quality, delay and algorithm complexity, and achieve optimal settings.

[0049] 2. PSNR (Peak Signal to Noise Ratio): Peak signal to noise ratio, an objective standard for evaluating images.

[0050] 3. ERP (Equi-Rectangular Projection): Equi-Rectangular Projection, a projection method for mapping spherical information to a single plane.

[0051] 4. CMP (Cube Map Projection): Cube Map Projection, a projection method for placing a sphere in a cube and mapping it to six independent faces.

[0052] The technical solutions of the present application will be further described below in conjunction with the drawings.

[0053] A panoramic video enhancement method with a stable viewpoint of the present application comprises the following steps:

[0054] S1, in order to extract the same background of continuous frame, need to ERP format panoramic video perspective projection, get the same field of view angle of the user has the window; assume that the video segment for V T , wherein T represents the panoramic video frame number of the panoramic video segment, wherein the t frame is represented as V t ;

[0055] The perspective projection operation is performed on V t , including the following sub-steps:

[0056] S11, each pixel point in the plane to be projected is converted into three-dimensional coordinates: set the horizontal and vertical size of the user's field of view angle as fov x And fov y , the size of the resolution of the window obtained by projection is set as port w And port h ; the focal length f x And f y of the X and Y directions of the plane to be projected is calculated by formula (1) and (2):

[0057]

[0058]

[0059] According to the above information, the intrinsic matrix of the camera is obtained:

[0060]

[0061] Among them, (1, 3) and (2, 3) in the matrix represent the coordinate position of the principal point on the image plane, that is, the position of the camera in the intrinsic matrix of the camera;

[0062] A grid with the same size as the window obtained by projection is created, and two one-dimensional indexes u mesh And v mesh corresponding to the grid are obtained; by formula (4), 1 and u mesh , v mesh are spliced to obtain matrix e:

[0063]

[0064] Among them, 1 represents that the z axis in the three-dimensional coordinates is 1; then the homogeneous projection coordinates q are obtained:

[0065] q=K -1 e (5)

[0066] After the normalization operation of the projection coordinates q is performed, the following is obtained:

[0067]

[0068] and multiply by a diagonal matrix to flip the z-axis coordinate to get the non-homogeneous projection coordinate grid:

[0069]

[0070] S12, according to the set window latitude and longitude Calculate the pixel position needed in the projection plane in the sphere: first calculate the rotation matrix according to the given latitude and longitude point:

[0071]

[0072] Then rotate and translate the projection coordinate grid by formula (9) to make the grid rotate to the corresponding position:

[0073] E = RP T (9)

[0074] Where P T represents the transpose of matrix P;

[0075] Convert E to the latitude and longitude format in radian form:

[0076]

[0077]

[0078] Where E1, E2 and E3 represent the 1st, 2nd and 3rd row data in matrix E, respectively;

[0079] Finally, due to the geometric distortion existing in the ERP format, the obtained pixel position needs to be corrected:

[0080]

[0081]

[0082] Finally, a grid of size (port w , port h ) corresponding to the viewpoint position is obtained:

[0083]

[0084] Assign the pixels in the ERP format video frame to the grid; the information in the grid represents the mapping relationship between the position in the projection result plane grid and the pixel in the ERP image. Through the mapping relationship, the corresponding pixel in the ERP image is taken out and placed in the corresponding position in the grid. And bilinear interpolation is used to fill the pixels with positional offset;

[0085] S2, collect a plurality of video sequences of the same viewpoint, as shown in Figure 1 , and perform the processing of step S1 on the video sequences as a training set; from Figure 1 , it can be seen from the comparison of the two rows of pictures that the background in the picture in the first row has not changed, and only the object in the lower left corner has moved. The mountains in the image in the second row have moved to a certain extent, and this has also caused a certain loss of the moving object. From the comparison of the two groups of pictures, it can be known that the complexity of enhancing the moving object in the image in the first row is lower than that in the image in the second row. Based on this, a plurality of video sequences of the same viewpoint are collected for training and subsequent testing.

[0086] S3, taking 2T+1 perspective projection frames (T frames in time order and T frames in reverse time order, and the target frame) with the same viewpoint as the input of the enhancement network model, the structure of the enhancement network model is as shown in Figure 2 , the Convolution of S1 plays the role of feature extraction, the Convolution of S2 represents convolution and plays the role of down-sampling, the Deconvolution of S2 represents deconvolution and has the role of up-sampling, and the addition represents the operation of summing two matrices in pixel units. The enhancement network model performs feature extraction, down-sampling and up-sampling through a structure similar to U-Net, and then performs a splicing operation on the features with the same dimension through a skip connection to obtain a group of features (i.e. the offset field in the figure); the group of features is divided into two parts, one part is the offset, and the other part is the importance of the offset; the offset is composed of two parts, which are the horizontal and vertical coordinates of the offset; then the offset, the importance of the offset and the multi-frame information of the original input are fused through deformable convolution; then after L layers of convolution with a step size of 1, the mask Mask is obtained, and finally the addition operation with the target frame A t in pixel units is performed to obtain the final quality enhancement result.

[0087] The loss function calculation method of the enhancement network model is: the purpose of this network is to make the quality of the compressed frame close to the quality of the original frame, so the sum of squared errors is used as the loss function of the model:

[0088]

[0089] wherein is the predicted result, E t is the original frame. The difference between the predicted value and the label value is calculated by the loss function, and then the gradient is returned according to the difference and the parameters in the quality enhancement network are updated according to the gradient.

[0090] The enhancement effect of the present application is further tested by experiments.

[0091] Using the VQA ODV dataset, which contains 10 groups, each group has 6 videos, we randomly selected 46 from 60 videos as training videos, and the remaining 14 as test videos. All sequences are compressed using H.265. In order to form a comparison, the method of the present application is compared with the V-DNN method, and only at the position where the latitude and longitude are both 0, a sequence is extracted for each training video; while the TOMM method is to collect the view sequence predicted by the model. The same training is carried out in these two sequence sets respectively, and the comparison is carried out, as shown in Table 1.

[0092] Table 1

[0093]

[0094] Table 1 gives the specific performance of the two models on PSNR of each test sequence under the condition of QP37 compression. It can be seen that V-DNN has a negative gain in sequence No. 4, while the method proposed by us can have a positive gain in all sequences. In addition, from the variance data in the second last row, the stability of our method is better than V-DNN, and in the average performance in the last row, the present application has an improvement of 18.2%.

[0095] Figure 3 The average ΔPSNR on all frames of each test sequence is given. It can be seen that the method proposed by us is better than V-DNN, the average PSNR is 0.248655, which is 0.053515 better than V-DNN, with an improvement of 27.4%. Both methods reach the best at QP32, and the improvement at QP27 reaches 45%. Overall, the method of the present application has more stable and better performance than V-DNN in the test QP.

[0096] Those skilled in the art will appreciate that the embodiments described herein are presented for the purpose of helping the reader to understand the principles of the present application, and should be understood as not limiting the scope of protection of the present application to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations according to the technical inspiration disclosed in the present application without departing from the essence of the present application, and these modifications and combinations are still within the scope of protection of the present application.

Claims

1. A method for enhancing a panoramic video with a stable viewpoint, characterized in that, Comprising the following steps: S1, perspective projection is performed on the panoramic video in the ERP format to obtain a view window with the same field of view as that of the user; it is assumed that the video segment is wherein represents the number of panoramic video frames possessed by the video segment, wherein the first frame is represented as ;​ right Performing perspective projection includes the following sub-steps: S11, converting each pixel point in the plane to be projected into a three-dimensional coordinate: set the horizontal and vertical sizes of the user's field of view angle as and , and set the size of the resolution of the window obtained by projection as and ; calculate the focal length of the X and Y directions of the plane to be projected by formulas (1) and (2) and : ; ; According to the above information, the intrinsic matrix of the camera is obtained: ; Wherein, the two data of (1, 3) and (2, 3) in the matrix represent the coordinate position of the principal point on the image plane, that is, the position of the camera in the intrinsic matrix of the camera; Create a grid with the same size as the window size of the projection, and get two one-dimensional indexes corresponding to the grid and ; splice the matrix , by formula (4) to get the matrix : ; where 1 represents that the z axis in the three-dimensional coordinates is all 1; and then the homogeneous projection coordinates are obtained : ; The projection coordinates are then obtained as follows The normalization operation is performed to obtain: ; And multiply a diagonal matrix to flip the z-axis coordinate to obtain a non-homogeneous projection coordinate grid: ; S12. Calculate the pixel position in the sphere that the projection plane needs according to the set window longitude and latitude , ) Calculate the pixel position in the sphere that the projection plane needs according to the set window longitude and latitude: First, calculate the rotation matrix according to the given longitude and latitude point: ; Then the projection coordinate grid is rotated and translated by formula (9) to rotate the grid to the corresponding position: ; wherein denotes the transpose of the matrix ; Convert to latitude-longitude format in radians: Convert to latitude-longitude format in radians: ; ; wherein , and represent the 1st, 2nd and 3rd row data in the matrix , respectively; Correct the obtained pixel position: ; ; A mesh of size is finally obtained for the view position against which it is meshed: ; Assign the pixels in the ERP format video frame to the grid, and use bilinear interpolation to fill the pixels with position offset; S2, collect multiple frames of video sequences of the same viewpoint, and process the video sequences according to step S1 as a training set; S3, 2T+1 perspective projection frames with the same viewpoint are taken as the input of the enhancement network model, respectively down-sampled and up-sampled, then a group of features are obtained after the features with the same dimension are spliced; the group of features are divided into two parts, one part is offset, and the other part is the importance of the offset; the offset is composed of two parts, which are the horizontal and vertical coordinates of the offset; then the offset, the importance of the offset and the multi-frame information of the original input are fused through deformable convolution; then after L layers of convolution with a step of 1, the mask Mask is obtained, and finally the mask Mask is added to the target frame in pixel units to obtain the final quality enhancement result.

Citation Information

Patent Citations

  • Image or video quality enhancement method based on convolution neural networks

    CN107481209A

  • Video quality enhancement method based on adaptive separable convolution

    CN108900848A

  • Blurred video super-resolution method and system based on deep learning

    CN110458756A