Immersive Video Enhancement Method and Device Based on Frequency-Domain Boundary Collaborative Optimization

Through the immersive video enhancement method based on frequency domain boundary collaborative optimization, the problem of distortion easily introduced in immersive videos during compression is solved, and the video quality improvement and artifact reduction is achieved, which significantly improves the user experience.

CN119850441BActive Publication Date: 2025-06-24HUAQIAO UNIVERSITY +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510317059.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-24
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

The data volume of immersive videos is huge, involving ultra-high resolution and complex multi-view processing, which leads to the introduction of compression distortion and artifacts during acquisition, compression, transmission and rendering, affecting the user's viewing experience.

Method used

A method of immersive video enhancement based on frequency domain boundary collaborative optimization is proposed. By constructing an immersive video enhancement model, dynamically fuse frequency domain features, boundary features and spatiotemporal information, use the frequency domain enhancement module and boundary enhancement module to optimize video quality, reduce compression artifacts, and further improve video quality through the spatiotemporal deformable convolution module and quality enhancement module.

Benefits of technology

It effectively improves the quality of immersive videos, reduces compression artifacts and boundary artifacts, improves the time and space consistency and detail recovery capabilities of the video, and significantly improves the user's viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119850441B_ABST
    Figure CN119850441B_ABST
Patent Text Reader

Abstract

The present invention discloses an immersive video enhancement method and apparatus based on frequency-domain boundary collaborative optimization, which relates to the field of video processing and includes: obtaining a compressed multi-view texture plus depth video sequence to be reconstructed and inputting it into a trained immersive video enhancement model; the current video frame to be enhanced first passes through a feature extraction module to respectively extract high-frequency features and low-frequency features; the high-frequency features and low-frequency features pass through a frequency-domain enhancement module to obtain a frequency-domain enhanced image; the frequency-domain enhanced image and the current video frame to be enhanced are input into a boundary enhancement module to obtain a fused image; the fused image and adjacent video frames of the current video frame to be enhanced are input into a spatio-temporal deformable convolution module to obtain an aligned fused image, and the aligned fused image passes through a quality enhancement module to predict an enhancement residual and generate a corresponding reconstructed video. The present invention solves problems such as compression artifacts, boundary artifacts, and low quality of immersive videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video processing, and in particular to an immersive video enhancement method and device based on frequency-domain boundary collaborative optimization. Background Art

[0002] Immersive video is a type of emerging media video that can provide a high degree of freedom, and can provide users with a highly immersive and realistic visual experience. By fusing multi-viewpoint texture and depth information, immersive video supports users to view three-dimensional scenes with 6 degrees of freedom, enabling users to perceive the spatial hierarchy and three-dimensional effect of the scene from any perspective. However, the data volume of immersive video is huge, involving ultra-high resolution and complex multi-viewpoint processing, and problems such as compression distortion and artifacts are easily introduced during the processes of acquisition, compression, transmission, and rendering, thus affecting the viewing experience of users. This poses higher requirements for video quality improvement technologies and is also a key challenge for promoting the development of immersive video applications.

[0003] Multiview Video plus Depth (MVD) is a mainstream immersive video data format, which consists of texture videos collected from multiple viewpoints and corresponding depth videos obtained through depth estimation or depth cameras, etc. Due to the complexity and diversity of MVD videos in the spatial, temporal, and viewpoint dimensions, various compression distortions such as blurring and artifacts in the compressed MVD videos are more significant, seriously affecting the immersive experience of users. Therefore, it is particularly important to propose an algorithm that conforms to the human visual characteristics and can accurately and quickly enhance the quality of immersive videos.

[0004] Currently, research on immersive video quality enhancement mostly focuses on single-frame or single-viewpoint video reconstruction, ignoring the role of multi-viewpoint fusion and inter-frame information in quality restoration. In addition, current research generally lacks sufficient consideration of the unique characteristics of immersive videos. Therefore, designing a video quality enhancement algorithm that combines human visual characteristics with the characteristics of immersive videos can not only provide support for improving the quality of immersive videos theoretically, but also has significant practical application value, and is of great significance for promoting the popularization and industrialization of immersive video technologies. Summary of the Invention

[0005] The purpose of this application is to propose an immersive video enhancement method and device based on frequency-domain boundary collaborative optimization for the above-mentioned technical problems.

[0006] In a first aspect, the present invention provides an immersive video enhancement method based on frequency-domain boundary collaborative optimization, including the following steps:

[0007] Build an immersive video enhancement model and train it to obtain a trained immersive video enhancement model. The immersive video enhancement model includes a feature extraction module, a frequency domain enhancement module, a boundary enhancement module, a spatio-temporal deformable convolution module, and a quality enhancement module connected in sequence;

[0008] Obtain a compressed multi-view texture plus depth video sequence to be reconstructed and input it into the trained immersive video enhancement model; for each current video frame to be enhanced in the video sequences to be enhanced in the compressed multi-view texture plus depth video sequence to be reconstructed, first perform frequency domain transformation and feature extraction through the feature extraction module to respectively obtain high-frequency features and low-frequency features; the high-frequency features and low-frequency features pass through the frequency domain enhancement module to be respectively assigned corresponding high-frequency weights and low-frequency weights, and then perform weighted fusion and inverse transformation to obtain a frequency domain enhanced image; the frequency domain enhanced image and the current video frame to be enhanced are input into the boundary enhancement module to obtain a fused image; the fused image and the adjacent video frames of the current video frame to be enhanced in the video sequence to be enhanced are input into the spatio-temporal deformable convolution module to align the time information using the inter-frame complementary relationship to obtain an aligned fused image, and the aligned fused image passes through the quality enhancement module to predict an enhancement residual and generate a corresponding reconstructed video.

[0009] Preferably, for each current video frame to be enhanced in the video sequences to be enhanced in the compressed multi-view texture plus depth video sequence to be reconstructed, first perform frequency domain transformation and feature extraction through the feature extraction module to respectively obtain high-frequency features and low-frequency features, specifically including:

[0010] For the video sequence to be enhanced , denotes the current video frame to be enhanced, denotes the th frames before and after the current video frame to be enhanced, denotes the adjacent video frame of the th frame before the current video frame to be enhanced, denotes the adjacent video frame of the th frame after the current video frame to be enhanced;

[0011] Perform Fourier transform on the current video frame to be enhanced to obtain the corresponding frequency domain features, as shown in the following formula:

[0012] ;

[0013] where, denotes the pixel value at the pixel coordinate in the current video frame to be enhanced, denotes the frequency domain feature corresponding to the current video frame to be enhanced at the frequency domain coordinate , and respectively represent the width and height of the current video frame to be enhanced, represents the Fourier transform;

[0014] The low-frequency features and high-frequency features are separated from the frequency-domain features through Gaussian filtering, as shown in the following formula:

[0015] ;

[0016] ;

[0017] where, represents the Gaussian filtering function, represents the low-frequency features corresponding to the current video frame to be enhanced located at the frequency-domain coordinates ; represents the high-frequency features corresponding to the current video frame to be enhanced located at the frequency-domain coordinates ;

[0018] Preferably, the high-frequency features and low-frequency features are respectively given corresponding high-frequency weights and low-frequency weights through the frequency-domain enhancement module, and then weighted fusion and inverse transformation are performed to obtain the frequency-domain enhanced image, which specifically includes:

[0019] For the low-frequency features, the low-frequency weights are calculated using the following formula:

[0020] ;

[0021] where, represents the low-frequency weights located at the frequency-domain coordinates ; is the low-frequency enhancement factor, , is the smoothing term, represents the Gaussian filtering function, represents the frequency-domain features corresponding to the current video frame to be enhanced located at the frequency-domain coordinates ; represents the view point importance weights located at the frequency-domain coordinates , and its expression is as follows:

[0022] ;

[0023] where, represents the frequency-domain coordinates of the center frequency of the user's view point area, represents the parameter of the user's view point area, and its expression is: , represents the empirical factor, represents the maximum value in the frequency-domain features corresponding to the current video frame to be enhanced located at any frequency-domain coordinate;

[0024] For high-frequency features, the high-frequency weight is calculated using the following formula:

[0025] ;

[0026] where is the high-frequency weight at the frequency-domain coordinate , is the high-frequency retention factor , represents the motion saliency weight at the frequency-domain coordinate , and its expression is as follows:

[0027] ;

[0028] where represents the pixel change amplitude between the current video frame to be enhanced and its previous video frame at the pixel coordinate , represents the pixel change amplitude between the current video frame to be enhanced and its previous video frame at any pixel coordinate , and max represents taking the maximum value;

[0029] The low-frequency features and high-frequency features are weighted using the low-frequency weight and high-frequency weight respectively to obtain the enhanced frequency-domain features, as shown in the following formula:

[0030] ;

[0031] where represents the enhanced frequency-domain features corresponding to the current video frame to be enhanced;

[0032] The enhanced frequency-domain features are returned to the spatial domain through inverse Fourier transform to obtain the frequency-domain enhanced image, as shown in the following formula:

[0033] ;

[0034] where represents the pixel value at the pixel coordinate in the frequency-domain enhanced image corresponding to the current video frame to be enhanced, and represent the width and height of the current video frame to be enhanced respectively.

[0035] Preferably, the frequency-domain enhanced image and the current video frame to be enhanced are input into the boundary enhancement module to obtain a fused image, specifically including:

[0036] Use the Canny edge detection algorithm to extract the significant boundary regions in the frequency-domain enhanced image to obtain boundary features, as shown in the following formula:

[0037] ;

[0038] Among them, represents the pixel coordinates corresponding to the current video frame to be enhanced of the boundary feature, represents the Canny edge detection algorithm;

[0039] Calculate the pixel interpolation of the neighboring region in the current video frame to be enhanced, and estimate the artifact intensity, as shown in the following formula:

[0040] ;

[0041] Among them, represents the artifact intensity at the pixel coordinates in the current video frame to be enhanced, represents the pixel value at the pixel coordinates in the current video frame to be enhanced, represents the smoothed estimate of the neighboring pixels at the pixel coordinates in the current video frame to be enhanced, and its expression is as follows:

[0042] ;

[0043] Among them, represents the around the pixel coordinates neighborhood, the neighboring pixel coordinates at the pixel coordinates in the current video frame to be enhanced of the pixel value, represents taking the median of all pixel values in the neighborhood;

[0044] Weight the boundary feature and the artifact intensity to calculate the boundary weight, as shown in the following formula:

[0045] ;

[0046] Among them, is the boundary enhancement factor, represents the artifact influence factor;

[0047] Multiply the boundary weight by the frequency domain enhanced image corresponding to the current video frame to be enhanced to obtain the fused image, as shown in the following formula:

[0048] ;

[0049] Among them, represents the fused image corresponding to the current video frame to be enhanced at the pixel coordinates The pixel value.

[0050] Preferably, the adjacent video frames of the fused image and the current video frame to be enhanced in the video sequence to be enhanced are input into the spatio-temporal deformable convolution module, and the temporal information is aligned using the inter-frame complementary relationship to obtain the aligned fused image. The aligned fused image passes through the quality enhancement module to predict the enhancement residual and generate the corresponding reconstructed video, specifically including:

[0051] The image features pass through the spatio-temporal deformable convolution module to obtain the aligned fused image, as shown in the following formula:

[0052]

[0053] Where, represents the size of the convolution kernel, represents the th channel of the convolution kernel, represents the aligned fused image corresponding to the current video frame to be enhanced, represents the fused image corresponding to the current video frame to be enhanced, represents any spatial position, represents the th channel of the convolution kernel corresponding to the regular sampling offset, represents the deformable offset corresponding to the time position being the Tth frame and the spatial position being , which is extracted by jointly extracting the offset fields at all time positions and spatial positions within the current video frame to be enhanced. The expression of the offset field is as follows:

[0054]

[0055] Where, represents the U-net offset prediction network, represents the th frame adjacent video frame before the current video frame to be enhanced, represents the th frame adjacent video frame after the current video frame to be enhanced;

[0056] Input the aligned fused image into the quality enhancement module composed of multi-level residual dense channel attention blocks to predict the enhancement residual , and add the enhancement residual element-wise to the current video frame to be enhanced to generate the reconstructed video frame , as shown in the following formula:

[0057] ;

[0058] Where, Denotes element-wise addition. All the reconstructed video frames constitute a reconstructed video sequence and are synthesized into a reconstructed video.

[0059] Preferably, the total loss function used in the training process of the immersive video enhancement model is the weighted result of the reconstruction loss and the boundary loss, as shown in the following formula:

[0060] ;

[0061] Wherein, Denotes the total loss function, and Denote the weights of the reconstruction loss and the boundary loss respectively, and Denote the reconstruction loss and the boundary loss respectively;

[0062] The calculation formula of the reconstruction loss is as follows:

[0063] ;

[0064] Wherein, Denotes the total number of pixels, and Denote the pixel value of the i-th pixel of the reconstructed video frame and the pixel value of the i-th pixel of the original image respectively;

[0065] The calculation formula of the boundary loss is as follows:

[0066] ;

[0067] Wherein, Denotes the gradient, which is obtained by the Sobel edge detection operator, and Denote the gradient value of the i-th pixel of the reconstructed video frame and the gradient value of the i-th pixel of the original image respectively.

[0068] In a second aspect, the present invention provides an immersive video enhancement device based on frequency-domain boundary collaborative optimization, including:

[0069] A model construction module configured to construct and train an immersive video enhancement model to obtain a trained immersive video enhancement model. The immersive video enhancement model includes a feature extraction module, a frequency-domain enhancement module, a boundary enhancement module, a spatio-temporal deformable convolution module, and a quality enhancement module connected in sequence;

[0070] The reconstruction module is configured to obtain a compressed multi-view texture plus depth video sequence to be reconstructed and input it into a trained immersive video enhancement model; for each current video frame to be enhanced in the video sequences to be enhanced in the compressed multi-view texture plus depth video sequence to be reconstructed, frequency domain transformation and feature extraction are first performed by the feature extraction module to respectively obtain high-frequency features and low-frequency features; the high-frequency features and low-frequency features are respectively given corresponding high-frequency weights and low-frequency weights by the frequency domain enhancement module, and weighted fusion and inverse transformation are performed to obtain a frequency domain enhanced image; the frequency domain enhanced image and the current video frame to be enhanced are input into the boundary enhancement module to obtain a fused image; the fused image and adjacent video frames of the current video frame to be enhanced in the video sequence to be enhanced are input into the spatio-temporal deformable convolution module, and the frame-interframe complementary relationship is used to align the time information to obtain an aligned fused image, and the aligned fused image passes through the quality enhancement module to predict an enhancement residual and generate a corresponding reconstructed video.

[0071] In a third aspect, the present invention provides an electronic device, including one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method described in any implementation manner of the first aspect.

[0072] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.

[0073] In a fifth aspect, the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.

[0074] Compared with the prior art, the present invention has the following beneficial effects:

[0075] (1) The immersive video enhancement method based on frequency domain boundary collaborative optimization proposed by the present invention effectively improves the quality of immersive videos by constructing an immersive video enhancement model to dynamically fuse frequency domain features, boundary features, and spatio-temporal information.

[0076] (2) The immersive video enhancement method based on frequency domain boundary collaborative optimization proposed by the present invention uses the frequency domain enhancement module to extract important information related to immersive video compression distortion from high-frequency features and low-frequency features, optimizes global structure perception and motion details, effectively reduces compression artifacts, and uses edge detection and gradient information in the boundary enhancement module to focus on processing the foreground and background junction regions and strengthen the reconstruction ability of geometric boundary details, effectively suppressing boundary artifacts.

[0077] (3)The immersive video enhancement method based on frequency-domain boundary collaborative optimization proposed by the present invention realizes the collaborative enhancement of the frequency domain and the boundary by dynamically aggregating the frequency-enhanced image and the boundary weight, combines the inter-frame complementary relationship and the view-point weighting strategy, optimizes the spatio-temporal consistency and detail restoration of the multi-viewpoint immersive video, comprehensively improves the video quality and visual immersion, and effectively improves the immersive video quality enhancement performance and the overall quality of the video through the deep fusion of the view-point information and the dynamic alignment of the inter-frame complementary relationship. Description of the Drawings

[0078] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0079] Figure 1 It is a schematic flowchart of the immersive video enhancement method based on frequency-domain boundary collaborative optimization according to the embodiment of the present application;

[0080] Figure 2 It is a schematic diagram of the immersive video enhancement model of the immersive video enhancement method based on frequency-domain boundary collaborative optimization according to the embodiment of the present application;

[0081] Figure 3 It is a schematic diagram of the immersive video enhancement device based on frequency-domain boundary collaborative optimization according to the embodiment of the present application;

[0082] Figure 4 It is a schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. Detailed Embodiments

[0083] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0084] Figure 1 There is shown an immersive video enhancement method provided by the embodiment of the present application, including the following steps:

[0085] S1. Construct an immersive video enhancement model and train it to obtain a trained immersive video enhancement model. The immersive video enhancement model includes a feature extraction module, a frequency-domain enhancement module, a boundary enhancement module, a spatio-temporal deformable convolution module, and a quality enhancement module connected in sequence.

[0086] Specifically, the immersive video enhancement model proposed in the embodiments of the present application is composed of a feature extraction module, a frequency domain enhancement module, a boundary enhancement module, a spatio-temporal deformable convolution module, and a quality enhancement module. First, the feature extraction module performs a frequency domain transformation on the obtained compressed multi-viewpoint texture plus depth video (MVD) sequence to extract high-frequency features and low-frequency features respectively. Then, in the frequency domain enhancement module, according to the immersive video distortion characteristics, the low-frequency features are optimized based on viewpoint weighting to enhance the global structure perception, and the high-frequency weights corresponding to the high-frequency features are dynamically adjusted in combination with the saliency of the motion region, and finally a frequency domain enhanced image is obtained. In the boundary enhancement module, the boundary features are enhanced to strengthen the detail restoration of the foreground and background junction region, and the frequency domain enhanced image and the boundary features are dynamically aggregated to generate a fused image. Further, the fused image is input into the spatio-temporal deformable module to align the time information by using the inter-frame complementary relationship to generate an aligned fused image; finally, the aligned fused image is input into the quality enhancement module to predict the enhancement residual, and a reconstructed video with high-quality enhancement corresponding to the compressed multi-viewpoint texture plus depth video (MVD) sequence is output. The spatio-temporal deformable convolution module and the quality enhancement module are existing modules, and their specific calculation processes and principles will not be elaborated here.

[0087] In a specific embodiment, the total loss function used in the training process of the immersive video enhancement model is the weighted result of the reconstruction loss and the boundary loss, as shown in the following formula:

[0088] ;

[0089] Among them, represents the total loss function, and represent the weights of the reconstruction loss and the boundary loss respectively, and represent the reconstruction loss and the boundary loss respectively;

[0090] The calculation formula of the reconstruction loss is as follows:

[0091] ;

[0092] Among them, represents the total number of pixels, and represent the pixel value of the i-th pixel of the reconstructed video frame and the pixel value of the i-th pixel of the original image respectively;

[0093] The calculation formula of the boundary loss is as follows:

[0094] ;

[0095] Among them, Denote the gradient, which is obtained by the Sobel edge detection operator. and respectively reconstruct the gradient value of the i-th pixel of the video frame and the gradient value of the i-th pixel of the original image.

[0096] Specifically, to guide the immersive video enhancement model to optimize the video quality during training, a total loss function that comprehensively considers the reconstruction loss and the boundary loss is designed to recover more detailed information during training. Among them, the reconstruction loss function is designed by measuring the pixel-level difference between the reconstructed video frame and the original video frame, and the boundary loss function is designed by measuring the pixel-level difference at the edge region between the reconstructed video frame and the original video frame. A trained immersive video enhancement model is obtained based on the total loss function.

[0097] S2. Obtain the compressed multi-view texture plus depth video sequence to be reconstructed and input it into the trained immersive video enhancement model; the current video frame to be enhanced in each video sequence to be enhanced in the compressed multi-view texture plus depth video sequence to be reconstructed first undergoes frequency domain transformation and feature extraction through the feature extraction module, and the high-frequency feature and the low-frequency feature are respectively extracted; the high-frequency feature and the low-frequency feature are respectively given corresponding high-frequency weights and low-frequency weights through the frequency domain enhancement module, and are weighted and fused and inverse-transformed to obtain the frequency domain enhanced image; the frequency domain enhanced image and the current video frame to be enhanced are input into the boundary enhancement module to obtain the fused image; the fused image and the adjacent video frame of the current video frame to be enhanced in the video sequence to be enhanced are input into the spatio-temporal deformable convolution module, and the frame-to-frame complementary relationship is used to align the time information to obtain the aligned fused image, and the aligned fused image passes through the quality enhancement module to predict the enhancement residual and generate the corresponding reconstructed video.

[0098] Specifically, by deploying the trained immersive video enhancement model, the compressed multi-view texture plus depth video sequence to be reconstructed can be input into the trained immersive video enhancement model for video reconstruction.

[0099] In a specific embodiment, the current video frame to be enhanced in each video sequence to be enhanced in the compressed multi-view texture plus depth video sequence to be reconstructed first undergoes frequency domain transformation and feature extraction through the feature extraction module, and the high-frequency feature and the low-frequency feature are respectively extracted, which specifically includes:

[0100] For the video sequence to be enhanced , denote the current video frame to be enhanced, denote the frames before and after the current video frame to be enhanced, denote the adjacent video frame of the current video frame to be enhanced before the current video frame to be enhanced, Indicates the th adjacent video frame after the current video frame to be enhanced;

[0101] The current video frame to be enhanced is subjected to Fourier transform to obtain the corresponding frequency domain features, as shown in the following formula:

[0102] ;

[0103] where, represents the pixel value at the pixel coordinates in the current video frame to be enhanced, represents the frequency domain feature at the frequency domain coordinates corresponding to the current video frame to be enhanced, and respectively represent the width and height of the current video frame to be enhanced, represents the Fourier transform;

[0104] The low-frequency features and high-frequency features are separated from the frequency domain features through Gaussian filtering, as shown in the following formula:

[0105] ;

[0106] ;

[0107] where, represents the Gaussian filtering function, represents the low-frequency feature at the frequency domain coordinates corresponding to the current video frame to be enhanced, represents the high-frequency feature at the frequency domain coordinates corresponding to the current video frame to be enhanced.

[0108] Specifically, referring to Figure 2 , if there are M video sequences to be enhanced in the compressed multi-view texture plus depth video sequence to be reconstructed, each video sequence to be enhanced is composed of the current video frame to be enhanced and the previous and subsequent R video frames. In the feature extraction module, the current video frame to be enhanced is subjected to Fourier transform (FFT) to obtain the corresponding frequency domain features, and through Gaussian filtering, the low-frequency features representing the global structural information (such as smooth regions) and the high-frequency features representing the local detail information (such as edges and textures) are separated.

[0109] In a specific embodiment, the high-frequency features and low-frequency features are respectively given corresponding high-frequency weights and low-frequency weights through the frequency domain enhancement module, and weighted fusion and inverse transformation are performed to obtain the frequency domain enhanced image, specifically including:

[0110] For the low-frequency features, the low-frequency weights are calculated using the following formula:

[0111] ;

[0112] Among them, represents the low-frequency weight located at the frequency-domain coordinate of the low-frequency enhancement factor, is the low-frequency enhancement factor, , is the smoothing term, represents the Gaussian filtering function, represents the frequency-domain feature located at the frequency-domain coordinate corresponding to the current video frame to be enhanced, represents the viewpoint importance weight located at the frequency-domain coordinate , and its expression is as follows:

[0113] ;

[0114] Among them, represents the frequency-domain coordinate of the center frequency of the user's viewpoint area, represents the parameter of the user's viewpoint area, and its expression is: , represents the empirical factor, represents the maximum value among the frequency-domain features corresponding to the current video frame to be enhanced located at any frequency-domain coordinate;

[0115] For high-frequency features, the high-frequency weight is calculated using the following formula:

[0116] ;

[0117] Among them, is the high-frequency weight located at the frequency-domain coordinate , is the high-frequency retention factor, , represents the motion saliency weight located at the frequency-domain coordinate , and its expression is as follows:

[0118] ;

[0119] Among them, represents the pixel change amplitude between the current video frame to be enhanced and its previous video frame at the pixel coordinate , represents the pixel change amplitude between the current video frame to be enhanced and its previous video frame at any pixel coordinate , and max represents taking the maximum value;

[0120] The low-frequency features and high-frequency features are weighted using the low-frequency weight and high-frequency weight respectively to obtain the enhanced frequency-domain features, as shown in the following formula:

[0121] ;

[0122] Wherein, represents the enhanced frequency-domain feature corresponding to the current video frame to be enhanced;

[0123] The enhanced frequency-domain feature is returned to the spatial domain through inverse Fourier transform to obtain a frequency-domain enhanced image, as shown in the following formula:

[0124] ;

[0125] Wherein, represents the pixel value at the pixel coordinate in the frequency-domain enhanced image corresponding to the current video frame to be enhanced, and respectively represent the width and height of the current video frame to be enhanced.

[0126] Specifically, in the frequency-domain enhancement module, according to the human visual characteristics and immersive video characteristics, different importance low-frequency weights and high-frequency weights are assigned to the low-frequency and high-frequency features and dynamically weighted. For low-frequency features, to better adapt to the global characteristics of immersive video distortion, a viewpoint importance weight is introduced to assign higher weights to the user's viewpoint area of current interest. The viewpoint importance weight represents the weight distribution of low-frequency features in the user's viewpoint area and can be generated by the Gaussian model of the user's viewpoint area. For high-frequency features, since moving objects in immersive videos are usually accompanied by significant changes in high-frequency details, a motion saliency weight is introduced to dynamically adjust the weights in combination with motion information. After weighted fusion of low-frequency and high-frequency features through low-frequency and high-frequency weights, the corresponding enhanced frequency-domain features are obtained, and the enhanced frequency-domain features are returned to the spatial domain through inverse Fourier transform (IFFT) to obtain a frequency-domain enhanced image.

[0127] In a specific embodiment, the frequency-domain enhanced image and the current video frame to be enhanced are input into a boundary enhancement module to obtain a fused image, which specifically includes:

[0128] Use the Canny edge detection algorithm to extract the significant boundary regions in the frequency-domain enhanced image to obtain boundary features, as shown in the following formula:

[0129] ;

[0130] Wherein, represents the boundary feature at the pixel coordinate corresponding to the current video frame to be enhanced, Denote the Canny edge detection algorithm;

[0131] Calculate the pixel interpolation of the neighboring region in the current video frame to be enhanced, and estimate the artifact intensity, as shown in the following formula:

[0132] ;

[0133] Where, Denote the artifact intensity at the pixel coordinate in the current video frame to be enhanced, Denote the pixel value at the pixel coordinate in the current video frame to be enhanced, Denote the smooth estimate of the neighboring pixel at the pixel coordinate in the current video frame to be enhanced, and its expression is as follows:

[0134] ;

[0135] Where, Denote the neighborhood around the pixel coordinate , The pixel value at the neighboring pixel coordinate at the pixel coordinate in the current video frame to be enhanced, Denote taking the median of all pixel values in the neighborhood;

[0136] Weight the boundary feature and the artifact intensity, and calculate the boundary weight, as shown in the following formula:

[0137] ;

[0138] Where, Is the boundary enhancement factor, Denote the artifact influence factor;

[0139] Multiply the boundary weight by the frequency-domain enhanced image corresponding to the current video frame to be enhanced to obtain the fused image, as shown in the following formula:

[0140] ;

[0141] Where, Denote the pixel value at the pixel coordinate in the fused image corresponding to the current video frame to be enhanced.

[0142] Specifically, in the boundary enhancement module, the Canny edge detection algorithm is used to extract the significant boundary regions in the frequency-domain enhanced image to obtain boundary features. The boundary features are dynamically weighted to give higher attention to the boundary regions. The frequency-enhanced image and the boundary weights are fused to obtain a fused image.

[0143] In a specific embodiment, the fused image and the adjacent video frames of the current video frame to be enhanced in the video sequence to be enhanced are input into the spatio-temporal deformable convolution module. The inter-frame complementary relationship is utilized to align the time information, and the aligned fused image is obtained. The aligned fused image passes through the quality enhancement module to predict the enhancement residual and generate the corresponding reconstructed video, which specifically includes:

[0144] The image features pass through the spatio-temporal deformable convolution module to obtain the aligned fused image, as shown in the following formula:

[0145]

[0146] Among them, represents the size of the convolution kernel, represents the th channel of the convolution kernel, represents the aligned fused image corresponding to the current video frame to be enhanced, represents the fused image corresponding to the current video frame to be enhanced, represents any spatial position, represents the regular sampling offset corresponding to the th channel of the convolution kernel, represents the deformable offset corresponding to the time position being the Tth frame and the spatial position being , which is extracted by jointly extracting the offset fields of all time positions and spatial positions within the current video frame to be enhanced. The expression of the offset field is as follows:

[0147]

[0148] Among them, represents the U-net offset prediction network, represents the th adjacent video frame before the current video frame to be enhanced, represents the th adjacent video frame after the current video frame to be enhanced;

[0149] The aligned fused image is input into the quality enhancement module composed of multi-level residual dense channel attention blocks to predict the enhancement residual , and the enhancement residual is added to the current video frame to be enhanced element by element to generate the reconstructed video frame , as shown in the following formula:

[0150] ;

[0151] where represents element-wise addition, and all the reconstructed video frames constitute a reconstructed video sequence and synthesize a reconstructed video.

[0152] Specifically, to obtain temporal motion compensation, the video sequence is input into the spatio-temporal deformable convolutional module (STDC) to fuse adjacent frame information, aggregate the frequency-domain features according to the variability offset, and obtain the aligned fused image . The aligned fused image is input into the quality enhancement module composed of multi-level residual dense channel attention blocks to obtain the predicted enhanced residual, and then the enhanced residual is added to the current video frame to be enhanced element-wise to generate the reconstructed video frame. When all the video frames to be enhanced are generated into corresponding reconstructed video frames in the above manner, a reconstructed video sequence can be obtained, and finally the reconstructed video is synthesized.

[0153] The following is illustrated by specific experiments.

[0154] The experiment uses the MVD dataset provided by the MIV standard. The dataset contains a total of 16 multi-view test sequences with different resolutions, including 10 computer-generated contents (Computer generation, CG) and 6 natural contents (Natural Content, NC). The texture video format of the video sequence is YUV 4:2:0 10-bit, and the depth video format is YUV 4:2:0 16-bit. The projection formats include the equirectangular projection format (Equirectangular Projection, ERP) and the perspective projection format (Perspective Projection). Table 1 lists the detailed video parameters of the test sequences. The QP parameter combinations specified by the MPEG standard test conditions are also given, and the specific numerical settings are shown in Table 2.

[0155] Table 1 Experimental sequence information:

[0156]

[0157] Table 2 Video sequence quantization parameter setting table:

[0158]

[0159] The experimental dataset uses 16 multi-viewpoint MVD sequences improved by MPEG. The number of viewpoints in each sequence ranges from 9 to 25, and each viewpoint video contains 300 frames. In the experiment, 12 sequences are used as the training set. At the same time, two of the training sequences (P and T) are selected, and the last 100 frames of each sequence are removed as additional test sequences. All videos of the remaining 4 sequences (A, E, G, O) are used as the test set. The experiment compares the enhancement effects of MFQE 1.0, MFQE 2.0, STDF, RFDA, and the immersive video enhancement method based on frequency-domain boundary collaborative optimization proposed in this application under two groups of quantization parameter combinations of QP3 and QP4. The objective image quality evaluation indicators used are video peak signal-to-noise ratio (Peak Signal to Noise Ratio, PSNR) and immersive video peak signal-to-noise ratio (Immersive Video Peak Signal to Noise Ratio, IV-PSNR).

[0160] Table 3 shows when the compression parameter is QP3 PSNR (dB) and IV-PSNR (dB) comparison results:

[0161]

[0162] As shown in Table 3, the enhancement results when the compression parameter is QP3 are as follows. In all test scenarios, the IV-PSNR value increment of the immersive video enhancement method based on frequency-domain boundary collaborative optimization proposed in this application is significantly better than the existing methods. For low-quality scenarios (such as sequences O and P), the immersive video enhancement method based on frequency-domain boundary collaborative optimization proposed in this application realizes a large degree of detail restoration and artifact suppression by dynamically fusing visual saliency information, and the IV-PSNR increment is increased by about 0.07 to 0.10 dB. In other high-quality scenarios, the immersive video enhancement method based on frequency-domain boundary collaborative optimization proposed in this application can still ensure better enhancement effects than the existing methods by optimizing boundary details and spatio-temporal information alignment.

[0163] Further referring to Figure 3 , as an implementation of the methods shown in the above figures, this application provides an embodiment of an immersive video enhancement device based on frequency-domain boundary collaborative optimization. This device embodiment corresponds to Figure 1 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0164] This application embodiment provides an immersive video enhancement device based on frequency-domain boundary collaborative optimization, including:

[0165] The model construction module 1 is configured to construct and train an immersive video enhancement model to obtain a trained immersive video enhancement model. The immersive video enhancement model includes a feature extraction module, a frequency domain enhancement module, a boundary enhancement module, a spatio-temporal deformable convolution module, and a quality enhancement module that are connected in sequence;

[0166] The reconstruction module 2 is configured to obtain a compressed multi-view texture plus depth video sequence to be reconstructed and input it into the trained immersive video enhancement model. For each current video frame to be enhanced in the video sequences to be enhanced in the compressed multi-view texture plus depth video sequence to be reconstructed, it first undergoes frequency domain transformation and feature extraction through the feature extraction module to respectively extract high-frequency features and low-frequency features. The high-frequency features and low-frequency features are respectively given corresponding high-frequency weights and low-frequency weights through the frequency domain enhancement module, and then weighted fusion and inverse transformation are performed to obtain a frequency domain enhanced image. The frequency domain enhanced image and the current video frame to be enhanced are input into the boundary enhancement module to obtain a fused image. The fused image and the adjacent video frames of the current video frame to be enhanced in the video sequence to be enhanced are input into the spatio-temporal deformable convolution module, and the frame-interframe complementary relationship is used to align the time information to obtain an aligned fused image. The aligned fused image passes through the quality enhancement module to predict an enhancement residual and generate a corresponding reconstructed video.

[0167] Figure 4 It is a schematic hardware structure diagram of the electronic device provided by the embodiment of the present invention. As Figure 4 shown, the electronic device of this embodiment includes: a processor 401 and a memory 402; wherein the memory 402 is used to store computer execution instructions; the processor 401 is used to execute the computer execution instructions stored in the memory to implement the various steps executed by the electronic device in the above embodiment. Specifically, reference can be made to the relevant descriptions in the foregoing method embodiments.

[0168] Optionally, the memory 402 can be either independent or integrated with the processor 401.

[0169] When the memory 402 is independently set, the electronic device further includes a bus 403 for connecting the memory 402 and the processor 401.

[0170] The embodiment of the present invention also provides a computer storage medium, in which computer execution instructions are stored. When the processor 401 executes the computer execution instructions, the above method is implemented.

[0171] The embodiment of the present invention also provides a computer program product, including a computer program. When the computer program is executed by the processor 401, the above method is implemented.

[0172] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be indirect couplings or communication connections through some interfaces, devices or modules, and can be in electrical, mechanical or other forms.

[0173] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.

[0174] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in one unit. The units formed by the above modules can be implemented in the form of hardware or in the form of a combination of hardware and software functional units.

[0175] The integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor 401 to execute some steps of the methods in various embodiments of the present application.

[0176] It should be understood that the above processor 401 can be a central processing unit (Central Processing Unit, abbreviated as CPU), and can also be other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as DSP), application specific integrated circuits (Application Specific Integrated Circuit, abbreviated as ASIC), etc. The general-purpose processor can be a microprocessor or the processor 401 can also be any conventional processor 401, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by the hardware processor 401, or executed by a combination of the hardware and software modules in the processor 401.

[0177] The memory 402 may include high-speed RAM memory and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.

[0178] The bus 403 can be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus 403 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the bus 403 in the accompanying drawings of this application is not limited to only one bus 403 or one type of bus 403.

[0179] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0180] An exemplary storage medium is coupled to the processor 401, enabling the processor 401 to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor 401. The processor 401 and the storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor 401 and the storage medium can also exist as discrete components in an electronic device or a master control device.

[0181] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: ROM, RAM, magnetic disk, or optical disc and other media that can store program codes.

[0182] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An immersive video enhancement method based on frequency domain boundary collaborative optimization, characterized in that: The following steps are involved: An immersive video enhancement model is constructed and trained to obtain a trained immersive video enhancement model, wherein the immersive video enhancement model includes a feature extraction module, a frequency domain enhancement module, a boundary enhancement module, a spatiotemporal deformable convolution module, and a quality enhancement module connected in sequence; A compressed multi-view texture plus depth video sequence to be reconstructed is obtained and input into the trained immersive video enhancement model; a current video frame to be enhanced in each video sequence to be enhanced in the compressed multi-view texture plus depth video sequence to be reconstructed is first subjected to frequency domain transformation and feature extraction by the feature extraction module to respectively extract high-frequency features and low-frequency features; The high-frequency features and low-frequency features are respectively assigned corresponding high-frequency weights and low-frequency weights by the frequency domain enhancement module, and weighted fusion and inverse transformation are performed to obtain a frequency domain enhanced image, specifically including: For the low-frequency features, the low-frequency weight is calculated using the following formula: Among them, W low (u, v) represents the low-frequency weight at the frequency domain coordinate (u, v), α is the low-frequency enhancement factor, α>1, ∈ is the smoothing term, H(u, v) represents the Gaussian filter function, F t (u, v) represents the frequency domain feature at the frequency domain coordinate (u, v) corresponding to the current video frame to be enhanced, W viewport (u,v) represents the importance weight of the viewpoint located at the frequency domain coordinate (u,v), and its expression is as follows: Among them, u c, v c represents the frequency domain coordinate of the center frequency of the user's viewpoint area, σ represents the parameter of the user's viewpoint area, and its expression is: σ = λ·f max , λ represents the empirical factor, f max Indicates the maximum value of the frequency domain feature located at any frequency domain coordinate corresponding to the current video frame to be enhanced; For the high-frequency features, the high-frequency weight is calculated using the following formula: Among them, W high (u, v) is the high-frequency weight at the frequency domain coordinate (u, v), β is the high-frequency retention factor, β < 1, W motion (u,v) represents the motion saliency weight at the frequency domain coordinate (u,v), which is expressed as follows: Among them, ΔI t (x, y) represents the pixel change amplitude at pixel coordinates (x, y) between the current video frame to be enhanced and the previous video frame, ΔI t (x', y') represents the pixel change amplitude at any pixel coordinate (x', y') between the current video frame to be enhanced and the previous video frame, and max represents the maximum value; The low-frequency weight and high-frequency weight are used to weight the low-frequency features and high-frequency features respectively to obtain the enhanced frequency domain features, as shown in the following formula: F t '(u,v)=W low (u,v)·F low (u,v)+W high (u,v)·F high (u,v); Among them, F t '(u,v) represents the enhanced frequency domain features corresponding to the current video frame to be enhanced; The enhanced frequency domain features are returned to the spatial domain through inverse Fourier transform to obtain a frequency domain enhanced image, as shown in the following formula: Among them, I weighted(t) (x, y) represents the pixel value at the pixel coordinate (x, y) in the frequency domain enhanced image corresponding to the current video frame to be enhanced, W and H represent the width and height of the current video frame to be enhanced respectively; the frequency domain enhanced image and the current video frame to be enhanced are input into the boundary enhancement module to obtain a fused image; the fused image and the adjacent video frames of the current video frame to be enhanced in the video sequence to be enhanced are input into the spatiotemporal deformable convolution module, and the time information is aligned using the complementary relationship between frames to obtain an aligned fused image; the aligned fused image passes through the quality enhancement module to predict the enhanced residual and generate the corresponding reconstructed video.

2. The immersive video enhancement method based on frequency domain boundary collaborative optimization according to claim 1, characterized in that: The current video frame to be enhanced in each video sequence to be enhanced in the compressed multi-view texture plus depth video sequence to be reconstructed is first subjected to frequency domain transformation and feature extraction by the feature extraction module to respectively extract high-frequency features and low-frequency features, specifically including: For the video sequence to be enhanced {I t-R ,…,I t ,…,I t+R }, I t represents the current video frame to be enhanced, R represents the Rth frame before and after the current video frame to be enhanced, I t-R represents the Rth adjacent video frame before the current video frame to be enhanced, I t+R represents the Rth adjacent video frame after the current video frame to be enhanced; The current video frame to be enhanced is subjected to Fourier transform to obtain the corresponding frequency domain features, as shown in the following formula: Among them, I t (x, y) represents the pixel value at the pixel coordinate (x, y) in the current video frame to be enhanced, F t (u,v) represents the frequency domain feature at the frequency domain coordinate (u,v) corresponding to the current video frame to be enhanced, W and H represent the width and height of the current video frame to be enhanced, respectively. represents Fourier transform; The low-frequency features and high-frequency features are separated from the frequency domain features by Gaussian filtering, as shown in the following formula: F low (u,v)=H(u,v)·F t (u,v); F high (u,v)=(1-H(u,v))·F t (u,v); Among them, H(u,v) represents the Gaussian filter function, F low (u, v) represents the low-frequency feature at the frequency domain coordinate (u, v) corresponding to the current video frame to be enhanced, F high (u, v) represents the high-frequency feature located at the frequency domain coordinate (u, v) corresponding to the current video frame to be enhanced.

3. The immersive video enhancement method based on frequency domain boundary collaborative optimization according to claim 1, characterized in that: The frequency domain enhanced image and the current video frame to be enhanced are input into the boundary enhancement module to obtain a fused image, which specifically includes: The Canny edge detection algorithm is used to extract the significant boundary area in the frequency domain enhanced image to obtain the boundary feature, as shown in the following formula: B t (x,y)=EdgeDetector(I weighted(t) (x,y)); Among them, B t (x, y) represents the boundary feature of the pixel coordinate (x, y) corresponding to the current video frame to be enhanced, and EdgeDetector represents the Canny edge detection algorithm; Calculate the pixel interpolation of the neighboring area in the current video frame to be enhanced, and estimate the artifact intensity, as shown in the following formula: Where A(x,y) represents the artifact intensity at the pixel coordinate (x,y) in the current video frame to be enhanced, and I t (x, y) represents the pixel value at the pixel coordinate (x, y) in the current video frame to be enhanced. represents the smoothed estimate of the neighboring pixels at the pixel coordinates (x, y) in the current video frame to be enhanced, and its expression is as follows: Among them, N k' represents the k'×k' neighborhood around the pixel coordinate (x, y), I(x+i, y+j) is the pixel value of the neighboring pixel coordinate (x+i, y+j) located at the pixel coordinate (x, y) in the current video frame to be enhanced, and Median represents the median of all pixel values ​​in the neighborhood; The boundary feature and the artifact strength are weighted to calculate the boundary weight, as shown in the following formula: W edge(t) (x,y)=γ·B t (x,y)+δ·A t (x,y); Among them, γ is the boundary enhancement factor, δ represents the artifact impact factor; The boundary weight is multiplied by the frequency domain enhanced image corresponding to the current video frame to be enhanced to obtain a fused image, as shown in the following formula: I fused(t) (x,y)=W edge (x,y)·I weighted(t) (x,y); Among them, I fused(t) (x, y) represents the pixel value at the pixel coordinate (x, y) in the fused image corresponding to the current video frame to be enhanced.

4. The immersive video enhancement method based on frequency domain boundary collaborative optimization according to claim 1, characterized in that: The fused image and adjacent video frames of the current video frame to be enhanced in the video sequence to be enhanced are input into the spatiotemporal deformable convolution module, and the time information is aligned by using the complementary relationship between frames to obtain an aligned fused image. The aligned fused image is passed through the quality enhancement module to predict the enhanced residual and generate the corresponding reconstructed video, specifically including: The image features are passed through the spatiotemporal deformable convolution module to obtain an aligned fused image, as shown in the following formula: Among them, K represents the size of the convolution kernel, W k represents the kth channel of the convolution kernel, I ft represents the aligned fused image corresponding to the current video frame to be enhanced, I fused(t) represents the fused image corresponding to the current video frame to be enhanced, p represents any spatial position, and p k represents the conventional sampling offset corresponding to the kth channel of the convolution kernel, δ (T,p) It represents the deformable offset corresponding to the time position of the Tth frame and the spatial position of p, which is extracted by combining the offset fields of all time positions and spatial positions in the current video frame to be enhanced. The expression of the offset field Δ is as follows: in, represents the U-net offset prediction network, I t-R represents the Rth adjacent video frame before the current video frame to be enhanced, I t+R represents the Rth adjacent video frame after the current video frame to be enhanced; The aligned fused image is input into a quality enhancement module composed of multi-level residual dense channel attention blocks to predict the enhanced residual RC t , will enhance the residual RC t Add the current video frame to be enhanced element by element to generate a reconstructed video frame As shown below: in, It represents element-by-element addition, and all reconstructed video frames constitute a reconstructed video sequence and synthesize the reconstructed video.

5. The immersive video enhancement method based on frequency domain boundary collaborative optimization according to claim 1, characterized in that: The total loss function used in the training process of the immersive video enhancement model is the weighted result of the reconstruction loss and the boundary loss, as shown in the following formula: L total =λ content L content +λ boundary L boundary ; Among them, L total represents the total loss function, λ content and λ boundary Represent the weight of reconstruction loss and the weight of boundary loss respectively, L content and L boundary denote reconstruction loss and boundary loss respectively; The calculation formula of the reconstruction loss is as follows: Where N represents the total number of pixels, and denote the pixel value of the i-th pixel of the reconstructed video frame and the pixel value of the i-th pixel of the original image respectively; The calculation formula of the boundary loss is as follows: in, Represents the gradient, obtained by the Sobel edge detection operator, and The gradient value of the i-th pixel of the video frame and the gradient value of the i-th pixel of the original image are reconstructed respectively.

6. An immersive video enhancement device based on frequency domain boundary collaborative optimization, characterized in that: include: A model building module is configured to build and train an immersive video enhancement model to obtain a trained immersive video enhancement model, wherein the immersive video enhancement model includes a feature extraction module, a frequency domain enhancement module, a boundary enhancement module, a spatiotemporal deformable convolution module, and a quality enhancement module connected in sequence; The reconstruction module is configured to obtain a compressed multi-view texture plus depth video sequence to be reconstructed and input it into the trained immersive video enhancement model; the current video frame to be enhanced in each video sequence to be enhanced in the compressed multi-view texture plus depth video sequence to be reconstructed is first subjected to frequency domain transformation and feature extraction by the feature extraction module to respectively extract high-frequency features and low-frequency features; The high-frequency features and low-frequency features are respectively assigned corresponding high-frequency weights and low-frequency weights by the frequency domain enhancement module, and weighted fusion and inverse transformation are performed to obtain a frequency domain enhanced image, specifically including: For the low-frequency features, the low-frequency weight is calculated using the following formula: Among them, W low (u, v) represents the low-frequency weight at the frequency domain coordinate (u, v), α is the low-frequency enhancement factor, α>1, ∈ is the smoothing term, H(u, v) represents the Gaussian filter function, F t (u, v) represents the frequency domain feature at the frequency domain coordinate (u, v) corresponding to the current video frame to be enhanced, W viewport (u,v) represents the importance weight of the viewpoint located at the frequency domain coordinate (u,v), and its expression is as follows: Among them, u c ,v c represents the frequency domain coordinate of the center frequency of the user's viewpoint area, σ represents the parameter of the user's viewpoint area, and its expression is: σ = λ·f max , λ represents the empirical factor, f max Indicates the maximum value of the frequency domain feature located at any frequency domain coordinate corresponding to the current video frame to be enhanced; For the high-frequency features, the high-frequency weight is calculated using the following formula: Among them, W high (u, v) is the high-frequency weight at the frequency domain coordinate (u, v), β is the high-frequency retention factor, β < 1, W motion (u,v) represents the motion saliency weight at the frequency domain coordinate (u,v), which is expressed as follows: Among them, ΔI t (x, y) represents the pixel change amplitude at pixel coordinates (x, y) between the current video frame to be enhanced and the previous video frame, ΔI t (x', y') represents the pixel change amplitude at any pixel coordinate (x', y') between the current video frame to be enhanced and the previous video frame, and max represents the maximum value; The low-frequency weight and high-frequency weight are used to weight the low-frequency features and high-frequency features respectively to obtain the enhanced frequency domain features, as shown in the following formula: F t '(u,v)=W low (u,v)·F low (u,v)+W high (u,v)·F high (u,v); Among them, F t '(u,v) represents the enhanced frequency domain features corresponding to the current video frame to be enhanced; The enhanced frequency domain features are returned to the spatial domain through inverse Fourier transform to obtain a frequency domain enhanced image, as shown in the following formula: Among them, I weighted(t) (x, y) represents the pixel value at the pixel coordinate (x, y) in the frequency domain enhanced image corresponding to the current video frame to be enhanced, W and H represent the width and height of the current video frame to be enhanced respectively; the frequency domain enhanced image and the current video frame to be enhanced are input into the boundary enhancement module to obtain a fused image; the fused image and the adjacent video frames of the current video frame to be enhanced in the video sequence to be enhanced are input into the spatiotemporal deformable convolution module, and the time information is aligned using the complementary relationship between frames to obtain an aligned fused image; the aligned fused image passes through the quality enhancement module to predict the enhanced residual and generate the corresponding reconstructed video.

7. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Constant code rate compressed video quality enhancement method based on double-domain learning

    CN115131254A

  • Image quality enhancement method and system based on cloud game

    CN119367760A