A compressed HDR video enhancement method and system based on deep learning

By building a compressed HDR video dataset and deep learning model with rich scenes, image-feature conversion and multiple rounds of enhancement processing, the shortcomings of compressed HDR video enhancement in the existing technology are solved, and bit rate saving and video quality improvement are achieved.

CN116193121BActive Publication Date: 2025-05-13UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310210249.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-07
Publication Date
2025-05-13
Estimated Expiration
2043-03-07

AI Technical Summary

Technical Problem

The prior art lacks deep learning models and data sets for compressed HDR video enhancement, and cannot effectively improve the quality of compressed HDR video.

Method used

By building a compressed HDR video dataset with rich scene information, using deep learning models for image-feature conversion, downsampling, video enhancement module recovery, upsampling and feature fusion, the quality of compressed HDR video is gradually improved.

Benefits of technology

It achieves efficient enhancement of compressed HDR videos, saves bit rate, improves video quality, and is suitable for improving user viewing experience and expanding research directions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116193121B_ABST
    Figure CN116193121B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for enhancing compressed HDR video based on deep learning, and the enhancing method comprises the following steps: S1, constructing a compressed HDR video data set; S2, performing image-feature conversion; S3, performing downsampling; S4, using a video enhancement module to perform a first recovery of the downsampled HDR video features; S5, performing a second downsampling, using a video enhancement module to perform a second inter-scale recovery, and then fusing the features with the features before the second enhancement; S6, performing an upsampling operation, and then using a video enhancement module to perform enhancement, and fusing the features with the features after the first enhancement; S7, performing an upsampling operation, and fusing the features obtained by upsampling with the features obtained by S2; S8, obtaining an enhanced HDR video through a feature-image conversion module. The method of the present invention can save bit rate and improve the quality of compressed HDR video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of video coding and decoding, and in particular relates to a compressed HDR video enhancement method and a compressed HDR video enhancement system based on deep learning. Background Art

[0002] A complex system has a wide range of brightness, which brings colorful visual details. However, in the past few years, 8-bit or lower bit-depth displays still dominated the market, which made many scholars' research still stay on compressed standard dynamic range (SDR) video enhancement. With the development of science and technology and the improvement of people's living standards, displays that support 10-bit or even higher bit-depths have become popular, and it is urgent to study the methods of enhancing compressed high dynamic range (HDR) videos. In the field of film photography, high dynamic range videos save scenes similar to those of human eyes, which greatly improves people's viewing experience. In high-level tasks of computer vision, enhanced HDR videos can help improve the accuracy of target detection tasks, classification tasks, etc. due to their richer semantic features. In the field of medicine, enhanced HDR videos with higher dynamic range can help telemedicine diagnosis more clearly and reliably, and the low latency brought by bit rate savings can help doctors analyze and diagnose patients' conditions more clearly and promptly. In general, HDR videos have inherent advantages over SDR videos. Research on HDR videos, especially research on compressed HDR video enhancement, is an urgent need in today's era.

[0003] At present, the research at home and abroad is mainly focused on the enhancement of compressed SDR videos. There are many mature models and a large number of data sets in this direction. However, as far as we know, there are no special deep learning models and supporting data sets designed for the research of compressed HDR video enhancement at home and abroad.

[0004] Patent application number 202110319011.6 discloses a method for blindly enhancing compressed video quality based on QP estimation. This method is for SDR video enhancement and blindly enhances SDR video by estimating the quality factor (QP). However, this method is only for SDR video and does not take HDR video into consideration. Therefore, it is not suitable for compressed HDR video enhancement.

[0005] Patent application number 202010026179.3 discloses a deep learning-based dynamic scene HDR reconstruction method. This method uses three standard dynamic range images with different exposures as input, aligns the images in the exposure stack using the LK optical flow method, and then reconstructs the HDR image through a residual attention module and a UNet-like network. This method only considers the mapping from SDR images to HDR images, which does not match the problem of compressed HDR video enhancement that we currently need to solve. Summary of the invention

[0006] The purpose of the present invention is to overcome the shortcomings of the prior art, provide a compressed HDR video enhancement method based on deep learning that can save bit rate and improve the quality of compressed HDR video, and provide a compressed HDR video enhancement system based on deep learning.

[0007] The objective of the present invention is achieved through the following technical solution: A compressed HDR video enhancement method based on deep learning, comprising the following steps:

[0008] S1. Build a compressed HDR video dataset with rich scene information;

[0009] S2, obtaining a compressed HDR video feature map by performing image-feature conversion on the video frames in the compressed HDR video dataset;

[0010] S3, downsampling the compressed HDR video feature map using a downsampling operation;

[0011] S4, using the video enhancement module to perform a first restoration on the downsampled HDR video features to obtain the initially restored HDR video features;

[0012] S5, after performing a second downsampling on the initially restored HDR video features, a video enhancement module is used to perform a second inter-scale restoration to obtain a second reconstructed HDR video feature, and then the feature is fused with the feature before the second enhancement;

[0013] S6, decoding the HDR video features after fusion in S5, firstly using the upsampling module to upsample the fused HDR video features, then using the video enhancement module to enhance them, and fusing the enhanced features with the first enhanced features;

[0014] S7, upsampling the HDR video features fused by S6, and fusing the upsampled features with the features obtained by S2;

[0015] S8. The HDR video features fused in S7 are converted into enhanced HDR video through a feature-image conversion module.

[0016] The specific implementation method of step S1 is: screen out videos including indoor, outdoor, light, shadow, portrait, and animal scene categories, extract one frame from these videos every 60 frames, and then screen again to discard pictures with similar scenes, to obtain 2023 frames of HDR video data with rich scenes, with a resolution of 3840*2160;

[0017] The HDR video data is compressed by HM16.9 at QP of 22, 27, 32, and 37 to obtain four compressed data sets. Each data set is then cropped every 400 pixels to a size of 512*512, resulting in a total of 101,150 video frames, which constitute the compressed HDR video data set.

[0018] The video enhancement module is composed of four secondary residual blocks connected together, which can be formulated as follows:

[0019]

[0020]

[0021] E(H F )=Cat(RC 1 (H F ),RC 2 (H F ), RC 3 (H F ),RC 4 (H F ))

[0022] Among them, Conv 3x3 represents a 3x3 convolutional block, ReLU() represents a ReLU activation function, and D(·) represents a residual block; 0 represents a convolution operation; H F represents the compressed HDR video features to be processed, RC(·) represents the secondary residual connection block, and the Cat(·) operation is used to connect the inputs of the secondary residual connection blocks together; the output of each secondary residual block will be used as the input of another secondary residual block, and then the outputs of the four secondary residuals will be concatenated together.

[0023] Another object of the present invention is to provide a compressed HDR video enhancement system based on deep learning, comprising three video enhancement blocks, an image-feature conversion module, two upsampling modules, two downsampling modules, three skip-connected structures, and a feature-image conversion module; the three video enhancement modules are respectively recorded as a first video enhancement module, a second video enhancement module, and a third video enhancement module; the two upsampling modules are respectively recorded as a first upsampling module and a second upsampling module; the two downsampling modules are respectively recorded as a first downsampling module and a second downsampling module; the three skip-connected structures are respectively recorded as a first skip-connected structure, a second skip-connected structure, and a third skip-connected structure;

[0024] The image-feature conversion module, the first down-sampling module, the first video enhancement module, the second down-sampling module, the second video enhancement module, the first up-sampling module, the third video enhancement module, and the second up-sampling module adopt a series structure; the input of the image-feature conversion module is a compressed HDR video data set;

[0025] The first skip connection structure passes the output of the image-feature conversion module to the output of the second upsampling module, and the output of the image-feature conversion module and the output of the second upsampling module are added as the input of the feature-image conversion module;

[0026] The second jump connection structure transmits the output of the first video enhancement module to the output of the first up-sampling module, and the output of the first video enhancement module and the output of the first up-sampling module are added to serve as the input of the third video enhancement module;

[0027] The third skip connection structure passes the output of the second down-sampling module to the output of the second video enhancement module, and the output of the second down-sampling module and the output of the second video enhancement module serve as the input of the first up-sampling module.

[0028] The beneficial effects of the present invention are as follows: the present invention provides a compressed HDR video enhancement method based on deep learning, which can save bit rate and improve the quality of compressed HDR video. It is the first to propose the first compressed HDR video enhancement method based on deep learning at home and abroad, breaking through the previous research that only focused on compressed SDR video enhancement and ignored compressed HDR video enhancement, and provides a good foundation for further improving the user viewing experience and expanding the research direction. At the same time, it has a certain cutting-edge technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a flowchart of the compressed HDR video enhancement method based on deep learning of the present invention;

[0030] Figure 2 It is a schematic diagram of the structure of the video enhancement module of the present invention. DETAILED DESCRIPTION

[0031] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.

[0032] like Figure 1 As shown, a compressed HDR video enhancement method based on deep learning of the present invention comprises the following steps:

[0033] S1. Construct a compressed HDR video dataset with rich scene information. The specific implementation method is as follows: by analyzing the properties of HDR videos, such as scenes, color gamut and bit depth, 20 long video sets on the Internet that meet the requirements are selected, totaling 20,000 frames. Then, these videos are screened to select videos including indoor, outdoor, lighting, shadow, portrait, and animal scene categories. After extracting one frame every 60 frames, these videos are screened again to discard pictures with similar scenes. 2023 frames of HDR video data with rich scenes are obtained, with a resolution of 3840*2160, covering commonly used high dynamic range video standards such as HDR10 and HLG.

[0034] The HDR video data is compressed by HM16.9 at QP of 22, 27, 32, and 37 to obtain four compressed data sets. Each data set is then cropped every 400 pixels to a size of 512*512, resulting in a total of 101,150 video frames, which constitute the compressed HDR video data set.

[0035] S2, obtaining a compressed HDR video feature map by performing image-feature conversion on the video frames in the compressed HDR video dataset;

[0036] S3, downsampling the compressed HDR video feature map using a downsampling operation;

[0037] S4, use the video enhancement module to perform the first recovery of the downsampled HDR video features to obtain the initially recovered HDR video features; the video enhancement model is the basic module for compressed HDR video enhancement, and its main functions are two: first, to ensure that the compressed HDR video information is not lost, and second, to enhance the compressed HDR video. Based on these two points, the present invention constructs the module based on 3x3 convolution and ReLU activation function, such as Figure 2 As shown in Figure 1, the video enhancement module consists of four secondary residual blocks connected together, which are formulated as follows:

[0038]

[0039]

[0040] E(H F )=Cat(RC 1 (H F ), RC 2 (HF ), RC 3 (H F ), RC 4 (H F ))

[0041] Among them, Conv 3x3 represents a 3x3 convolutional block, ReLU() represents a ReLU activation function, and D(·) represents a residual block; 0 represents a convolution operation; H F Represents the compressed HDR video features to be processed, RC(·) represents the secondary residual connection block, and the Cat(·) operation is used to link the inputs of the secondary residual connection blocks together; the output of each secondary residual block will be used as the input of another secondary residual block, and then the outputs of the four secondary residuals will be spliced ​​together. The advantage of this is that the input compressed HDR video features are as lossless as possible, and these HDR features can be enhanced with the help of convolution blocks.

[0042] S5, after performing a second downsampling on the initially restored HDR video features, a video enhancement module is used to perform a second inter-scale restoration to obtain a second reconstructed HDR video feature, and then the feature is fused with the feature before the second enhancement;

[0043] S6, decoding the HDR video features after fusion in S5, firstly using the upsampling module to upsample the fused HDR video features, then using the video enhancement module to enhance them, and fusing the enhanced features with the first enhanced features;

[0044] S7, upsampling the HDR video features fused by S6, and fusing the upsampled features with the features obtained by S2;

[0045] S8. The HDR video features fused in S7 are converted into enhanced HDR video through a feature-image conversion module.

[0046] There are three problems with the HDR video enhancement block constructed by the present invention when it is directly used for compressed HDR video reconstruction: 1) only one video enhancement block is not accurate enough for HDR video enhancement tasks; 2) multi-scale information is not taken into account; 3) the amount of calculation for directly accumulating multiple video enhancement blocks is large. Based on this, the present invention adopts a UNet-like structure to solve this problem. First, the compressed HDR video frame is converted into an HDR video feature by an image feature conversion module at the head of the model, where the image-feature conversion module is implemented using a 3x3 convolution. Then, a 3x3 convolution is used for downsampling operation, followed by a video enhancement module for preliminary enhancement, and then the obtained HDR video feature is subjected to a second downsampling operation and then enhanced again by the video enhancement module, followed by an upsampling operation consisting of a 3x3 convolution and pixelShuffle for upsampling the reconstructed HDR video feature, and a video enhancement module is used again for the upsampled HDR video feature, and then an upsampling operation is performed again. So far, the final enhanced HDR video feature map is obtained. Finally, the enhanced video feature map is converted into a video frame by a feature-image conversion module constructed by a 3x3 convolution. The structure of the compressed HDR video enhancement system based on deep learning is as follows Figure 1 As shown, it includes three video enhancement blocks, an image-feature conversion module, two upsampling modules, two downsampling modules, three skip structures, and a feature-image conversion module; the three video enhancement modules are respectively recorded as the first video enhancement module, the second video enhancement module, and the third video enhancement module; the two upsampling modules are respectively recorded as the first upsampling module and the second upsampling module; the two downsampling modules are respectively recorded as the first downsampling module and the second downsampling module; the three skip structures are respectively recorded as the first skip structure, the second skip structure, and the third skip structure; in the figure, from left to right, the first 160*160*64 represents the output size of the image-feature conversion module, the second 160*160*64 represents the input size of the feature-image conversion module, the first 80*80*64 represents the output size of the first video enhancement block, the second 80*80*64 represents the output size of the third video enhancement block, and the 40*40*64 represents the output size of the third video enhancement block.

[0047] The image-feature conversion module, the first down-sampling module, the first video enhancement module, the second down-sampling module, the second video enhancement module, the first up-sampling module, the third video enhancement module, and the second up-sampling module adopt a series structure; the input of the image-feature conversion module is a compressed HDR video data set;

[0048] The first skip connection structure passes the output of the image-feature conversion module to the output of the second upsampling module, and the output of the image-feature conversion module and the output of the second upsampling module are added as the input of the feature-image conversion module;

[0049] The second jump connection structure transmits the output of the first video enhancement module to the output of the first up-sampling module, and the output of the first video enhancement module and the output of the first up-sampling module are added to serve as the input of the third video enhancement module;

[0050] The third skip connection structure passes the output of the second down-sampling module to the output of the second video enhancement module, and the output of the second down-sampling module and the output of the second video enhancement module serve as the input of the first up-sampling module.

[0051] The compressed HDR video dataset constructed in step S1 is used for training. At the same time, in order to verify the effectiveness of our model, we use 4 standard HDR video test sequences with a total of 1319 frames, which contain different scene information. In the test, we set QP to 22, 27, 32, and 37 respectively.

[0052] The present invention uses ADAM optimizer and sets the learning rate to 2e -4 And it is reduced by half every 100,000 iterations. A total of 500,000 iterations. All models are built using the Pytorch framework and trained on NVIDIA GeForce RTX 2080SUPER. The total training time is 17 hours.

[0053] Table 1 shows the bit rate savings after adopting the present invention, using the test set compressed by VTM12.1 as a benchmark for comparison. From the table, it can be seen that the present invention can effectively save bit rate, especially in the Y component, achieving a bit rate savings of 1.74%.

[0054] Table 1

[0055] name Y U V Hurdles -1.71% -0.33% -0.86% BalloonFestival -1.60% -0.52% -0.52% ShowGirl2TeaserClip -2.49% -1.32% -1.28% Cosmos -1.17% -1.48% -2.03% average -1.74% -0.91% -1.17%

[0056] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.

Claims

1. A compressed HDR video enhancement method based on deep learning, characterized in that: The following steps are involved: S1. Build a compressed HDR video dataset with rich scene information; S2, obtaining a compressed HDR video feature map by performing image-feature conversion on the video frames in the compressed HDR video dataset; S3, downsampling the compressed HDR video feature map using a downsampling operation; S4, using the video enhancement module to perform a first restoration on the downsampled HDR video features to obtain the initially restored HDR video features; S5, after performing a second downsampling on the initially restored HDR video features, a video enhancement module is used to perform a second inter-scale restoration to obtain a second reconstructed HDR video feature, and then the feature is fused with the feature before the second enhancement; S6, decoding the HDR video features fused in S5, firstly using an upsampling module to upsample the fused HDR video features, then fusing the upsampled features with the features after the first enhancement, and then using a video enhancement module to enhance them; S7, upsampling the HDR video features fused by S6, and fusing the upsampled features with the features obtained by S2; S8. The HDR video features fused in S7 are converted into enhanced HDR video through a feature-image conversion module.

2. The method for compressed HDR video enhancement based on deep learning according to claim 1, characterized in that: The specific implementation method of step S1 is: screen out videos including indoor, outdoor, light, shadow, portrait, and animal scene categories, extract one frame from these videos every 60 frames, and then screen again to discard pictures with similar scenes, to obtain 2023 frames of HDR video data with rich scenes, with a resolution of 3840*2160; The HDR video data is compressed by HM16.9 at QP of 22, 27, 32, and 37 to obtain four compressed data sets. Each data set is then cropped every 400 pixels to a size of 512*512, resulting in a total of 101,150 video frames, which constitute the compressed HDR video data set.

3. The method for compressed HDR video enhancement based on deep learning according to claim 1, characterized in that: The video enhancement module is composed of four sequentially connected secondary residual blocks, which are formulated as follows: E(H F )=Cat(RC 1 (H F ),RC 2 (H F ),RC 3 (H F ),RC 4 (H F ) Among them, Conv 3x3 represents a 3x3 convolutional block, ReLU() represents the ReLU activation function, and D(·) represents the residual block; represents the convolution operation; H F represents the compressed HDR video features to be processed, RC(·) represents the secondary residual connection block, and the Cat(·) operation is used to connect the inputs of the secondary residual connection blocks together; the output of each secondary residual block will be used as the input of another secondary residual block, and then the outputs of the four secondary residuals will be concatenated together.

4. A compressed HDR video enhancement system based on deep learning, characterized in that: It includes three video enhancement modules, an image-feature conversion module, two upsampling modules, two downsampling modules, three skip-connection structures, and a feature-image conversion module; the three video enhancement modules are respectively recorded as a first video enhancement module, a second video enhancement module, and a third video enhancement module; the two upsampling modules are respectively recorded as a first upsampling module and a second upsampling module; the two downsampling modules are respectively recorded as a first downsampling module and a second downsampling module; the three skip-connection structures are respectively recorded as a first skip-connection structure, a second skip-connection structure, and a third skip-connection structure; The image-feature conversion module, the first down-sampling module, the first video enhancement module, the second down-sampling module, the second video enhancement module, the first up-sampling module, the third video enhancement module, and the second up-sampling module adopt a series structure; the input of the image-feature conversion module is a compressed HDR video data set; The first skip connection structure passes the output of the image-feature conversion module to the output of the second upsampling module, and the output of the image-feature conversion module and the output of the second upsampling module are added as the input of the feature-image conversion module; The second jump connection structure transmits the output of the first video enhancement module to the output of the first up-sampling module, and the output of the first video enhancement module and the output of the first up-sampling module are added to serve as the input of the third video enhancement module; The third skip connection structure passes the output of the second down-sampling module to the output of the second video enhancement module, and the output of the second down-sampling module and the output of the second video enhancement module serve as the input of the first up-sampling module.

Citation Information

Patent Citations

  • Dynamic scene HDR reconstruction method based on deep learning

    CN111242883A

  • Compressed video quality blind enhancement method based on QP estimation

    CN115134598A

  • Quality enhancement method for fixed-code-rate compressed video

    CN114125460A

  • Ball machine monitoring anomaly detection method based on PSPNet-RCNN

    CN114298948A