Video Bit Depth Enhancement Method and Device Based on Pyramid Multi-Level Information Fusion
By adopting the pyramid multi-level information fusion method in video bit depth enhancement, and using technologies such as deformation convolution and step convolution, the problem of insufficient inter-frame information utilization in the existing methods is solved, and high-quality high-bit depth video sequence reconstruction is achieved.
Patent Information
- Application Number
- CN202210282812.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-22
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-03-22
AI Technical Summary
The existing video bit depth enhancement methods are difficult to effectively utilize inter-frame information, resulting in the existence of inter-frame jitter and redundant information, and fail to fully consider some special distortions in bit depth enhancement, such as pseudo-contour distortion and color distortion.
The video bit depth enhancement method based on pyramid multi-level information fusion is adopted, and implicit alignment is used through the feature alignment module to reduce the redundant information and jitter phenomenon between frames; in the pyramid feature extraction and fusion module, multi-level spatio-temporal features are extracted by step convolution and transposed convolution, and residual fusion is performed to fully explore the spatio-temporal information.
Improve the quality of the reconstructed high-bit depth target frames, reduce inter-frame jitter and redundant information, and improve the visual experience of video sequences.
Smart Images

Figure CN114663306B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video bit-depth enhancement, and in particular, to a method and device for video bit-depth enhancement based on pyramid multi-level information fusion. Background Art
[0002] To meet the needs of people's pursuit of high-quality visual experience, the field of bit-depth enhancement plays a decisive role. Traditional standard dynamic range (SDR) images all use 8 bits to represent the pixel values at each position in the channel. The greater the bit-depth of each channel in the image, the wider the brightness range and color range that can be represented, and the higher the subjective quality in terms of vision. Therefore, compared with their corresponding low-bit-depth content, high-bit-depth images or video frames can have richer colors and more natural transitions, indirectly improving the human visual experience.
[0003] Bit-depth enhancement can be regarded as an inverse quantization operation on the image. By inputting a low-bit-depth image or video frame, a corresponding high-bit-depth image or video frame is reconstructed. Most of the existing methods are based on image bit-depth enhancement methods. For traditional methods, there is the Minimum Risk based Classification (MRC) [1] and the Intensity Potential for Adaptive De-quantization (IPAD) [2] ; for deep learning methods, there is the Bit-Depth Enhancement via Convolutional Neural Network (BE-CNN) [3] and the Lighter but Efficient Bit-Depth Expansion Network (LBDEN) [4] . Image-based algorithms only utilize spatial information. However, for continuous video sequences, in addition to spatial information, temporal information between frames also needs to be considered. Currently, there are few video-based bit-depth enhancement algorithms, such as the Spatiotemporal Symmetric Convolutional Neural Network for Video Bit-Depth Enhancement (SSCNN) [5] .
[0004] For the video bit-depth enhancement method, the alignment of inter-frame information is particularly important. Through alignment, redundant inter-frame information can be reduced, and inter-frame jitter distortion can be minimized. In addition, some special distortions in bit-depth enhancement need to be considered, such as pseudo-contour distortion and color aberration. Therefore, the video bit-depth enhancement method needs to utilize the spatio-temporal information between adjacent frames to reconstruct a high-quality high-bit-depth video sequence. Summary of the Invention
[0005] The present invention provides a video bit-depth enhancement method and apparatus based on pyramid multi-level information fusion. The present invention constructs a feature alignment and pyramid feature extraction module. In the feature alignment module, the deformable convolution is used to implicitly align the inter-frame information of adjacent frames and the target frame, reducing redundant inter-frame information and inter-frame jitter phenomena. In the pyramid feature extraction and fusion module, the strided convolution with a convolution stride of 2 is used to upsample the feature blocks, and the transposed convolution with the same stride of 2 (i.e., the inverse process of convolution) is used when downsampling to restore the feature block scale, and the spatio-temporal information is fully exploited by combining with the shared dense module. The present invention improves the quality of the reconstructed high-bit-depth target frame, as described in detail below:
[0006] In a first aspect, a video bit-depth enhancement method based on pyramid multi-level information fusion, the method comprising:
[0007] Input consecutive zero-padded low-bit-depth video frames, and align the adjacent frames with the target frame through the feature alignment module to generate aligned features;
[0008] Input the aligned features into the pyramid feature extraction and fusion module, extract multi-level spatio-temporal features, and perform residual fusion;
[0009] Send the fused spatio-temporal features into the high-bit-depth reconstruction module to output a predicted residual map, and finally add it to the input zero-padded low-bit-depth target frame image to obtain the reconstructed high-bit-depth target frame image;
[0010] Subtract the input zero-padded low-bit target frame image from the real high-bit-depth target frame image to obtain a real residual map, and use the mean square error between the network predicted residual map and the real residual map as the loss function.
[0011] Wherein, the feature alignment module is an implicit alignment operation composed of two layers of deformable convolution,
[0012] Convolve the input low-bit-depth adjacent frame and the target frame to obtain the corresponding feature maps f t+i and f t , and send them into two convolutional layers through a concatenation operation to obtain preliminary offset features Send the preliminary offset features and the adjacent frame feature map ft+i They are sent into the deformable convolution together to obtain the preliminary aligned features
[0013] Then, the preliminary aligned features and the target frame feature map f t are continued to be concatenated to obtain the final offset features The final offset features and the preliminary aligned features are sent into the second - layer deformable convolution to obtain the final aligned features
[0014] Furthermore, the pyramid feature extraction and fusion module includes:
[0015] Extraction part: After aligning the input low - bit - depth feature maps with the target frame feature map respectively, they are concatenated into aligned feature blocks according to the channel dimension; two strided convolution operations with a stride of 2 are performed on the aligned feature blocks respectively to obtain feature blocks with two downsampling multiples, and the hierarchical feature blocks of three different scales are sent into the shared dense unit to extract spatio - temporal information features;
[0016] Fusion part: After sending the aligned feature map of each frame into the residual unit, it is concatenated with the extracted spatio - temporal feature map, and then added to the original aligned feature map to obtain the fused spatio - temporal information features.
[0017] Among them, the shared dense unit is composed of three identical residual units in cascade. There are two convolutional units inside each residual unit, and the extracted features are upsampled back to the initial scale through transposed convolution to obtain the final spatio - temporal feature map.
[0018] In a second aspect, a video bit - depth enhancement device based on pyramid multi - level information fusion, characterized in that the device includes: a processor and a memory, and program instructions are stored in the memory. The processor calls the program instructions stored in the memory to enable the device to execute the method steps described in any one of the first aspect.
[0019] In a third aspect, a computer - readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor is enabled to execute the method steps described in any one of the first aspect.
[0020] The beneficial effects of the technical solution provided by the present invention are:
[0021] 1. The present invention performs implicit alignment through the feature alignment module composed of deformable convolution, better mines the inter - frame information, and reduces the inter - frame jitter;
[0022] 2. The present invention utilizes the pyramid feature extraction and fusion module to introduce multi-level information of different scales in cooperation with the shared dense unit step, fully extracting spatio-temporal information features, so that the reconstructed high-bit-depth target frame is closer to the true value. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 FIG. is a flowchart of a video bit-depth enhancement method based on pyramid multi-level information fusion;
[0024] Figure 2 FIG. is the overall framework of the video bit-depth enhancement network;
[0025] Figure 3 FIG. is a schematic diagram of the feature alignment module;
[0026] Figure 4 FIG. is a schematic diagram of the pyramid feature extraction and fusion module;
[0027] Figure 5 FIG. is a schematic structural diagram of a video bit-depth enhancement device based on pyramid multi-level information fusion. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] To make the objectives, technical solutions and advantages of the present invention clearer, the following further describes the embodiments of the present invention in detail.
[0029] Embodiment 1
[0030] The embodiment of the present invention proposes a video bit-depth enhancement method based on pyramid multi-level information fusion. Referring to Figure 1 , the method includes the following steps:
[0031] 101: Make a low-bit video sequence dataset. Randomly select 1000 groups of video sequences from the Sintel high-bit-depth continuous video sequence database as the training set. Obtain the corresponding low-bit-depth video sequence through quantization processing, and perform zero-padding on it to obtain continuous zero-padded low-bit-depth video frames, which are used as the input part of the network;
[0032] Among them, the Sintel high-bit-depth continuous video sequence database is well-known to those skilled in the art, and the embodiments of the present invention will not elaborate on this.
[0033] 102: Input continuous zero-padded low-bit-depth video frames. Align adjacent frames with the target frame (middle frame) through the feature alignment module to generate aligned features;
[0034] Among them, the overall framework of the network constructed in the embodiments of the present invention is shown in Figure 2As shown, the input is five consecutive zero-padded low-bit-depth video frames, with the middle frame as the target frame, and the output is the high-bit-depth target frame corresponding to the zero-padded low-bit-depth target frame. For the feature alignment module, refer to Figure 3 , and align the input consecutive video frames with the target frame through two deformable convolutions respectively to improve the alignment efficiency and reduce inter-frame jitter.
[0035] Among them, the feature alignment module consists of an implicit alignment operation composed of two layers of deformable convolutions, which convolve the adjacent low-bit-depth frames of the input with the target frame to obtain the corresponding feature maps f t+i and f t , and send them into two convolutional layers through a concatenation operation to obtain the preliminary offset features Send the preliminary offset features and the adjacent frame feature map f t+i into the deformable convolution together to obtain the preliminary aligned features Then send the preliminary aligned features and the target frame feature map f t to continue the concatenation to obtain the final offset features Send the final offset features and the preliminary aligned features into the second layer of deformable convolution to obtain the final aligned features
[0036] 103: Input the aligned features into the pyramid feature extraction and fusion module to extract multi-level spatio-temporal features and perform residual fusion;
[0037] Refer to Figure 4 , and the pyramid feature extraction and fusion module uses strided convolution and transposed convolution to upsample and downsample the feature blocks respectively, cooperate with the shared dense unit to fully exploit the spatio-temporal information of the video frames, and then perform spatio-temporal information fusion through the residual unit.
[0038] Among them, the pyramid feature extraction and fusion module is mainly divided into two parts: extraction and fusion. For the extraction part, after aligning the input low-bit-depth feature map with the target frame feature map respectively, they are concatenated into an aligned feature block according to the channel dimension. Input the aligned feature block into the pyramid feature extraction and fusion module to perform two strided convolution operations with a stride of 2 respectively to obtain two feature blocks with downsampling multiples of 2 and 4. Send the three hierarchical feature blocks with different scales into the shared dense unit to extract spatio-temporal information features. Among them, the shared dense unit is composed of three identical residual units in cascade. Upsample the extracted features back to the original scale through a transposed convolution with the same stride (for example: stride equals 2) to obtain the final spatio-temporal feature map.
[0039] Among them, the pyramid mainly performs two downsamplings on the initial scale after the upsampling operation. The feature blocks of different scales form a pyramid shape with a "large - small - large" structure. Multiple levels are equivalent to multiple scales, and each level represents a scale. That is, by using upsampling and downsampling operations, the scale of the feature blocks is changed to form a pyramid - like structure of "large - medium - small - medium - large". And for the feature blocks of different scales in each level, sufficient extraction and fusion are carried out to achieve the information extraction and fusion operation at multiple levels. Compared with the single - level network that only considers a single scale, the multi - level operation has a greater improvement in the reconstruction performance of the network.
[0040] For the fusion part, the aligned feature maps of each frame are input into the residual fusion module. After passing through a residual unit, they are concatenated with the spatio - temporal feature maps to obtain the concatenated feature maps. After passing through a convolutional unit, they are added to the original aligned feature maps of each frame, which is the output of the residual fusion module.
[0041] The strided convolution described above is a type of convolution. By changing the stride size, the downsampling operation is achieved. The transposed convolution is the inverse operation of convolution, and the upsampling is achieved by changing the stride. The embodiments of the present invention do not elaborate on this.
[0042] 104: Send the fused spatio - temporal features into the reconstruction high - bit - depth module to output a residual map, and finally add it to the input zero - padded low - bit - depth target frame image to obtain the reconstructed high - bit - depth target frame image; see Figure 2 , the reconstruction high - bit - depth module is mainly composed of a convolutional unit CNN and a dense unit (without sharing learning parameters) DENSE. The dense unit DENSE is composed of three identical residual units RES in cascade. The spatio - temporal features output by the pyramid feature extraction and fusion module are concatenated and sent into the reconstruction high - bit - depth module. The number of channels is reduced through convolution. After passing through the dense unit, the feature map is reduced to a residual map with 3 channels for output. Finally, the residual map is added to the input zero - padded low - bit - depth target frame image to obtain the reconstructed high - bit - depth target frame image.
[0043] This reconstructed high - bit - depth target frame image is the high - bit target frame image finally output by the network. For example: input 5 low - bit frames, output one reconstructed high - bit image frame, and display it on a high - bit monitor. In practical applications, displaying low - bit images on a high - bit monitor has distortion. Therefore, through the network model designed in the embodiments of the present invention, the reconstructed high - bit image frame is output, and the distortion is reduced on the high - bit monitor.
[0044] 105: Subtract the input zero - padded low - bit target frame image from the real high - bit - depth target frame image to obtain the real residual map, and use the mean square error between the network - predicted residual map and the real residual map as the loss function.
[0045] In the embodiment of the present invention, the learning rate is initially set to 0.00005. After every 30 training iterations, the learning rate is reduced to half of the initial value, and the Adam optimizer is used [6] to optimize the network model parameters.
[0046] In summary, the embodiment of the present invention designs a video bit-depth enhancement method based on pyramid multi-level information fusion. This method uses a preprocessed low-bit video sequence of 5 consecutive zero-padded frames as the network input, adopts a feature alignment module dominated by deformable convolution for implicit alignment, then enters the pyramid feature extraction and fusion module to extract spatio-temporal features, and finally sends the fused spatio-temporal features into the reconstruction high-bit depth module, where the reconstruction quality of the high-bit depth target frame is improved through residual learning.
[0047] Embodiment 2
[0048] The solution in Embodiment 1 is further introduced below in conjunction with the accompanying drawings and examples. See the following description for details:
[0049] 201: Sintel database [7] is a CG animated short film with more than 20,000 consecutive high-bit depth (16-bit) images, and the resolution of each frame is 436×1024. In the production of the data set, 1000 groups of video sequences are randomly selected, and each group contains 5 consecutive video frames. Each group of high-bit depth (16-bit) is quantized to low-bit depth (4-bit). For the training set, each group of video sequences is randomly cropped 24 times to images with a resolution of 128×128 for convenient training.
[0050] 202: The network input is an image obtained by zero-padding the low-bit depth video frame to high-bit and then normalizing it. And before entering the feature alignment module, the 5 video frames are converted into 5 feature maps f={f t-n ,…,f t-1 ,f t ,f t+1 ,…,f t+n} with the number of channels changed to 64 through convolution operations;
[0051] The convolution size in the network is all 3*3, and a PReLU activation function is added later. The feature alignment module, see Figure 3 , is mainly composed of two deformable convolutions. The adjacent frame feature maps of the input low-bit depth and the target frame feature map are sent to two convolutional layers through splicing to obtain preliminary offset features The offset features and the adjacent frame feature maps are spliced and then sent to the deformable convolution to obtain preliminary aligned features which is expressed as:
[0052]
[0053]
[0054] Among them, CNN1 and CNN2 respectively represent two convolutional units, f t and f t+i respectively represent the target frame and the adjacent frame feature maps, [,] is the concatenation operation according to the channel dimension, and DCN is the deformable convolutional unit, is the preliminary offset feature of the two feature maps.
[0055] To obtain fully aligned video frame information, the same operation is used to and the target frame feature map f t for the second alignment to obtain the final which is expressed as:
[0056]
[0057]
[0058] Among them, is the final offset feature.
[0059] 203: The pyramid feature extraction and fusion module is mainly divided into two parts: extraction and fusion. For the extraction part, after aligning the input low-bit-depth feature map with the target frame feature map respectively, they are concatenated into an aligned feature block F align along the channel dimension. The aligned feature block is input into the pyramid feature extraction and fusion module for two strided convolutional operations with a stride of 2 to obtain two downsampled feature blocks After performing channel reduction convolution operations on the three feature blocks with different scales, they are sent to a shared dense unit to extract spatio-temporal information features. The extracted features are upsampled back to the original scale through transposed convolution to obtain the final spatio-temporal feature map f pyramid which is expressed as:
[0060]
[0061]
[0062]
[0063] f dense = DENSE(CNN(F align ))
[0064]
[0065]
[0066]
[0067] Among them, CNN S=2 and CNN T=2 are strided convolution and transposed convolution with a stride of 2 respectively. DENSE represents shared dense units, and CNN represents convolutional units for reducing the number of channels. and represent the feature blocks when the aligned feature blocks are downsampled to 2 and 4 respectively. f dense 、 and represent the spatio-temporal information features of different scales extracted through shared dense units. represents the target frame feature map.
[0068] For the fusion part, the aligned feature map f of each frame align is input into the residual fusion module and concatenated with the spatio-temporal feature map f pyramid through a residual unit. After passing through a convolutional unit and adding it to the original aligned feature map of each frame, it is the output of the residual fusion module.
[0069] The shared dense unit is in the form of three residual units connected in series, namely "residual-residual-residual", where each residual unit is composed of two convolutional units inside:
[0070] RES1(x) = x + CNN2(CNN1(x))
[0071] DENSE(x) = RES3(RES2(RES1(x)))
[0072] Among them, x is the input feature map, CNN1 and CNN2 represent two convolutional units, RES1, RES2 and RES3 represent three residual units, RES1(x) represents the output of the residual unit, and DENSE(x) represents the output of the shared dense unit.
[0073] During the network training process, the introduction of residuals can improve efficiency and enhance network performance. In the pyramid feature extraction and fusion module, the learning method of shared parameters is carried out in the shared dense unit of each layer, so that the parameters of the three shared dense units are the same, thereby accelerating the network training speed, reducing network parameters, and fully extracting spatio-temporal information features.
[0074] 204: What the network learns is the residual map between the reconstructed high-bit-depth target frame and the input zero-padded low-bit-depth target frame. By adding the residual map and the input target frame image pixel by pixel, the high-bit-depth target frame image can be obtained.
[0075] Among them, the above steps are mainly achieved through a long skip connection. See Figure 2. During training, two residual maps are also used, namely the real residual map and the residual map predicted by the network, as the input of the mean square error loss function to improve the quality of the training network and optimize the network model. The real residual map is obtained by subtracting the input zero-padded target frame image from the real high-bit-depth target frame image.
[0076] Embodiment 3
[0077] The following combines specific experimental data to evaluate the effectiveness of the solutions in Embodiments 1 and 2, as detailed in the following description:
[0078] In the embodiments of the present invention, 50 groups of Sintel [7] datasets and 30 groups of TOS [7] video sequences in the dataset are randomly selected as the test set, where the TOS [7] dataset is the same as the Sintel [7] dataset and has continuous high-bit-depth (16-bit) video frames. The evaluation metrics are respectively selected as PSNR and SSIM [8] , and the higher these two types of metrics, the better the quality of the reconstruction. The present invention is compared with the bit-depth enhancement method.
[0079] In the experiment, the present method is compared with 6 bit-depth enhancement methods, including 5 image bit-depth enhancement methods and 1 video bit-depth enhancement method.
[0080] Table 1 lists the test results of the present method and the other 6 comparison methods on the Sintel [7] test set and the TOS [7] test set. From the above experimental data, it can be seen that a video bit-depth enhancement method based on pyramid multi-level information fusion proposed in the embodiments of the present invention exceeds other bit-depth enhancement algorithms and can well reconstruct high-bit-depth video frames with high quality.
[0081] Table 1
[0082]
[0083] Embodiment 4
[0084] A video bit-depth enhancement device based on pyramid multi-level information fusion, see Figure 5 , the device includes: a processor and a memory, and program instructions are stored in the memory. The processor calls the program instructions stored in the memory to enable the device to execute the following method steps:
[0085] Input continuous zero-padded low-bit-depth video frames, and align adjacent frames with the target frame through the feature alignment module to generate aligned features;
[0086] Input the aligned features into the pyramid feature extraction and fusion module to extract multi-level spatio-temporal features and perform residual fusion;
[0087] Feed the fused spatio-temporal features into the reconstruction high-bit depth module to output the predicted residual map, and finally add it to the input zero-padded low-bit depth target frame image to obtain the reconstructed high-bit depth target frame image;
[0088] Subtract the input zero-padded low-bit target frame image from the real high-bit depth target frame image to obtain the real residual map, and use the mean square error between the network predicted residual map and the real residual map as the loss function.
[0089] Among them, the feature alignment module is: an implicit alignment operation composed of two layers of deformable convolution,
[0090] Convolve the input low-bit depth adjacent frame and the target frame to obtain the corresponding feature maps f t+i and f t , and send them into two convolutional layers through a splicing operation to obtain the preliminary offset features Send the preliminary offset features and the adjacent frame feature map f t+i into the deformable convolution together to obtain the preliminary alignment features
[0091] Then send the preliminary alignment features and the target frame feature map f t to continue splicing to obtain the final offset features Send the final offset features and the preliminary alignment features into the second layer of deformable convolution to obtain the final alignment features
[0092] Furthermore, the pyramid feature extraction and fusion module includes:
[0093] Extraction part: After aligning the input low-bit depth feature map with the target frame feature map respectively, splice them into an aligned feature block according to the channel dimension; perform two strided convolution operations with a stride of 2 on the aligned feature block respectively to obtain two feature blocks with different downsampling multiples, and send the three hierarchical feature blocks with different scales into the shared dense unit to extract spatio-temporal information features;
[0094] Fusion part: After sending the aligned feature map of each frame into the residual unit, splice it with the extracted spatio-temporal feature map, and then add it to the original aligned feature map to obtain the fused spatio-temporal information feature.
[0095] Among them, the shared dense unit is composed of three identical residual units connected in cascade. Each residual unit contains two convolutional units inside. The extracted features are upsampled back to the initial scale through transposed convolution to obtain the final spatio-temporal feature map.
[0096] It should be noted here that the device descriptions in the above embodiments correspond to the method descriptions in the embodiments, and the embodiments of the present invention will not be elaborated herein.
[0097] The execution subjects of the above-mentioned processor 1 and memory 2 can be devices with computing functions such as a computer, a single-chip microcomputer, a microcontroller, etc. In specific implementation, the embodiments of the present invention do not limit the execution subject and can be selected according to the needs in actual applications.
[0098] Data signals are transmitted between the memory 2 and the processor 1 through the bus 3, and the embodiments of the present invention will not be elaborated herein.
[0099] Based on the same inventive concept, the embodiments of the present invention also provide a computer-readable storage medium. The storage medium includes a stored program that controls the device where the storage medium is located to execute the method steps in the above embodiments when the program runs.
[0100] The computer-readable storage medium includes, but is not limited to, flash memory, hard disk, solid-state drive, etc.
[0101] It should be noted here that the description of the readable storage medium in the above embodiments corresponds to the method description in the embodiments, and the embodiments of the present invention will not be elaborated herein.
[0102] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part.
[0103] The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium or a semiconductor medium, etc.
[0104] References
[0105] [1] Mittal G, Jakhetiya V, Jaiswal S P, et al. Bit-depth expansion using minimum risk based classification[C] / / 2012 Visual Communications and Image Processing. IEEE, 2012: 1-5.
[0106] [2] Liu J, Zhai G, Liu A, et al. IPAD: Intensity potential for adaptive de-quantization[J]. IEEE Transactions on Image Processing, 2018, 27(10): 4860-4872.
[0107] [3] Liu J, Sun W, Liu Y. Bit-depth enhancement via convolutional neural network[C] / / International Forum on Digital TV and Wireless Multimedia Communications. Springer, Singapore, 2017: 255-264.
[0108] [4] Zhao Y, Wang R, Chen Y, et al. Lighter but Efficient Bit-Depth Expansion Network[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2020, PP(99): 1-1.
[0109] [5] Liu J, Liu P, Su Y, et al. Spatiotemporal symmetric convolutional neural network for video bit-depth enhancement[J]. IEEE Transactions on Multimedia, 2019, 21(9): 2397-2406.
[0110] [6]Kingma D P,Ba J.Adam:A method for stochastic optimization[J].arXivpreprint arXiv:1412.6980,2014.
[0111] [7]Foundation X.Xiph.Org,https: / / www.xiph.org / ,2016.
[0112] [8]Wang Z,Bovik AC,Sheikh H R,et al.Image quality assessment:fromerror visibility to structural similarity[J].IEEE transactions on imageprocessing,2004,13(4):600-612.
[0113] [9]Liu J,Sun W,Su Y,et al.BE-CALF:bit-depth enhancement byconcatenating all level features of DNN[J].IEEE Transactions on ImageProcessing,2019,28(10):4926-4940.
[0114] In the embodiments of the present invention, unless otherwise specified for the models of each device, the models of other devices are not limited, and any device that can perform the above functions can be used.
[0115] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred embodiment, and the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0116] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A video bit-depth enhancement method based on pyramid multi-level information fusion, characterized in that, The method includes: Inputting consecutive zero-padded low-bit-depth video frames, aligning adjacent frames with a target frame through a feature alignment module to generate aligned features; Inputting the aligned features into a pyramid feature extraction and fusion module to extract multi-level spatio-temporal features and perform residual fusion; Feeding the fused spatio-temporal features into a reconstruction high-bit-depth module to output a predicted residual map, and finally adding it to the input zero-padded low-bit-depth target frame image to obtain the reconstructed high-bit-depth target frame image; Subtracting the input zero-padded low-bit target frame image from the real high-bit-depth target frame image to obtain a real residual map, and using the mean square error between the network predicted residual map and the real residual map as a loss function; Among them, the feature alignment module is an implicit alignment operation composed of two layers of deformable convolution; Convolve the input adjacent frames with low bit depth and the target frame to obtain the corresponding feature map f t+i and f t , and send them into two convolutional layers through a concatenation operation to obtain preliminary offset features Send the preliminary offset features and the adjacent frame feature map f t+i into the deformable convolution together to obtain preliminary alignment features Then, the preliminary alignment features and the target frame feature map f t are continued to be concatenated to obtain the final offset feature The final offset feature and the preliminary alignment features are fed into the second layer of deformable convolution to obtain the final alignment feature 2. The video bit-depth enhancement method based on pyramid multi-level information fusion according to claim 1, characterized in that, The pyramid feature extraction and fusion module includes: Extraction part: After aligning the input low-bit-depth feature map with the target frame feature map respectively, concatenating them into an aligned feature block according to the channel dimension; performing two strided convolution operations with a stride of 2 on the aligned feature block to obtain two feature blocks with different downsampling multiples respectively, and sending the three hierarchical feature blocks with different scales into a shared dense unit to extract spatio-temporal information features; Fusion part: After sending the aligned feature map of each frame into a residual unit, concatenating it with the extracted spatio-temporal feature map, and then adding it to the original aligned feature map to obtain the fused spatio-temporal information feature.
3. The video bit-depth enhancement method based on pyramid multi-level information fusion according to claim 1, characterized in that, The shared dense unit is composed of three cascaded identical residual units, and each residual unit has two convolutional units inside. The extracted features are upsampled back to the initial scale through transposed convolution to obtain the final spatio-temporal feature map.
4. A video bit-depth enhancement device based on pyramid multi-level information fusion, characterized in that, The device includes: a processor and a memory. Program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the device to execute the method steps described in any one of claims 1-3.
5. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by the processor, the processor is enabled to execute the method steps described in any one of claims 1-3.
Citation Information
Patent Citations
Video bit enhancement method based on efficient spatio-temporal information fusion
CN113066022A
Power monitoring video deblurring method based on depth separable residual network
CN113888426A