An image upsampling method for SUVC codec
Through Legall 5/3 wavelet transformation and super-resolution neural network SR trained by super-resolution, image detail components are separated and stitched to generate high-resolution image enhancement layers, solving the problem of high-frequency detail restoration in ultra-high-definition video editing and improving editing efficiency.
Patent Information
- Application Number
- CN202210353737.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-06
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2042-04-06
AI Technical Summary
Existing image upsampling methods fail to efficiently and accurately restore high-frequency details such as edges and textures in ultra-high-definition videos, resulting in inefficient editing.
The super-resolution neural network SR with super-resolution training is used to separate the image into approximate components, horizontal detail components, vertical detail components and diagonal detail components through Legall 5/3 wavelet transformation, and residual processing and stitching are performed to generate the SUVC enhancement layer of high-resolution images.
It saves the memory overhead of the super-score model when predicting ultra-high-definition scenes, improves the prediction speed of the model, and realizes efficient video editing.
Smart Images

Figure CN114710673B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of ultra-high-definition video coding and decoding, and more specifically, to an image upsampling method for SUVC coding and decoding. Background Art
[0002] When it comes to editing 4K or 8K ultra-high-definition videos, the mainstream editing methods, namely original bit rate editing and proxy bit rate editing, are difficult to achieve the requirements of low latency and high efficiency.
[0003] Layered coding and decoding technology (proposed by the applicant of this application, referred to as SUVC format) is mainly aimed at ultra-high-definition video (8K, 16K), etc. Its advantage is that it can reduce the bandwidth for reading files on the editing side. While retaining high image quality, the editing side can directly extract the basic layer code stream for efficient editing, solving the performance problems caused by original bit rate editing and proxy bit rate editing under ordinary coding and decoding methods. For details, see Chinese patent CN113271467B.
[0004] Current image upsampling methods, such as interpolation, cannot efficiently and accurately restore more high-frequency details of the image, such as edges and textures. Summary of the Invention
[0005] The purpose of the present invention is to overcome the shortcomings of the existing technology and provide an image upsampling method for SUVC codec, which can save the huge memory overhead of the super-resolution model when predicting ultra-high-definition scenes (such as 4K to 8K), improve the prediction speed of the model, and realize efficient video editing with diverse resolution requirements.
[0006] The object of the present invention is achieved through the following solutions:
[0007] An image upsampling method for SUVC encoding and decoding includes the following steps:
[0008] The original decoded low-resolution image is inferred using a supervised trained super-resolution neural network SR to obtain a high-resolution image approximation component, horizontal detail component, vertical detail component and diagonal detail component separated in the frequency domain; and the low-frequency information component predicted by the trained super-resolution neural network SR is subjected to residual processing with the information of the original decoded low-resolution image to generate an SUVC enhancement layer of the high-resolution image.
[0009] Furthermore, the training process of the supervised super-resolution neural network SR includes utilizing gradient descent and supervised learning methods, wherein the label of the training data set is high-resolution image wavelet frequency domain information, and the high-resolution image wavelet frequency domain information is transformed by Legall 5 / 3 wavelet to obtain the approximate component, horizontal detail component, vertical detail component, and diagonal detail component, which are first concatenated to form the training set label of the super-resolution neural network SR.
[0010] Furthermore, the generation of the SUVC enhancement layer of the high-resolution image includes the sub-steps of: splicing the image after the residual processing with the horizontal detail component, vertical detail component and diagonal detail component predicted by the super-resolution neural network SR to form the upsampled high-resolution image enhancement layer information, and then performing entropy coding to form an enhancement layer code stream.
[0011] Furthermore, the splicing includes splicing in order from top to bottom and from left to right.
[0012] Furthermore, the splicing together includes splicing together from left to right and from top to bottom.
[0013] Furthermore, the original decoded low-resolution image includes a decoded 4K ultra-high-definition video low-resolution image.
[0014] Furthermore, the entropy coding includes any one of flow coding, Huffman coding, and arithmetic coding.
[0015] Furthermore, the super-resolution neural network SR includes any one of the neural network structures of SESR, FSRCNN, and RCAN.
[0016] Furthermore, the high-resolution image wavelet frequency domain information includes 8K Legall 5 / 3 wavelet frequency domain information.
[0017] Furthermore, the 8K Legall 5 / 3 wavelet frequency domain information includes an 8K ultra-high-definition Legall 5 / 3 wavelet approximation component, a horizontal detail component, a vertical detail component, and a diagonal detail component.
[0018] The beneficial effects of the present invention include:
[0019] The present invention uses wavelet transform to separate low-frequency and high-frequency images. If the super-resolution SR network pays more attention to the high-frequency part of the image, the wavelet transform has the ability to reduce the image format. The lower format also brings lower data throughput, which can save the huge memory overhead of the super-resolution model when predicting ultra-high-definition scenes (such as 4K to 8K) and improve the prediction speed of the model.
[0020] The wavelet super-resolution model proposed in the present invention is superior to the spatial domain training model in terms of indicators such as video memory consumption and single-card prediction speed. The neural network training model based on wavelet transform helps to reduce the calculation and throughput of the model.
[0021] This invention is applied to the SUVC video codec framework, which is based on layered coding. 4K or 8K ultra-high-definition images undergo wavelet transform pre-processing to separate approximation, horizontal detail, vertical detail, and diagonal detail components. These components are then processed through residual, splicing, and entropy coding to form base layer and enhancement layer streams. These two streams are distinguished during encapsulation, preserving high image quality while allowing the editor to directly extract the base layer stream for efficient editing, enabling efficient video editing for a variety of resolution requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0023] Figure 1 This is a flow chart of a method for implementing image upsampling based on a neural network of Legall 5 / 3 wavelet transform for SUVC encoding and decoding according to an embodiment of the present invention;
[0024] Figure 2 Flowchart of the steps of an embodiment of the present invention. DETAILED DESCRIPTION
[0025] All features disclosed in all embodiments in this specification, or steps in all methods or processes implicitly disclosed, except for mutually exclusive features and / or steps, can be combined and / or expanded or replaced in any manner.
[0026] like Figure 1 As shown, the embodiment of the present invention is aimed at the decoded low-resolution image, and proposes a method for generating a high-resolution image enhancement layer code stream in SUVC codec based on Legall 5 / 3 wavelet transform and neural network. The original decoded 4K ultra-high-definition video is inferred by the supervised trained super-resolution neural network SR to obtain the high-resolution image approximate component, horizontal detail component, vertical detail component and diagonal detail component separated in the frequency domain, and the SUVC enhancement layer of the high-resolution image is generated in combination with residual processing. This technology combines Legall 5 / 3 wavelet transform to save the huge memory overhead of the SR model when predicting ultra-high-definition scenes, improves the prediction speed of the model, and realizes the upsampling of the image from the base layer to reconstruct the enhancement layer. As shown Figure 2As shown, the specific process of the embodiment of the present invention is described as follows:
[0027] S1 decodes 4K images at their original bitrate.
[0028] S2 passes the decoded 4K image as input to the trained super-resolution SR neural network, and obtains the 8K Legall 5 / 3 wavelet frequency domain information after image upsampling through neural network inference, namely the 8K ultra-high-definition Legall 5 / 3 wavelet approximation component, horizontal detail component, vertical detail component, and diagonal detail component. The specific subdivision steps are as follows:
[0029] The architecture of the S21 super-resolution SR neural network includes but is not limited to SESR, FSRCNN, RCAN and other neural network structures.
[0030] The S22 super-resolution neural network training process utilizes gradient descent and supervised learning. The labels of the training dataset—the wavelet frequency domain information of the high-resolution image—are obtained by performing a Legall 5 / 3 wavelet transform to obtain the approximate, horizontal, vertical, and diagonal detail components. These components are then concatenated from top to bottom and left to right to form the training set labels.
[0031] S3 performs residual processing on the low-frequency information LL component predicted by the SR neural network and the information after decoding the original low-resolution image video to reduce the size of the data.
[0032] S4 splices the image after residual of the detailed components with the horizontal detail component LH, vertical detail component HL and diagonal detail component HH predicted by the SR neural network in S2 in a left-to-right and top-to-bottom manner to form the upsampled high-resolution image enhancement layer information, and then forms the enhancement layer code stream through entropy coding (such as process coding, Huffman coding, arithmetic coding, etc.).
[0033] Example 1
[0034] An image upsampling method for SUVC encoding and decoding includes the following steps:
[0035] The original decoded low-resolution image is inferred using a supervised trained super-resolution neural network SR to obtain a high-resolution image approximation component, horizontal detail component, vertical detail component and diagonal detail component separated in the frequency domain; and the low-frequency information component predicted by the trained super-resolution neural network SR is subjected to residual processing with the information of the original decoded low-resolution image to generate an SUVC enhancement layer of the high-resolution image.
[0036] Example 2
[0037] Based on Example 1, the training process of the supervised super-resolution neural network SR includes using gradient descent and supervised learning methods, the label of the data set participating in the training is the high-resolution image wavelet frequency domain information, and the high-resolution image wavelet frequency domain information is obtained by Legall 5 / 3 wavelet transform to obtain the approximate component, horizontal detail component, vertical detail component and diagonal detail component, which are first spliced to form the training set label of the super-resolution neural network SR.
[0038] Example 3
[0039] Based on Example 1, the SUVC enhancement layer of the high-resolution image is generated, including the sub-steps of: splicing the image after the residual processing with the horizontal detail component, vertical detail component and diagonal detail component predicted by the super-resolution neural network SR to form the upsampled high-resolution image enhancement layer information, and then performing entropy coding to form an enhancement layer code stream.
[0040] Example 4
[0041] Based on Example 2, the splicing includes splicing from top to bottom and from left to right.
[0042] Example 5
[0043] Based on Example 3, the splicing together includes splicing together from left to right and from top to bottom.
[0044] Example 6
[0045] Based on Example 1, the original decoded low-resolution image includes a decoded 4K ultra-high-definition video low-resolution image.
[0046] Example 7
[0047] Based on Example 3, the entropy coding includes any one of flow coding, Huffman coding, and arithmetic coding.
[0048] Example 8
[0049] Based on Example 1, the super-resolution neural network SR includes any one of the neural network structures of SESR, FSRCNN, and RCAN.
[0050] Example 9
[0051] Based on Example 2, the high-resolution image wavelet frequency domain information includes 8KLegall5 / 3 wavelet frequency domain information.
[0052] Example 10
[0053] Based on Example 9, the 8K Legall 5 / 3 wavelet frequency domain information includes 8K ultra-high-definition Legall 5 / 3 wavelet approximation components, horizontal detail components, vertical detail components and diagonal detail components.
[0054] The units involved in the embodiments of the present invention may be implemented in software or hardware, and the units described may also be provided in a processor. In some cases, the names of these units do not limit the units themselves.
[0055] According to one aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described above.
[0056] As another aspect, the present application further provides a computer-readable medium, which may be included in the electronic device described in the above embodiments, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device implements the method described in the above embodiments.
[0057] The parts not involved in the present invention are the same as the existing technology or can be implemented by using the existing technology.
[0058] The above technical solution is only one embodiment of the present invention. For those skilled in the art, it is easy to make various types of improvements or modifications based on the application methods and principles disclosed in the present invention, and it is not limited to the method described in the above specific embodiment of the present invention. Therefore, the method described above is only preferred and does not have a restrictive meaning.
[0059] In addition to the above examples, those skilled in the art may obtain other embodiments based on the above disclosure or by utilizing knowledge or technology in related fields to make modifications. The features of each embodiment may be interchangeable or replaced. The modifications and changes made by those skilled in the art do not depart from the spirit and scope of the present invention and should be within the scope of protection of the claims attached to the present invention.
Claims
1. An image upsampling method for SUVC codec, characterized in that: The steps include: The original decoded low-resolution image is inferred using a supervised trained super-resolution neural network (SR) to obtain a high-resolution image approximation component, a horizontal detail component, a vertical detail component, and a diagonal detail component separated in the frequency domain; and the low-frequency information component predicted by the trained super-resolution neural network (SR) is subjected to residual processing with the information of the original decoded low-resolution image to generate an SUVC enhancement layer of the high-resolution image; The supervised training super-resolution neural network (SR) training process includes utilizing gradient descent and supervised learning methods, wherein the labels of the training data set are high-resolution image wavelet frequency domain information, and the high-resolution image wavelet frequency domain information is transformed by Legall 5 / 3 wavelet to obtain the approximate component, horizontal detail component, vertical detail component, and diagonal detail component, which are first concatenated to form the training set labels of the super-resolution neural network (SR); The SUVC enhancement layer of the high-resolution image is generated, comprising the sub-steps of: splicing the image after the residual processing with the horizontal detail component, vertical detail component, and diagonal detail component predicted by the super-resolution neural network SR to form the upsampled high-resolution image enhancement layer information, and then performing entropy coding to form an enhancement layer code stream.
2. The image upsampling method for SUVC encoding and decoding according to claim 1, characterized in that The first splicing includes splicing in order from top to bottom and from left to right.
3. The image upsampling method for SUVC encoding and decoding according to claim 1, characterized in that The splicing together includes splicing together from left to right and from top to bottom.
4. The image upsampling method for SUVC encoding and decoding according to claim 1, characterized in that The original decoded low-resolution image includes a decoded 4K ultra-high-definition video low-resolution image.
5. The image upsampling method for SUVC encoding and decoding according to claim 1, characterized in that: The entropy coding includes any one of flow coding, Huffman coding, and arithmetic coding.
6. The image upsampling method for SUVC encoding and decoding according to claim 1, characterized in that: The super-resolution neural network SR includes any one of the neural network structures of SESR, FSRCNN, and RCAN.
7. The image upsampling method for SUVC encoding and decoding according to claim 1, characterized in that: The high-resolution image wavelet frequency domain information includes 8K Legall 5 / 3 wavelet frequency domain information.
8. The image upsampling method for SUVC encoding and decoding according to claim 7, characterized in that: The 8K Legall 5 / 3 wavelet frequency domain information includes an 8K ultra-high-definition Legall 5 / 3 wavelet approximation component, a horizontal detail component, a vertical detail component, and a diagonal detail component.
Citation Information
Patent Citations
A Layered Encoding and Decoding Method for Ultra-High-Definition Video that Supports Efficient Editing
CN113271467B
Image super-resolution reconstruction method based on wavelet transformation and convolutional neural network
CN106991648A