Video encoding method, apparatus, terminal device, medium, and program product
By using a first downsampling network model trained on a super-resolution network model, the problem of insufficient high-frequency information retention in traditional downsampling methods is solved, achieving higher quality video coding results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- VIVO MOBILE COMM CO LTD
- Filing Date
- 2025-01-07
- Publication Date
- 2026-07-07
AI Technical Summary
Traditional downsampling methods have limited ability to retain high-frequency information when reducing image resolution, making it difficult for downsampled images to be restored to high-quality, high-resolution images in super-resolution processing, thus affecting video coding quality.
A first downsampling network model trained based on a super-resolution network model is adopted. By downsampling video frames, image blocks that retain more high-frequency information are generated, and residual blocks are generated by combining the prediction blocks to improve the video coding quality.
By coupling the super-resolution network model, the first downsampling network model can output high-frequency information that is easier to recover by the super-resolution network model, thereby improving the overall video coding quality.
Smart Images

Figure CN122349022A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of video encoding and decoding technology, specifically relating to a video encoding method, apparatus, terminal equipment, medium, and program product. Background Technology
[0002] In the field of video encoding and decoding, in order to reduce encoding complexity and improve encoding efficiency, the original high-resolution video frames are usually downsampled.
[0003] Currently, traditional downsampling methods, such as average pooling and bilinear interpolation, can reduce image resolution, but their ability to preserve high-frequency details is limited. This results in the loss of a significant amount of high-frequency information in the low-resolution image obtained after downsampling, especially in areas with complex textures. Consequently, it becomes difficult to restore the video frame to a high-quality, high-resolution image in subsequent super-resolution processing. Therefore, how to retain more high-frequency information during downsampling to improve the quality of video encoding has become an urgent problem to be solved. Summary of the Invention
[0004] This application provides a video encoding method, apparatus, terminal device, medium, and program product that enables the first downsampling network model to output image blocks that retain more high-frequency information that is beneficial for the super-resolution network model to recover, thereby improving the overall video encoding quality.
[0005] In a first aspect, a video coding method is provided, the method comprising: inputting a first video frame into a first downsampling network model and outputting at least one first image block, wherein the first downsampling network model is trained based on a super-resolution network model and is used to reduce the resolution of the video frame; generating a first residual block based on the first image block and a first prediction block, wherein the first image block is an image block among the at least one image block, and the first prediction block is obtained by predicting the first image block.
[0006] In a second aspect, a video encoding apparatus is provided, comprising: a processing module; the processing module being configured to input a first video frame into a first downsampling network model and output at least one image block, wherein the first downsampling network model is trained based on a super-resolution network model and is used to perform downsampling processing on the video frame; the processing module being further configured to generate a first residual block based on the first image block and a first prediction block, wherein the first image block is an image block among the at least one image block, and the first prediction block is obtained by predicting the first image block.
[0007] Thirdly, a terminal device is provided, the terminal including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the first aspect.
[0008] Fourthly, a terminal device is provided, including a processor and a communication interface, wherein the processor is used to input a first video frame into a first downsampling network model and output at least one first image block, the first downsampling network model being trained based on a super-resolution network model, the first downsampling network model being used to reduce the resolution of the video frame; and to generate a first residual block based on the first image block and a first prediction block, wherein the first image block is an image block among at least one image block, and the first prediction block is obtained by predicting the first image block.
[0009] Fifthly, a readable storage medium is provided, on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0010] In a sixth aspect, a chip is provided, the chip including a processor and a communication interface coupled to the processor, the processor being used to run programs or instructions to implement the steps of the method described in the first aspect.
[0011] In a seventh aspect, a computer program / program product is provided, the computer program / program product being stored in a storage medium, the computer program / program product being executed by at least one processor to implement the steps of the video encoding method as described in the first aspect.
[0012] In this embodiment, a first video frame is input into a first downsampling network model, which outputs at least one first image block. This first downsampling network model is trained based on a super-resolution network model and is used to downsample the video frame. Based on the at least one image block and at least one prediction block, at least one residual block is generated. The at least one prediction block is obtained by predicting the at least one first image block. This method reduces the resolution of the video data by downsampling the first video frame using the first downsampling network model. Since the first downsampling network is jointly trained with the super-resolution network model, it has good coupling with the super-resolution network model, enabling it to output image blocks that retain more high-frequency information that is easier for the super-resolution network model to recover, thereby improving the overall video coding quality. Attached Figure Description
[0013] Figure 1A schematic diagram of the architecture of an encoding / decoding system provided for some embodiments of this application;
[0014] Figure 2 A schematic block diagram of the encoder structure used in the video encoding method provided for some embodiments of this application;
[0015] Figure 3 Schematic block diagram of the structure of the decoder used in the video encoding method provided for some embodiments of this application;
[0016] Figure 4 A flowchart illustrating a video encoding method provided for some embodiments of this application;
[0017] Figure 5 A schematic diagram of a downsampling network model provided for some embodiments of this application;
[0018] Figure 6 A schematic diagram of a residual block in a downsampling network model provided for some embodiments of this application;
[0019] Figure 7 A framework diagram of downsampling joint training provided for some embodiments of this application;
[0020] Figure 8 One of the structural schematic diagrams of a video encoding apparatus provided for some embodiments of this application;
[0021] Figure 9 A second schematic diagram of the structure of a video encoding apparatus provided for some embodiments of this application;
[0022] Figure 10 This is a schematic diagram of the hardware structure of a terminal provided for some embodiments of this application. Detailed Implementation
[0023] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0024] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, the first object can be one or more. Furthermore, "or" in this application indicates at least one of the connected objects. For example, the scope of protection for "A or B" covers at least three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. In addition, the terms "A and / or B," "at least one of A and B," and "at least one of A or B" also cover at least the above three scenarios. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0025] The term "instruction" in this application can be either a direct instruction (or explicit instruction) or an indirect instruction (or implicit instruction). A direct instruction can be understood as the sender explicitly informing the receiver of specific information, the required operation, or the requested result in the instruction sent. An indirect instruction can be understood as the receiver determining the corresponding information based on the instruction sent by the sender, or making a judgment and determining the required operation or requested result based on the judgment result.
[0026] In the description of the embodiments of this application, "at least one (item)," "at least one of," etc., refer to any one, any two, or a combination of two or more of the included objects. For example, at least one (item) of a, b, and c can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two (items)" refers to two or more, and its meaning is similar to that of "at least one (item)."
[0027] In the description of the embodiments of this application, "multiple" means two or more. For example, multiple image blocks refer to two or more image blocks. "At least two" has a similar meaning to "multiple," and in some embodiments, the two can be used interchangeably.
[0028] In the description of embodiments of this application, the terms "including," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0029] The following explains the nouns or terms used in the embodiments of this application.
[0030] Video encoding technology:
[0031] Video sequences contain a series of redundant information, including spatial redundancy, temporal redundancy, visual redundancy, information entropy redundancy, structural redundancy, knowledge redundancy, and importance redundancy. To remove as much redundant information as possible from video sequences and reduce the amount of data representing the video, video coding techniques have been proposed to reduce storage space and save transmission bandwidth. Video coding techniques are also known as video compression techniques.
[0032] Internationally accepted video compression coding standards include, for example: Advanced Video Coding (AVC) in Part 10 of the MPEG-2 and MPEG-4 standards developed by the Motion Picture Experts Group (MPEG); H.263, H.264, and H.265 (also known as High Efficiency Video Coding standard (HEVC)) developed by the International Telecommunication Union-Telecommunication Standardization Sector (ITU-T); and H.266 (also known as Versatile Video Coding (VVC)) developed by the Joint Video Experts Team (JVET), which is composed of the Video Coding Experts Group (VCEG) under MPEG and ITU-T. H.266 / VVC, as a next-generation video coding standard, not only helps users store more high-definition video on their devices, thereby reducing network data traffic, but also supports high resolution, high dynamic range, and screen content encoding in the main10 profile. Compared to the previous generation standard H.265 / HEVC, the H.266 / VVC standard further improves compression performance, enabling users to reduce data size by 50% while maintaining the same subjective video quality.
[0033] It should be noted that in encoding algorithms based on a hybrid encoding architecture, the above compression encoding methods can be used in combination.
[0034] Video decoding technology:
[0035] Video decoding technology is a component of video encoding and decoding technology. Corresponding to video encoding technology, video decoding involves the process of restoring compressed video data into the original, playable video signal. Specifically, video encoding technology uses specific compression algorithms to compress the original video data into smaller files for easier storage and transmission; video decoding technology is the reverse operation, decompressing the compressed video data back into the original video signal for playback on display devices. This technology is of great significance to various industries such as IPTV, digital cinemas, distance education, and video conferencing, as it can significantly reduce the bandwidth and storage space required for storage and transmission while maintaining video quality.
[0036] Reference Picture Resampling (RPR):
[0037] RPR (Real-Time Profiler) is a crucial tool in VVC (Video Capability Control) for real-time video encoding. It can adaptively change the resolution within the bitstream without inserting Instantaneous Decoding Refresh (IDR) frames or Intra Random Access Point (IRAP) frames. RPR adjusts the resolution adaptively based on network conditions. When network bandwidth is low, it can downsample and encode low-resolution (LR) frames; when network bandwidth improves, it can upsample and encode high-resolution (HR) original frames to avoid network congestion caused by excessively large IDR or IRAP frames. RPR technology enhances the stability of video transmission and user experience, and has significant application value, especially in real-time video communication and streaming media transmission.
[0038] Super-Resolution Network (SRN):
[0039] Super-resolution networks, also known as super-resolution networks, are deep learning-based models used to upscale low-resolution images to high-resolution images. By learning the mapping relationship between low-resolution and high-resolution images, super-resolution networks can recover high-frequency details in the images, thereby generating higher-quality images.
[0040] With the widespread application of high-resolution and high-frame-rate video, users' visual experience has become more refined and smoother, but this has also brought enormous data processing challenges. To cope with the transmission and storage pressure caused by the surge in video data volume, video coding technology effectively removes redundant information by compressing video signals. To further improve video coding efficiency and reduce bandwidth consumption, the next-generation video coding standard (Versatile Video Coding, VVC) proposed Reference Picture Resampling (RPR) technology. This technology can adaptively adjust the resolution of the reference image during the encoding process without introducing additional keyframes. This method can not only compress video data when bandwidth is limited, but also ensure the restoration of video quality when network conditions are good. RPR technology allows the encoder to encode images at a lower resolution, and the decoder to perform accurate image restoration as needed, thereby optimizing data transmission efficiency and avoiding network congestion caused by introducing additional keyframes.
[0041] Furthermore, to address the demands of real-time video communication, experts have combined image super-resolution with a reference image resampling framework. Image super-resolution technology plays a crucial role in improving video quality by reconstructing low-resolution images into high-resolution ones, thereby enhancing visual effects. With technological advancements, deep learning-based super-resolution neural network upsampling methods have been proposed and have gradually replaced traditional RPR upsampling methods. However, traditional downsampling methods, such as average pooling and bilinear interpolation, while reducing image resolution, have limited ability to preserve high-frequency details. This results in the loss of significant high-frequency information in the downsampled low-resolution image, particularly in areas with complex textures. Consequently, it becomes difficult to reconstruct high-quality, high-resolution images from video frames in subsequent super-resolution processing. Therefore, how to retain more high-frequency information during downsampling to improve video coding quality has become a pressing issue.
[0042] The system architecture used in the embodiments of this application is described below.
[0043] Figure 1 A schematic diagram of the architecture of the encoding / decoding system 10 used in an embodiment of this application is shown. Figure 1 As shown, the encoding / decoding system 10 may include a source device 110 and a destination device 120. The source device 110 is used to encode images; therefore, the source device 110 may be referred to as a video encoding apparatus (device). The destination device 120 is used to decode the encoded image data generated by the source device 110; therefore, the destination device 120 may be referred to as an image decoding apparatus (device) or a video decoding apparatus (device).
[0044] The source device 110 and the destination device 120 can take various forms, and this application embodiment does not specifically limit them. For example, the source device 110 and the destination device 120 can be desktop computers, mobile computing devices, laptops (e.g., laptops), tablet computers, set-top boxes, terminal devices (e.g., so-called "smartphones" or other handsets), televisions, cameras, display devices, digital media players, video game consoles, in-vehicle computers, or other similar devices.
[0045] Optionally, Figure 1 The source device 110 and the destination device 120 shown can be two separate devices. Alternatively, the source device 110 and the destination device 120 can also be a single device, meaning that the source device 110 or its corresponding functions and the destination device 120 or its corresponding functions can be integrated into the same device.
[0046] Optionally, the source device 110 and the destination device 120 may communicate. For example, the destination device 120 may receive encoded image data from the source device 110. In one example, the source device 110 and the destination device 120 may include one or more communication devices that can be used to transmit the encoded image data from the source device 110 to the destination device 120. These one or more communication devices may include routers, switches, base stations, or any other possible devices that facilitate communication from the source device 110 to the destination device 120, specifically determined according to actual usage requirements; this embodiment does not limit this.
[0047] like Figure 1 As shown, source device 110 may include encoder 112. Optionally, source device 110 may also include image preprocessor 111 and communication interface 113. Image preprocessor 111 can be used to perform preprocessing on the received image to be encoded. For example, preprocessing performed by image preprocessor 111 may include trimming, color format conversion (e.g., from RGB to YUV format), color correction, or noise reduction, or any other possible processing. Encoder 112 can be used to receive the image preprocessed by image preprocessor 111, process the preprocessed image using a correlation prediction mode, and output encoded image data. In some embodiments, encoder 112 can be used to perform the video encoding process described in the various embodiments below. Communication interface 113 can be used to transmit the encoded image data output by encoder 112 to destination device 120 or any other device (such as a storage device) for storage or direct reconstruction. Other devices can be any devices used for decoding or storage. Of course, in actual implementation, communication interface 113 can also encapsulate the encoded image data output by encoder 112 into a suitable format before transmission.
[0048] Optionally, the image preprocessor 111, encoder 112, and communication interface 113 may be hardware components in the source device 110, software programs in the source device 110, or a combination of hardware components and software programs in the source device 110. The specific details can be determined according to actual usage requirements, and this application embodiment does not limit this.
[0049] The destination device 120 may include a decoder 122. Optionally, the destination device 120 may also include a communication interface 121 and an image post-processor 123. The communication interface 121 may be used to receive encoded image data from the source device 110 or any other source device, such as a storage device. The communication interface 121 may also decapsulate the data transmitted by the communication interface 113 to obtain encoded image data. The decoder 122 is used to receive the encoded image data and output decoded image data (also referred to as reconstructed image data or reconstructed image data). In some embodiments, the decoder 122 may be used to perform the decoding process described in the various embodiments below. The image post-processor 123 may be used to perform post-processing on the decoded image data to obtain post-processed image data. The post-processing performed by the image post-processor 123 may include color format conversion (e.g., from YUV format to RGB format), color correction, retouching, or resampling, and any possible processing. The image post-processor 123 may also be used to transmit the post-processed image data to a display device for display.
[0050] Optionally, the aforementioned communication interface 121, decoder 122, and image post-processor 123 may be hardware components in the target device 120, software programs in the target device 120, or a combination of hardware components and software programs in the target device 120. The specific details can be determined according to actual usage requirements, and this application embodiment does not limit this.
[0051] Figure 2 This is a schematic block diagram illustrating a possible structure of an encoder 112 for implementing the video encoding method of this application, as provided in an embodiment of this application. Figure 2As shown, the encoder 112 may include a downsampling network 201, an intra-frame prediction unit 202, an inter-frame prediction unit 203, a residual calculation unit 204, a transform / quantization unit 205, an encoding unit 206, an inverse quantization unit (also called a dequantization unit) 207, an inverse transform unit 208, a reconstruction unit (or reconstruction unit) 209, a loop filter unit 210, and a decoded picture buffer (DPB) 211. Optionally, the encoder 112 may also include a super-resolution network 212. The intra-frame prediction unit 202 predicts the current block to generate a prediction block; the residual calculation unit 204 calculates the difference between the original image block and the prediction block generated by the intra-frame prediction unit to obtain a residual block; the transform / quantization unit 205 performs transform coding on the residual block and quantizes the transform coefficients, mapping continuous transform coefficients to a finite number of discrete values; the encoding unit 206 further encodes the quantized data, such as using entropy coding (e.g., Huffman coding, arithmetic coding, or CABAC) to reduce redundant information in the data, and the encoded data (bitstream) is sent to the decoder or storage medium; the inverse quantization unit 207 dequantizes the quantized data to recover the transformed data. The inverse transform unit 208 performs an inverse transform on these transform coefficients to recover the residual block; the reconstruction unit 209 adds the prediction block generated by the intra-frame prediction unit to the residual block recovered by the residual inverse transform unit to reconstruct the original image block; the loop filtering unit 210 filters the reconstructed image block to improve image quality and reduce visual artifacts such as block artifacts; the decoding image buffer 211 buffers the filtered reconstructed image block output by the loop filtering unit 210, or buffers the high-resolution reconstructed image block output by the super-resolution network 212; the inter-frame prediction unit 203 acquires the image block buffered in the decoding image buffer 211 and uses the image block as a reference block for subsequent motion estimation or motion compensation.
[0052] In one example, the input to encoder 112 is video data to be encoded. Encoder 112 may include a segmentation unit that can be used to segment the image to be encoded into multiple image blocks. Encoder 112 can complete the encoding of the video frame to be encoded by encoding multiple image blocks one by one. For example, encoder 112 can perform the encoding process for each image block separately to complete the encoding of the video frame to be encoded.
[0053] It should be noted that the downsampling network and downsampling network model in the embodiments of this application can be described interchangeably.
[0054] Figure 3 The diagram shown is a possible structural schematic of a decoder 122 used to implement the video encoding method of the embodiments of this application.
[0055] Decoder 122 can be used to receive, for example, image data encoded by encoder 112 (i.e., an encoded bitstream, for example, an encoded bitstream including image blocks and associated syntax elements) to obtain decoded image blocks.
[0056] like Figure 3 As shown, decoder 122 may include a bitstream parsing unit 301, an inverse quantization unit 302, an inverse transform unit 303, a prediction processing unit 304, a reconstruction unit 305, and a filter unit 306. In some instances, decoder 122 may perform a decoding process that is substantially the inverse of the encoding process described in encoder 112. Optionally, decoder 122 may also include a buffer and a filtered image buffer. The buffer can be used to buffer the reconstructed image blocks output by reconstruction unit 209, and the filtered image buffer can be used to buffer the filtered image blocks output by filter unit 306.
[0057] The bitstream parsing unit 301 can be used to decode the encoded bitstream to obtain quantized residual coefficients (or quantized residual values) and / or decoding parameters (e.g., decoding parameters may include any one or more of inter-frame prediction parameters, intra-frame prediction parameters, filter parameters, and / or other syntax elements performed on the encoding side). The inverse quantization unit 302 functions similarly to the inverse quantization unit 207 of the encoder 112, used to inverse quantize the quantized residual coefficients decoded by the bitstream parsing unit 301. The inverse residual transform unit 303 functions similarly to the inverse transform unit 208 of the encoder 112, used to perform an inverse transform on the aforementioned inverse quantized residual coefficients to obtain the reconstructed residual values. The block obtained after the inverse transform is the residual block of the reconstructed block to be decoded in the pixel domain. The reconstruction unit 305 (e.g., a summer) functions similarly to the reconstruction unit 209 of the encoder 112. The prediction processing unit 303 is used to receive or acquire encoded image data (e.g., the encoded bitstream of the current image block) and reconstructed image data. The prediction processing unit 303 can also receive or acquire relevant parameters of the prediction mode and / or information about the selected prediction mode from, for example, the bitstream parsing unit 302, and predict the current image block based on the relevant data and decoding parameters in the reconstructed image data to obtain a prediction block for the current image block. The reconstruction unit 305 can be used to add the reconstructed residual block to the prediction block to obtain a reconstructed block of the image to be decoded in the sample domain, for example, by adding the residual value in the reconstructed residual block to the predicted value in the prediction block. The filter unit 306 can be used to filter the reconstructed block to obtain a filtered block, which is the decoded image block.
[0058] It is understood that in the encoder 112 and decoder 122 provided in the embodiments of this application, the processing result of a certain stage may be further processed before being output to the next stage. For example, after the prediction, transformation or filtering stages, the processing result of the corresponding stage may be further processed by Clip or shift operations.
[0059] The execution entity of the video encoding method provided in this application embodiment can be Figure 1 The video encoding device in the video encoding device, or the device in the video encoding device used to perform video encoding-related functions, or the... Figure 3 The encoder in, or for Figure 3 The encoder in the document is a device used to perform video encoding-related functions. The following uses the video encoding device to perform the video encoding method as an example to explain the video encoding method provided in the embodiments of this application.
[0060] The video encoding provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings, through some examples and application scenarios.
[0061] Figure 4 This is a flowchart illustrating the video encoding method provided in an embodiment of this application, as shown below. Figure 4 As shown, the video encoding method may include the following steps 401 and 402:
[0062] Step 401: The video encoding device inputs the first video frame into the first downsampling network model and outputs at least one image block.
[0063] The first downsampling network model mentioned above is trained based on the super-resolution network model, and this first downsampling network model is used to reduce the resolution of video frames.
[0064] In some embodiments of this application, the first video frame described above may be a video frame in a video frame sequence.
[0065] In some embodiments of this application, the first video frame described above may be a video frame with a first resolution.
[0066] In some embodiments of this application, the first resolution may be 480p or higher, or the first resolution may be 720p or higher.
[0067] For example, the first video frame may be a video frame with a resolution of 480p, or a video frame with a resolution of 720p, or a video frame with a resolution of 1080p, etc.
[0068] It should be noted that the first video frame can be any high-resolution video frame. The above is only an example of the resolution of the first video frame. This application does not limit the resolution of the first video frame in any way.
[0069] In some embodiments of this application, the first video frame may be the original video frame; or the first video frame may be a video frame obtained by resampling the original video frame, i.e., the resampling result of the original video frame.
[0070] It should be noted that a raw video frame refers to video data captured directly from a video source (such as a camera or recorder) without any processing or modification. A raw video frame contains all the original information of the video, such as color, brightness, and contrast, as well as motion information and scene details. Raw video frames are typically captured and stored at a certain frame rate. The resampling result of a raw video frame refers to the video data obtained after resampling the raw video frame. Furthermore, resampling can include, but is not limited to, conversions in resolution, sampling rate, and color space. For example, the above resampling can involve downsampling the raw video frame to generate a lower-resolution video frame.
[0071] In some embodiments of this application, the first downsampling network model described above can be a network model that converts high-resolution images into low-resolution images, which can be constructed based on convolutional neural networks (CNNs) or other types of neural networks.
[0072] It should be noted that the high resolution in the embodiments of this application can be 480p or higher, or 720p or higher, and the low resolution in the embodiments of this application can be lower than 480p or lower than 720p.
[0073] In some embodiments of this application, the aforementioned at least one image block may be a Coding Tree Unit (CTU) or a Coding Unit (CU).
[0074] It should be noted that the image block can be an image unit of video coding, such as a CTU, CU, or a prediction unit (PU). The type of at least one image block output by the first downsampling network model can be determined according to the actual coding process or coding requirements, and this application embodiment does not limit it.
[0075] It should be noted that the first image block mentioned above can be a low-resolution reconstructed image block, i.e., a REC image block.
[0076] In this embodiment, by using a training method that better fits the coding environment, the super-resolution network model is integrated into the training process of the downsampling network, which improves the coupling between the downsampling neural network model and the super-resolution network model. In the actual coding process, the low-resolution video frames generated by the downsampling network contain more high-frequency components, making it easier for them to be restored into high-quality high-resolution video frames by the super-resolution network model during the coding process, thereby reducing the residual, reducing the bit rate, and improving the coding gain.
[0077] In some embodiments of this application, the super-resolution network model described above converts low-resolution images into high-resolution images.
[0078] It should be noted that the above-mentioned super-resolution network model can also be called a super-resolution network model or a super-resolution network, and the three terms can be used interchangeably in the embodiments of this application.
[0079] In some embodiments of this application, a downsampling network model is jointly trained based on a super-resolution network model to optimize parameters such as weights and biases of the downsampling network model. This enables the trained downsampling network model (i.e., the first downsampling network model) to output high-quality low-resolution video frames that are more easily upsampled by the super-resolution network model in the encoding environment.
[0080] Step 402: The video encoding device generates a first residual block based on the first image block and the first prediction block.
[0081] Wherein, the first image block is an image block among the at least one image block, and the first prediction block is obtained by predicting the first image block.
[0082] In some embodiments of this application, the video encoding apparatus can predict a first image block using an intra-frame prediction unit to generate a corresponding first prediction block, and calculate the difference between the first image block and the first prediction block generated by the intra-frame prediction unit using a residual calculation unit to obtain a first residual block.
[0083] In some embodiments of this application, after step 402 above, the following steps A1 and A2 may also be included:
[0084] Step A1: The video encoding device transforms and quantizes the first residual block to obtain quantization coefficients.
[0085] Step A2: The video encoding device writes the above quantization coefficients into the bitstream.
[0086] In some embodiments of this application, the video encoding device performs transform encoding on the first residual block through a transform / quantization unit to obtain transform coefficients, and performs quantization processing on the transform coefficients to map the transform coefficients to a finite number of discrete values to obtain quantization coefficients. Then, the quantization coefficients are further encoded through an entropy encoding unit, such as using entropy encoding (e.g., Huffman coding, arithmetic coding, or CABAC) to reduce redundant information in the data and output the bitstream.
[0087] It should be noted that the embodiments of this application only describe the processing of the first image block in at least one image block output by the first downsampling network model, and the processing of any image block in at least one image block is the same as that of the first image block.
[0088] The video coding method provided in this application provides an embodiment that inputs a first video frame into a first downsampling network model and outputs at least one first image block. The first downsampling network model is trained based on a super-resolution network model and is used to downsample the video frame. Based on the at least one image block and at least one prediction block, at least one residual block is generated. The at least one prediction block is obtained by predicting the at least one first image block. This method reduces the resolution of the video data by downsampling the first video frame using the first downsampling network model. Since the first downsampling network is jointly trained with the super-resolution network model, it has good coupling with the super-resolution network model, enabling the first downsampling network model to output image blocks that retain more high-frequency information that is easier for the super-resolution network model to recover, thereby improving the overall video coding quality.
[0089] In some embodiments of this application, after step 402 above, the video encoding method provided in this application may further include steps 403 and 404:
[0090] Step 403: The video encoding device generates a first reconstruction block based on the first residual block and the first prediction block.
[0091] Step 404: The video encoding device inputs the first reconstruction block into the super-resolution network model and outputs the second reconstruction block.
[0092] The first reconstruction block and the second reconstruction block are reconstruction blocks of the first image block with different resolutions, and the resolution of the second reconstruction block is greater than that of the first reconstruction block.
[0093] In some embodiments of this application, after the video encoding device generates the quantization coefficients corresponding to the first residual block through the transform / quantization unit, it performs inverse quantization on the quantization coefficients during reconstruction by the inverse quantization / inverse transform unit to recover the transform coefficients and performs inverse transform on the transform coefficients to recover the first residual block. Then, the reconstruction unit adds the first prediction block generated by the intra-frame prediction unit to the first residual block to reconstruct the first image block, i.e., the first reconstructed block.
[0094] In some embodiments of this application, by inputting the first reconstruction block into the super-resolution model, a reconstruction block with higher resolution, namely the second reconstruction block, can be obtained.
[0095] It is understandable that the first reconstructed block obtained after reconstruction is a low-resolution image block, so it needs to be further converted into a high-resolution image block in order to reconstruct the original video frame.
[0096] In this embodiment, by using a downsampling network model oriented towards super-resolution network models, richer high-frequency information can be better preserved, making it easier for the super-resolution network model to restore low-resolution video frames into high-quality high-resolution frames, thereby improving the quality of reconstructed video frames.
[0097] This application provides a training method, which may include the following steps 405 to 408:
[0098] Step 405: The video encoding device inputs video frame samples into the second downsampling network model and outputs at least one second image block.
[0099] Step 406: The video encoding device inputs at least one second image block into the super-resolution network model and outputs at least one third image block.
[0100] Step 407: The video encoding device calculates the loss value based on at least one second image block, at least one third image block, and video frame samples.
[0101] Step 408: The video encoding device trains the second downsampling network model based on the loss value to obtain the first downsampling network model.
[0102] In some embodiments of this application, the video frame sample described above may include one or more video frames.
[0103] In some embodiments of this application, the video frame sample is a video frame with a first resolution; or the video frame sample is a video frame with a second resolution obtained by downsampling a video frame with a first resolution, wherein the first resolution is greater than the second resolution.
[0104] In some embodiments of this application, the second resolution may be a resolution lower than 480p or lower than 720p.
[0105] It should be noted that the explanation of the first resolution can be found in the relevant description of the above embodiments, and will not be repeated here.
[0106] In some embodiments of this application, the aforementioned video frame samples may include multiple high-resolution video frames. Thus, by using a large number of high-resolution video sequence frames as a dataset, the generalization of the downsampling network model can be improved.
[0107] Understandably, the video frame samples are part of the training dataset, which contains a series of video frames, which can be the original high-resolution video frames or video frames that have been resampled from the original high-resolution video frames.
[0108] In some embodiments of this application, one video frame in a video frame sample may correspond to at least one second image block.
[0109] In some embodiments of this application, the parameters of the super-resolution network model are preset, including but not limited to basic quantization parameters, piecewise quantization parameters, and sampling factor. These parameters can be parameter information extracted from the training dataset.
[0110] In some embodiments of this application, video frame samples are input into a second downsampling network model. The second downsampling network model processes the video frame samples and outputs lower-resolution image patches. Then, a super-resolution network model processes one or more lower-resolution image patches output by the second downsampling network model to generate higher-resolution image patches. Then, based on the differences between the lower-resolution image patches output by the second downsampling network model and their corresponding video frame samples, and the differences between the higher-resolution image patches output by the super-resolution network model and their corresponding video frame samples, a loss value is calculated. Finally, the backpropagation algorithm is used to calculate the gradients of each parameter of the second downsampling network model, and these gradients are used to update the parameter values of the downsampling network model. This process is repeated until a predetermined convergence criterion is met. After training, the optimized parameters are saved for use in the subsequent inference stage.
[0111] It should be noted that the downsampling network model proposed in this application can be applied to different super-resolution network models. During training, the parameters of the super-resolution network model can be frozen to better optimize the downsampling network.
[0112] In this embodiment, a novel image downsampling method for super-resolution models is designed by combining the training processes of a deep learning-based super-resolution network and a downsampling network. During the training of the downsampling network, joint training with the super-resolution network is introduced. Specifically, during training, the output of the downsampling network serves as the input to the super-resolution network. The parameters of the super-resolution network are fixed, and its output and the true value are used to calculate a loss function, thereby updating the parameters of the downsampling network and training a more efficient downsampling network model.
[0113] In some embodiments of this application, the video encoding method provided in this application may further include steps 405 to 408 as described above.
[0114] The encoding method provided in this application can be applied to the RPR framework in codecs. Under the RPR framework, high-resolution video sequences are downsampled to obtain low-resolution video frames, which are then encoded and transmitted to the decoder. This method uses an image downsampling network model oriented towards super-resolution models. Through a training method that better fits the encoding environment, the super-resolution network is integrated into the training process of the downsampling network, improving the coupling between the neural network-based downsampling method and the super-resolution method. This results in better preservation of richer high-frequency information and improved texture detail retention, making it easier for the super-resolution network to recover high-quality images during the encoding process, thus achieving superior encoding performance. Simultaneously, the super-resolution network model at the decoder end more easily restores low-resolution video frames to high-quality high-resolution frames, further improving encoding efficiency. This method can be used for video and image encoding, decoding, streaming, and storage implementations. Therefore, compared to traditional video encoding and decoding technologies, the encoding method provided in this application improves video encoding and decoding performance.
[0115] The encoding method provided in this application is universal and can be applied to the joint training and optimization of downsampling methods and super-resolution methods based on downsampling network models.
[0116] In some embodiments of this application, step 405 may include steps 405a to 405c:
[0117] Step 405a: The video encoding device inputs video frame samples into the second downsampling network model to extract color feature information of the video frame samples.
[0118] The aforementioned color feature information includes at least one luminance component and at least one set of chromaticity components.
[0119] Step 405b: The video encoding device downsamples the above-mentioned at least one set of chroma components to obtain at least one set of processed chroma components.
[0120] Step 405c: The video encoding device outputs at least one second image block based on at least one luminance component and at least one set of processed chrominance components.
[0121] In some embodiments of this application, after inputting video frame samples into a second downsampling network model, the color feature information of the video frame samples can be extracted through the convolution module of the second downsampling network model.
[0122] In some embodiments of this application, the above-described convolutional module may include a 3×3 convolutional layer.
[0123] In some embodiments of this application, the luminance component may be the Y component; the at least one set of chromaticity components may include the U component and the V component, or the at least one set of chromaticity components may include the Cr component and the Cb component.
[0124] It should be noted that in practical applications, the YUV format uses a "luminance" component called Y (equivalent to grayscale) and two "chrominance" components U (blue projection) and V (red projection) to represent color. Here, Y represents the luminance component, which is the grayscale value, U (Cb) represents the chrominance component of the blue part, and V (Cr) represents the chrominance component of the red part.
[0125] In some embodiments of this application, after inputting video frame samples into the second downsampling network model, the Y channel features and UV channel features of the video frame samples can be extracted by the convolution module to obtain the Y component and UV component.
[0126] It should be noted that the above UV components are the U component and the V component.
[0127] Furthermore, when the video frame sample is a video frame obtained by downsampling a high-resolution video frame by a factor of two, the video frame sample is input into the second downsampling network model. The UV channel of the video frame is then upsampled by a factor of two, and the UV channel is coupled with the Y channel. Then, the Y channel and UV channel features are extracted by the convolution module to obtain the Y component and UV component.
[0128] In some embodiments of this application, the video encoding apparatus can downsample the UV components of a video frame sample to obtain lower-resolution U and V components. Further, the video encoding apparatus can downsample the U and V components by a factor of N, where N is an integer greater than 1. For example, downsampling the U and V components by a factor of two yields downsampled U and V components.
[0129] In some embodiments of this application, the video encoding device can process the Y component and the obtained U and V components respectively through M residual modules, and then downsample the processed U and V components through a convolutional layer with a stride of 2 and merge them with the processed Y component to output at least one low-resolution image block, i.e. at least one second image block, where M is a positive integer.
[0130] In some embodiments of this application, the above-mentioned M residual modules may include 10 residual modules.
[0131] In some embodiments of this application, the residual module may consist of one or more convolutional layers and one or more separable convolutional layers, and the result of the convolution is added to the input of the residual block to form a residual connection. Using different numbers of residual blocks to process the Y channel and UV channel respectively can extract the features of each channel more effectively, optimize the characteristics of each channel on complex image content, and thus improve the performance of the overall network.
[0132] For example, the residual module can consist of three 1×1 convolutional layers and one 3×3 separable convolutional layer.
[0133] It should be noted that the aforementioned residual module can also be called a residual block. It is understood that the residual block in the downsampling network model has a different meaning than the residual block generated based on the prediction block.
[0134] Figure 5 This is a schematic diagram of a downsampling network model provided in some embodiments of this application, such as... Figure 5 As shown, the downsampling network model takes the downsampling result of the high-resolution original frame image as input, upsamples its UV channels by a factor of two, and couples them with the Y channel. Initial features are then extracted through a 3×3 convolutional layer. Subsequently, the initial features are processed separately for the Y and UV channels based on the number of channels. The Y and UV channels each pass through multiple residual blocks, as shown... Figure 6 As shown, the residual block consists of three 1×1 convolutional layers and one 3×3 separable convolutional layer. The result of the convolution is added to the input of the residual block to form a residual connection. Finally, the UV components are downsampled by a convolutional layer with a stride of 2 and then merged with the Y component to output a low-resolution result. Optionally, this downsampling network model can also include 1×1 convolutional layers connected to the 3×3 convolutional layers, as well as activation numbers, to improve the model's performance and generalization ability.
[0135] In this embodiment of the application, by extracting the color feature information of video frame samples and downsampling the chroma components, good visual quality and coding efficiency are maintained while reducing the amount of data.
[0136] In some embodiments of this application, step 407 may include steps 407a to 407c:
[0137] Step 407a: The video encoding device calculates a first loss value based on at least one second image block and video frame samples.
[0138] Step 407b: The video encoding device calculates a second loss value based on at least one third image block and video frame sample.
[0139] Step 407c: The video encoding device calculates the loss value based on the first loss value and the second loss value mentioned above.
[0140] In some embodiments of this application, the above-mentioned loss value is calculated based on a loss function, which is either the absolute difference (SAD) function or the mean squared error (MSE) function.
[0141] It should be noted that the Sum of Absolute Differences (SAD) function calculates the sum of the absolute errors between the predicted and actual values, and is used to measure the average absolute difference between the two; the Mean Squared Error (MSE) function calculates the average of the squares of the differences between the predicted and actual values, and is used to quantify the accuracy of the model's prediction results.
[0142] In some embodiments of this application, the video encoding apparatus can calculate a first loss value based on the difference between at least one second image block output by the second downsampling network model and the corresponding video frame sample, and calculate a second loss value based on the difference between at least one third image block output by the super-resolution network model and the corresponding video frame sample. Then, it can calculate a joint loss value, i.e., the aforementioned loss value. Exemplarily, the difference can be pixel value difference, structural difference, etc.
[0143] In some embodiments of this application, a first loss value can characterize the low-resolution image guidance loss, and a second loss value can characterize the reconstruction loss.
[0144] In some embodiments of this application, the low-resolution image guidance loss refers to the loss calculated between the low-resolution image patch output by the downsampling network model and the downsampling result of the ground truth frame (or the original video frame), for example, the loss calculated between the low-resolution image patch output by the downsampling network model and the ground truth frame at twice the downsampling result; the reconstruction loss refers to the loss calculated between the output of the super-resolution network model and the high-resolution ground truth frame (or the original video frame).
[0145] It should be noted that the joint loss value can be a weighted sum of multiple individual loss values, used to comprehensively measure multiple different types of loss during training.
[0146] In some embodiments of this application, the joint loss value combines a first loss value and a second loss value to comprehensively measure the low-resolution image guidance loss and reconstruction loss during the training process. By calculating the joint loss value, a comprehensive metric for evaluating model performance can be obtained. During training, the parameters of the downsampling network model can be optimized by minimizing the joint loss value, thereby improving the performance of the downsampling network model. Furthermore, since the joint loss value combines the low-resolution image guidance loss and reconstruction loss, it can better optimize the overall performance of the model.
[0147] In some embodiments of this application, the process of step 407a described above may include steps 407a1 and 407a2:
[0148] Step 407a1: The video encoding device downsamples the video frame samples to obtain processed video frame samples.
[0149] Step 407a2: The video encoding device calculates a first loss value based on at least one second image block and processed video frame samples.
[0150] The first loss value is used to characterize the difference between the resolution corresponding to the at least one second image block and the resolution corresponding to the downsampled video frame sample.
[0151] In some embodiments of this application, the video encoding apparatus can downsample the UV components of a video frame sample to obtain a processed video frame sample.
[0152] In some embodiments of this application, the video encoding apparatus can calculate the difference information between at least one second image block and a downsampled video frame sample, and calculate a first loss value based on the difference information.
[0153] In some embodiments of this application, the aforementioned second loss value is used to characterize the difference between at least one third image block and the corresponding second video frame.
[0154] In some embodiments of this application, step 407c above can be implemented by step 407c1:
[0155] Step 407c1: The video encoding device linearly combines the first loss value and the second loss value to obtain the loss value.
[0156] It should be noted that linearly combining the first and second loss values means summing them according to certain weights to obtain a combined value. In other words, linear combination is used to combine the first and second loss values into a joint loss value so that both losses can be optimized simultaneously during training.
[0157] In some embodiments of this application, the video coding apparatus can perform a weighted summation of the first loss value and the second loss value according to certain weights to obtain a comprehensive loss value (i.e., joint loss value), so as to simultaneously optimize the low-resolution image guidance loss and reconstruction loss during the training process.
[0158] For example, assuming the first loss value is L1, the second loss value is L2, and the joint loss value L is a linear combination of L1 and L2, the formula for calculating the joint loss value L can be: L=αL1+βL2, where α and β are weighting coefficients.
[0159] It should be noted that the weight coefficients α and β can be adjusted according to task requirements or dataset characteristics to balance the impact of different loss values on model training. This application does not limit this.
[0160] In this embodiment, a comprehensive loss value is obtained by linearly combining the first loss value and the second loss value, thereby better guiding the optimization direction of the model during training. In addition, using multiple loss values for training can enhance the generalization ability of the model, enabling it to better adapt to different application scenarios and conditions, thereby more comprehensively evaluating and optimizing the performance of the model in video coding.
[0161] Figure 7 This is a framework diagram of downsampling joint training provided in the embodiments of this application, which is described below in conjunction with... Figure 7 The training method provided in the embodiments of this application will be described by way of example.
[0162] First, a large number of high-resolution video sequence frames are used as the dataset to improve the model's generalization ability. These original video frames are then input into a deep learning-based downsampling network model, which outputs low-resolution reconstructed image patches. These reconstructed image patches are then divided into YUV channels, which are used as inputs to the super-resolution network model. Other side information of the super-resolution network model (such as basic quantization parameters, piecewise quantization parameters, and sampling factors) is extracted from the training dataset and input as corresponding values. Next, the obtained features are used to calculate the joint loss, which includes low-resolution image guidance loss and reconstruction loss. The final joint loss is a linear combination of the guidance loss and reconstruction loss. Then, the backpropagation algorithm is used to calculate the gradients of each parameter of the downsampling network model, and these gradients are used to update the parameter values of the downsampling network model. This process is repeated until a predetermined convergence criterion is met, i.e., the loss function tends to stabilize. Finally, after training is complete, the optimized parameters are saved for use in the subsequent inference stage.
[0163] In this embodiment, the downsampling network model is trained by jointly training the super-resolution network model. This makes the training process of the downsampling network model more consistent with the actual coding environment and strengthens the connection between the downsampling network model and the super-resolution network model. During actual coding, the low-resolution video frames generated by the downsampling network contain more high-frequency components, which can be better restored into high-quality, high-resolution frames by the super-resolution network model. This reduces residuals, thereby lowering the bitrate and increasing coding gain.
[0164] The decoding method provided in this application embodiment can be applied to a video decoding device, and the decoding method may include the following steps 501 to 503:
[0165] Step 501: The video decoding device parses the bitstream and obtains the quantization coefficients.
[0166] The aforementioned quantization coefficients are obtained by transforming and quantizing the first residual block, which is generated based on the first image block.
[0167] Step 502: The video decoding device generates a first reconstructed block based on the quantization coefficients and the first prediction block.
[0168] The first prediction block is obtained by predicting the first image block.
[0169] Step 503: The video decoding device inputs the first reconstruction block into the super-resolution network model and outputs the second reconstruction block.
[0170] It should be noted that the explanation of this embodiment can be found in the relevant description of the encoding side method above, and will not be repeated here.
[0171] In this embodiment, the image downsampling network model used for super-resolution models integrates the super-resolution network model into the training process of the downsampling network through a training method that better fits the encoding environment. This improves the coupling between the neural network-based downsampling method and the super-resolution method, enabling better preservation of richer high-frequency information and enhancing the retention of texture details. This makes it easier for the super-resolution network to reconstruct high-quality images during the encoding process, resulting in superior encoding performance. Simultaneously, the super-resolution network model at the decoding end more easily restores low-resolution video frames to high-quality high-resolution frames, further improving encoding and decoding performance.
[0172] The video encoding method provided in this application can be executed by a video encoding device. This application uses a video encoding device executing the video encoding method as an example to illustrate the video encoding device provided in this application.
[0173] Accordingly, this application provides a video encoding device. Based on the above method example, the video encoding device can be divided into functional modules. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated modules can be implemented in hardware or as software functional modules. It should be noted that the module division in this application is illustrative and represents only one logical functional division; in actual implementation, other division methods may be used.
[0174] The video encoding device includes a processing module. This processing module can be implemented in software or hardware. When implemented in hardware, the processing module can be implemented by a processor. For example, the processor can include a general-purpose processor, a special-purpose processor, such as a Central Processing Unit (CPU), a microprocessor, a Digital Signal Processor (DSP), an Artificial Intelligence (AI) processor, a Graphics Processing Unit (GPU), an Application Specific Integrated Circuit (ASIC), a Network Processor (NP), a Field Programmable Gate Array (FPGA), or other programmable logic devices, gate circuits, transistors, discrete hardware components, etc. Optionally, the video encoding device may also include a receiving module and a transmitting module. The receiving module and the transmitting module can be implemented by a communication interface, which can include one or more of the following: a transceiver, pins, circuits, a bus, and a radio frequency unit.
[0175] When dividing each function into modules according to its corresponding function. Figure 8 A possible structural schematic diagram of the video encoding device 700 involved in the above embodiments is shown. For example... Figure 8 As shown, the video encoding apparatus 700 includes a processing module 701. The processing module 701 is configured to: input a first video frame into a first downsampling network model and output at least one image block; the first downsampling network model is trained based on a super-resolution network model and is used to reduce the resolution of the video frame; the processing module 701 is further configured to: generate a first residual block based on the first image block and a first prediction block; the first image block is an image block from at least one image block, and the first prediction block is obtained by predicting the first image block.
[0176] In some embodiments of this application, the processing module is further configured to generate a first reconstruction block based on the first residual block and the first prediction block after generating a first residual block based on the first image block and the first prediction block; the processing module is further configured to input the first reconstruction block into the super-resolution network model and output a second reconstruction block; wherein the first reconstruction block and the second reconstruction block are reconstruction blocks of different resolutions of the first image block, and the resolution of the second reconstruction block is greater than the resolution of the first reconstruction block.
[0177] In some embodiments of this application, the above-described processing module is further configured to: input video frame samples into a second downsampling network model and output at least one second image block; input the at least one second image block into a super-resolution network model and output at least one third image block; calculate a loss value based on the at least one second image block, the at least one third image block, and the video frame samples; and train the second downsampling network model based on the loss value to obtain the first downsampling network model.
[0178] In some embodiments of this application, the above-mentioned processing module is specifically used for: inputting video frame samples into a second downsampling network model, extracting color feature information of the video frame samples, wherein the color feature information includes at least one luminance component and at least one set of chrominance components; performing downsampling processing on the at least one set of chrominance components to obtain at least one set of processed chrominance components; and outputting at least one second image block based on the at least one luminance component and the at least one set of processed chrominance components.
[0179] In some embodiments of this application, the processing module is specifically used to: calculate a first loss value based on the at least one second image block and the video frame sample; calculate a second loss value based on the at least one third image block and the video frame sample; and calculate the loss value based on the first loss value and the second loss value.
[0180] In some embodiments of this application, the processing module is specifically used to: perform downsampling processing on the video frame samples to obtain processed video frame samples; calculate the first loss value based on the at least one second image block and the processed video frame samples; wherein the first loss value is used to characterize the difference between the resolution corresponding to the at least one second image block and the resolution corresponding to the downsampled video frame samples.
[0181] In some embodiments of this application, the above-mentioned processing module is specifically used to: linearly combine the above-mentioned first loss value and the above-mentioned second loss value to obtain the above-mentioned loss value.
[0182] In some embodiments of this application, the above-mentioned loss value is calculated based on a loss function, which is either the absolute difference (SAD) function or the mean squared error (MSE) function.
[0183] In some embodiments of this application, the video frame sample is a video frame with a first resolution; or the video frame sample is a video frame with a second resolution obtained by downsampling the video frame with the first resolution, wherein the first resolution is greater than the second resolution.
[0184] The video encoding apparatus provided in this application embodiment inputs a first video frame into a first downsampling network model and outputs at least one first image block. The first downsampling network model is trained based on a super-resolution network model and is used to downsample the video frame. Based on the at least one image block and at least one prediction block, at least one residual block is generated. The at least one prediction block is obtained by predicting the at least one first image block. This method reduces the resolution of the video data by downsampling the first video frame using the first downsampling network model. Since the first downsampling network is jointly trained with the super-resolution network model, it has good coupling with the super-resolution network model, enabling the first downsampling network model to output image blocks that retain more high-frequency information that is easier for the super-resolution network model to recover, thereby improving the overall video encoding quality.
[0185] The aforementioned video encoding device can be the aforementioned Figure 1 or Figure 2 The encoder mentioned above, or the encoder described above. Figure 1 In medium video encoding equipment, the device used to perform video encoding-related functions, or for Figure 3 The encoder 112 in the middle is a device for performing video encoding-related functions; or is a Figure 4 A device used to perform video encoding-related functions.
[0186] When using integrated units, Figure 9 A schematic diagram of another possible structure of the video encoding apparatus involved in the above embodiments is shown. For example... Figure 9 As shown, the video encoding apparatus provided in this application embodiment may include a processing unit 801, a communication unit 802, and a storage unit 803. The processing unit 801 can be used to control and manage the operation of the video encoding apparatus. For example, the processing unit 801 can be used to support the video encoding apparatus in executing steps 401 and 402 in the above method embodiments, and / or other processes used in the technology described herein. The communication unit 802 can be used to support communication between the video encoding apparatus and other network entities, such as communication with the reconstruction unit 208. The storage unit 803 is used to store the program code and data of the video encoding apparatus, such as storing encoded video images or bitstreams.
[0187] The processing unit 801 can be a processor, for example, a processor can be... Figure 3 The encoder 122 is included. The communication unit 802 can be a transceiver, transceiver circuit, or communication interface, for example... Figure 3 The communication interface 121 and the storage unit 803 can be a memory.
[0188] For more details on how the modules included in the video encoding device implement the above functions, please refer to the descriptions in the preceding method embodiments, which will not be repeated here.
[0189] Each module of the video encoding device described above can also be used to perform other actions in the above method embodiments. All relevant content of each step involved in the above method embodiments can be referred to the functional description of the corresponding functional module, and will not be repeated here.
[0190] Figure 9 The video encoding device shown can be the one described above. Figure 1 In medium video encoding equipment, the device used to perform video encoding-related functions, or for Figure 3 The encoder 112 is a device used to perform video encoding-related functions.
[0191] The video encoding apparatus provided in this application embodiment can implement all the processes implemented in the above-described video encoding method embodiment and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0192] This application embodiment also provides a terminal, which can be a terminal device, including a processor and a communication interface. The communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the steps in the above-described video encoding method embodiment. This terminal embodiment corresponds to the above-described terminal-side method embodiment, and all implementation processes and methods of the above method embodiments can be applied to this terminal embodiment and achieve the same technical effect. The terminal can be... Figure 8 The video encoding device shown. Specifically, Figure 10 A schematic diagram of the hardware structure of a terminal to implement an embodiment of this application.
[0193] The terminal 100 includes, but is not limited to, at least some of the following components: radio frequency unit 101, network module 102, audio output unit 103, input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, and processor 110.
[0194] Those skilled in the art will understand that the terminal 100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 10 The terminal structure shown does not constitute a limitation on the terminal. The terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0195] It should be understood that, in this embodiment, the input unit 104 may include a graphics processor 1041 and a microphone 1042. The graphics processor 1041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, joysticks, etc., which will not be described in detail here.
[0196] In this embodiment, after receiving downlink data from the network-side device, the radio frequency unit 101 can transmit it to the processor 110 for processing; in addition, the radio frequency unit 101 can send uplink data to the network-side device. Typically, the radio frequency unit 101 includes, but is not limited to, antennas, amplifiers, transceivers, couplers, low-noise amplifiers, duplexers, etc.
[0197] The memory 109 can be used to store software programs or instructions, as well as various data. The memory 109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 109 may include volatile memory or non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.
[0198] Processor 110 may include one or more processing units; optionally, processor 110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 110.
[0199] The processor 110 is configured to input a first video frame into a first downsampling network model and output at least one image block. The first downsampling network model is trained based on a super-resolution network model and is used to reduce the resolution of the video frame. The processor 110 is also configured to generate a first residual block based on the first image block and a first prediction block. The first image block is an image block among at least one image block, and the first prediction block is obtained by predicting the first image block.
[0200] In some embodiments of this application, the processor 110 is further configured to generate a first reconstruction block based on the first residual block and the first prediction block after generating a first residual block based on the first image block and the first prediction block; the processor 110 is further configured to input the first reconstruction block into the super-resolution network model and output a second reconstruction block; wherein the first reconstruction block and the second reconstruction block are reconstruction blocks of different resolutions of the first image block, and the resolution of the second reconstruction block is greater than the resolution of the first reconstruction block.
[0201] In some embodiments of this application, the processor 110 is further configured to: input video frame samples into a second downsampling network model and output at least one second image block; input the at least one second image block into a super-resolution network model and output at least one third image block; calculate a loss value based on the at least one second image block, the at least one third image block, and the video frame samples; and train the second downsampling network model based on the loss value to obtain the first downsampling network model.
[0202] In some embodiments of this application, the processor 110 is specifically configured to: input video frame samples into a second downsampling network model, extract color feature information of the video frame samples, the color feature information including at least one luminance component and at least one set of chrominance components; perform downsampling processing on the at least one set of chrominance components to obtain at least one set of processed chrominance components; and output the at least one second image block based on the at least one luminance component and the at least one set of processed chrominance components.
[0203] In some embodiments of this application, the processor 110 is specifically configured to: calculate a first loss value based on the at least one second image block and the video frame sample; calculate a second loss value based on the at least one third image block and the video frame sample; and calculate the loss value based on the first loss value and the second loss value.
[0204] In some embodiments of this application, the processor 110 is specifically configured to: perform downsampling processing on the video frame samples to obtain processed video frame samples; calculate the first loss value based on the at least one second image block and the processed video frame samples; wherein the first loss value is used to characterize the difference between the resolution corresponding to the at least one second image block and the resolution corresponding to the downsampled video frame samples.
[0205] In some embodiments of this application, the processor 110 is specifically used to: linearly combine the first loss value and the second loss value to obtain the loss value.
[0206] In some embodiments of this application, the above-mentioned loss value is calculated based on a loss function, which is either the absolute difference (SAD) function or the mean squared error (MSE) function.
[0207] In some embodiments of this application, the video frame sample is a video frame with a first resolution; or the video frame sample is a video frame with a second resolution obtained by downsampling the video frame with the first resolution, wherein the first resolution is greater than the second resolution.
[0208] The terminal provided in this application embodiment inputs a first video frame into a first downsampling network model and outputs at least one first image block. The first downsampling network model is trained based on a super-resolution network model and is used to downsample the video frame. Based on the at least one image block and at least one prediction block, at least one residual block is generated. The at least one prediction block is obtained by predicting the at least one first image block. Through this method, the terminal downsamples the first video frame using the first downsampling network model to reduce the resolution of the video data. Since the first downsampling network is jointly trained based on the super-resolution network model, it has good coupling with the super-resolution network model, enabling the first downsampling network model to output image blocks that retain more high-frequency information that is beneficial for the super-resolution network model to recover, thereby improving the overall video coding quality.
[0209] It is understood that the implementation process of each implementation method mentioned in this embodiment can refer to the relevant description of method embodiment 111 and achieve the same or corresponding technical effects. To avoid repetition, it will not be described again here.
[0210] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described video encoding method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0211] The processor mentioned above is either the processor in the terminal described in the above embodiments or the processor in the network-side device. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk. In some examples, the readable storage medium may be a non-transient readable storage medium.
[0212] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described video encoding method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0213] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0214] This application also provides a computer program / program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described video encoding method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0215] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0216] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When these computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center integrating one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital video discs (DVDs)), or semiconductor media (e.g., solid-state drives (SSDs)).
[0217] From the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of a computer software product plus the necessary general-purpose hardware platform, and of course, they can also be implemented by hardware. The computer software product is stored in a storage medium (such as ROM, RAM, magnetic disk, optical disk, etc.), and the computer software product includes several instructions to cause the video encoding device to execute the methods described in the various embodiments of this application.
[0218] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.
[0219] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other implementations under the guidance of this application without departing from the spirit and scope of the claims. All of these implementations are within the protection scope of this application.
Claims
1. A video encoding method, characterized in that, The method includes: The first video frame is input into the first downsampling network model, and at least one first image patch is output. The first downsampling network model is trained based on the super-resolution network model and is used to reduce the resolution of the video frame. A first residual block is generated based on a first image block and a first prediction block. The first image block is an image block among the at least one image block, and the first prediction block is obtained by predicting the first image block.
2. The method according to claim 1, characterized in that, After generating the first residual block based on the first image block and the first prediction block, the method further includes: A first reconstruction block is generated based on the first residual block and the first prediction block; The first reconstruction block is input into the super-resolution network model, and the second reconstruction block is output. Wherein, the first reconstruction block and the second reconstruction block are reconstruction blocks of the first image block at different resolutions, and the resolution of the second reconstruction block is greater than that of the first reconstruction block.
3. The method according to claim 1 or 2, characterized in that, The method further includes: Input video frame samples into the second downsampling network model and output at least one second image patch; The at least one second image block is input into the super-resolution network model, and at least one third image block is output. Calculate the loss value based on the at least one second image block, the at least one third image block, and the video frame samples; Based on the loss value, the second downsampling network model is trained to obtain the first downsampling network model.
4. The method according to claim 3, characterized in that, The step of inputting video frame samples into the second downsampling network model and outputting at least one second image patch includes: The video frame samples are input into the second downsampling network model to extract the color feature information of the video frame samples. The color feature information includes at least one luminance component and at least one set of chrominance components. The at least one set of chromaticity components is downsampled to obtain at least one set of chromaticity components after processing. Based on the at least one luminance component and the at least one set of processed chrominance components, the at least one second image block is output.
5. The method according to claim 3, characterized in that, The calculation of the loss value based on the at least one second image block, the at least one third image block, and the video frame samples includes: Calculate the first loss value based on the at least one second image patch and the video frame sample; The second loss value is calculated based on the at least one third image patch and the video frame sample; The loss value is calculated based on the first loss value and the second loss value.
6. The method according to claim 3, characterized in that, The calculation of the first loss value based on the at least one second image patch and the video frame samples includes: The video frame samples are downsampled to obtain processed video frame samples; The first loss value is calculated based on the at least one second image block and the processed video frame sample; The first loss value is used to characterize the difference between the resolution corresponding to the at least one second image patch and the resolution corresponding to the downsampled video frame sample.
7. The method according to claim 5, characterized in that, The step of calculating the loss value based on the first loss value and the second loss value includes: The loss value is obtained by linearly combining the first loss value and the second loss value.
8. The method according to any one of claims 3 to 7, characterized in that, The loss value is calculated based on a loss function, which is either the absolute difference (SAD) function or the mean squared error (MSE) function.
9. The method according to claim 3, characterized in that, The video frame sample is a video frame with a first resolution; or the video frame sample is a video frame with a second resolution obtained by downsampling a video frame with a first resolution, wherein the first resolution is greater than the second resolution.
10. A video encoding device, characterized in that, The device includes: a processing module; The processing module is used to input the first video frame into the first downsampling network model and output at least one image block. The first downsampling network model is trained based on the super-resolution network model and is used to reduce the resolution of the video frame. The processing module is further configured to generate a first residual block based on the first image block and the first prediction block, wherein the first image block is an image block among the at least one image block, and the first prediction block is obtained by predicting the first image block.
11. The apparatus according to claim 10, characterized in that, The processing module is further configured to generate a first reconstruction block based on the first residual block and the first prediction block after generating a first residual block based on the first image block and the first prediction block; The processing module is further configured to input the first reconstruction block into the super-resolution network model and output the second reconstruction block; Wherein, the first reconstruction block and the second reconstruction block are reconstruction blocks of the first image block at different resolutions, and the resolution of the second reconstruction block is greater than that of the first reconstruction block.
12. The apparatus according to claim 10 or 11, characterized in that, The processing module is further configured to: Input video frame samples into the second downsampling network model and output at least one second image patch; The at least one second image block is input into the super-resolution network model, and at least one third image block is output. Calculate the loss value based on the at least one second image block, the at least one third image block, and the video frame samples; Based on the loss value, the second downsampling network model is trained to obtain the first downsampling network model.
13. The apparatus according to claim 12, characterized in that, The processing module is specifically used for: The video frame samples are input into the second downsampling network model to extract the color feature information of the video frame samples. The color feature information includes at least one luminance component and at least one set of chrominance components. The at least one set of chromaticity components is downsampled to obtain at least one set of chromaticity components after processing. Based on the at least one luminance component and the at least one set of processed chrominance components, the at least one second image block is output.
14. The apparatus according to claim 12, characterized in that, The processing module is specifically used for: Calculate the first loss value based on the at least one second image patch and the video frame sample; The second loss value is calculated based on the at least one third image patch and the video frame sample; The loss value is calculated based on the first loss value and the second loss value.
15. The apparatus according to claim 12, characterized in that, The processing module is specifically used for: The video frame samples are downsampled to obtain processed video frame samples; The first loss value is calculated based on the at least one second image block and the processed video frame sample; The first loss value is used to characterize the difference between the resolution corresponding to the at least one second image patch and the resolution corresponding to the downsampled video frame sample.
16. The apparatus according to claim 14, characterized in that, The processing module is specifically used for: The loss value is obtained by linearly combining the first loss value and the second loss value.
17. The apparatus according to any one of claims 12 to 16, characterized in that, The loss value is calculated based on a loss function, which is either the absolute difference (SAD) function or the mean squared error (MSE) function.
18. The apparatus according to claim 12, characterized in that, The video frame sample is a video frame with a first resolution; or the video frame sample is a video frame with a second resolution obtained by downsampling a video frame with a first resolution, wherein the first resolution is greater than the second resolution.
19. A terminal device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the video encoding method as described in any one of claims 1 to 9.
20. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the video encoding method as described in any one of claims 1 to 9.
21. A computer program product, characterized in that, It includes computer program instructions that cause a computer to perform the steps of the video encoding method as described in any one of claims 1 to 9.
22. A chip, characterized in that, The chip includes a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the video encoding method as described in any one of claims 1 to 9.