Video coding method, video decoding method, and non-transitory computer-readable storage medium

WO2026194655A1PCT designated stage Publication Date: 2026-09-24ALIBABA (CHINA) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/081296
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-17
Filing Date
2026-03-04
Publication Date
2026-09-24

Smart Images

  • Figure CN2026081296_24092026_PF_FP_ABST
    Figure CN2026081296_24092026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present disclosure are a video coding method, a video decoding method, and a non-transitory computer-readable storage medium. The video decoding method comprises: after a bitstream of a video is received and decoded to obtain a reconstructed video, calculating distribution parameters of each color channel of a key reference frame in the reconstructed video, wherein the distribution parameters are used for representing color features of corresponding color channels; for the same color channel, using the distribution parameters of the key reference frame to calibrate the color distribution of subsequent frames in the reconstructed video to obtain calibrated subsequent frames; and, on the basis of the key reference frame and the calibrated subsequent frames, generating a target video. The present disclosure solves the technical problem in the prior art of color shift of reconstructed videos, achieving the technical effect of improving the quality of the reconstructed videos.
Need to check novelty before this filing date? Find Prior Art

Description

Video encoding and decoding methods and non-transitory computer-readable storage media

[0001] This disclosure claims priority to Chinese Patent Application No. 202510317007.4, filed on March 17, 2025 with the China National Intellectual Property Administration, entitled “Video Encoding, Decoding Method and Non-Transitory Computer-Readable Storage Medium”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This disclosure relates to the field of video technology, and in particular to video encoding and decoding methods and non-transitory computer-readable storage media. Background Technology

[0003] Generative Video Compression (GVC) frameworks generally include an encoding-decoding process. At the encoder, key reference frames are compressed using traditional encoding techniques, while subsequent frames are represented using compact transmission symbols and encoded into the output bitstream. At the decoder, the decoded key reference frames and compacted facial information are input into a synthesis model to reconstruct the video. This approach enables ultra-low bitrate and high-quality video communication. However, current GVC frameworks commonly suffer from the technical problem of color shift during video reconstruction. No effective solution has yet been proposed to address this issue. Summary of the Invention

[0004] This disclosure provides video encoding and decoding methods, as well as non-transitory computer-readable storage media, to solve one or more of the aforementioned technical problems.

[0005] In a first aspect, embodiments of this disclosure provide a video decoding method applied to a generative video compression framework, comprising: receiving a bitstream of video and decoding it to obtain a reconstructed video; calculating the distribution parameters of each color channel in a key reference frame in the reconstructed video, wherein the distribution parameters are used to characterize the color features of the corresponding color channel; calibrating the color distribution of subsequent frames in the reconstructed video using the distribution parameters of the key reference frame for the same color channel, thereby obtaining calibrated subsequent frames; and generating a target video based on the key reference frame and the calibrated subsequent frames.

[0006] Secondly, embodiments of this disclosure provide a video encoding method applied to a generative video compression framework, comprising: encoding key reference frames in an original video to obtain a first encoded bitstream; representing subsequent frames in the original video using compact transmission symbols and encoding the representation results to obtain a second encoded bitstream; determining the first encoded bitstream and the second encoded bitstream as the bitstream of the video corresponding to the original video, and sending the bitstream of the video, so that the decoder, after receiving the bitstream of the video, uses the above-described video decoding method to perform video decoding to obtain the target video.

[0007] Thirdly, embodiments of this disclosure provide a video decoding apparatus applied to a generative video compression framework, comprising: a calculation module, configured to receive a bitstream of video and decode it to obtain a reconstructed video, and then calculate the distribution parameters of each color channel of a key reference frame in the reconstructed video, wherein the distribution parameters are used to characterize the color features of the corresponding color channel; a first calibration module, configured to calibrate the color distribution of subsequent frames in the reconstructed video for the same color channel using the distribution parameters of the key reference frame, to obtain calibrated subsequent frames; and a first generation module, configured to generate a target video based on the key reference frame and the calibrated subsequent frames.

[0008] Fourthly, embodiments of this disclosure provide a video encoding apparatus applied to a generative video compression framework, comprising: a first encoding module for encoding key reference frames in an original video to obtain a first encoded bitstream; a second encoding module for representing subsequent frames in the original video using compact transmission symbols and encoding the representation results to obtain a second encoded bitstream; and a first transmission module for determining the first encoded bitstream and the second encoded bitstream as the bitstream of the video corresponding to the original video and transmitting the bitstream of the video, so that the decoder, after receiving the bitstream of the video, performs video decoding using the aforementioned video decoding method to obtain the target video.

[0009] Fifthly, embodiments of this disclosure provide a non-transitory computer-readable storage medium storing a bitstream of video, the non-transitory computer-readable storage medium being part of a computing device configured to execute a set of instructions to decode the bitstream of video according to operations including: after decoding the bitstream of video to obtain a reconstructed video, calculating distribution parameters of each color channel of a key reference frame in the reconstructed video, wherein the distribution parameters are used to characterize the color features of the corresponding color channel; calibrating the color distribution of subsequent frames in the reconstructed video using the distribution parameters of the key reference frame for the same color channel, obtaining calibrated subsequent frames; and generating a target video based on the key reference frame and the calibrated subsequent frames.

[0010] In this embodiment, after receiving and decoding the bitstream of a video to obtain a reconstructed video, the distribution parameters of each color channel in the key reference frame of the reconstructed video are calculated. These distribution parameters characterize the color features of the corresponding color channel. For the same color channel, the distribution parameters of the key reference frame are used to calibrate the color distribution of subsequent frames in the reconstructed video, resulting in calibrated subsequent frames. A target video is then generated based on the key reference frame and the calibrated subsequent frames. In other words, this embodiment uses a channel-by-channel calibration method, employing the distribution parameters of each color channel in the key reference frame to calibrate the color distribution of subsequent frames corresponding to that color channel. This solves the technical problem of color shift in reconstructed videos in related technologies, achieving the technical effect of improving the quality of the reconstructed video.

[0011] The above description is only an overview of the technical solution of this disclosure. In order to better understand the technical means of this disclosure, it can be implemented in accordance with the contents of the specification. In order to make the above and other objects, features and advantages of this disclosure more obvious and understandable, specific embodiments of this disclosure are given below. Attached Figure Description

[0012] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments according to this disclosure and should not be construed as limiting the scope of this disclosure.

[0013] Figure 1 illustrates a video compression framework in the related art provided in the embodiments of this disclosure that follows a prediction-transform architecture;

[0014] Figure 2 shows a schematic diagram of the basic framework of the end-to-end video compression depth model in the related art provided in the embodiments of this disclosure;

[0015] Figure 3 shows a schematic diagram of the basic framework of a deep learning video generation and compression scheme based on a first-order motion model in the related technologies provided in the embodiments of this disclosure;

[0016] Figure 4 shows a schematic diagram of the basic framework of a deep learning-based video generation and compression scheme based on compact feature representation in the related technologies provided in the embodiments of this disclosure;

[0017] Figures 5a and 5b show schematic diagrams of the data distribution of channel pixel values ​​in the original frame and the reconstructed frame at different resolutions, respectively.

[0018] Figures 6a-6c show schematic diagrams of color deviation between the original frame and reconstructed frames with different pixels;

[0019] Figure 7 shows a schematic flowchart of a video decoding method provided in an embodiment of this disclosure;

[0020] Figure 8 shows a schematic diagram of the encoding-decoding process of the GVC algorithm provided in the embodiments of this disclosure;

[0021] Figure 9 shows a schematic diagram of a color plane and corresponding pixel values ​​provided in an embodiment of this disclosure;

[0022] Figure 10 shows a schematic flowchart of a video color calibration method provided in an embodiment of this disclosure;

[0023] Figure 11 shows a structural block diagram of a video decoding device provided in an embodiment of the present disclosure;

[0024] Figure 12 shows a schematic flowchart of another video decoding method provided in an embodiment of this disclosure;

[0025] Figure 13 shows a structural block diagram of another video decoding device provided in an embodiment of this disclosure;

[0026] Figure 14 shows a schematic flowchart of a video encoding method provided in an embodiment of this disclosure;

[0027] Figure 15 shows a structural block diagram of a video encoding apparatus provided in an embodiment of the present disclosure;

[0028] Figure 16 shows a schematic flowchart of another video encoding method provided in an embodiment of this disclosure;

[0029] Figure 17 shows a structural block diagram of another video encoding apparatus provided in an embodiment of this disclosure. Detailed Implementation

[0030] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of this disclosure. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.

[0031] To facilitate understanding of the technical solutions of the embodiments of this disclosure, the related technologies of the embodiments of this disclosure are described below. The following related technologies are optional solutions and can be combined with the technical solutions of the embodiments of this disclosure in any way, and all of them fall within the protection scope of the embodiments of this disclosure.

[0032] The upcoming era of AI-Generated Content (AIGC) is witnessing the rapid development of video coding technology in smarter, more immersive, and interactive applications. Generative Video Coding (GVC) is one of the key technologies. GVC utilizes the powerful inference capabilities of deep generative models for visual data compression, achieving superior rate-distortion (RD) performance compared to traditional hybrid encoders such as High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC).

[0033] In particular, most existing generative video encoders have evolved from deep image animation methods, which can characterize high-dimensional visual input signals into compact representations and leverage powerful deep generative models to achieve high-quality signal reconstruction / animation. For example, the Deep Animation Codec utilizes 2D keypoint representations to achieve ultra-low bitrate video conferencing. Similarly, in free-viewpoint controlled face conversation video coding, 3D keypoints are used, and the feature matrix can more compactly represent the temporal trajectory of the face.

[0034] However, the capabilities of generative coding are limited by its feature design and generation scheme. On the one hand, these generative video encoders primarily use explicit feature representations with actual physical characteristics, leading to unnecessary compression redundancy. Simultaneously, this representation lacks sufficient expressiveness and generality, failing to handle more complex scenarios, such as moving human figures. In some techniques, explicit features (including landmarks, keypoints, and segmentation maps) are used for low-bandwidth video chat compression, and it has been observed that different feature representations result in significantly different bandwidth requirements.

[0035] Prior to this disclosure, a related technology was the conventional video compression standards, such as Advanced Video Coding (AVC), HEVC, and VVC, which, through meticulous development, have achieved excellent compression performance. These standards utilize a block-based hybrid video coding framework to take advantage of spatial, temporal, and entropy redundancy in the video. In this framework, the video compression encoder generates a bitstream based on the current input frame, while the decoder reconstructs the video frame from the received bitstream. The classic video compression framework shown in Figure 1 follows a prediction-transform architecture. Specifically, the input frame x... t It is divided into a series of blocks of the same size, i.e., square regions (e.g., 8×8 pixel blocks). The encoding process at the encoder end of a traditional video compression algorithm mainly includes the following steps: Step 1, Motion estimation: Estimate the current frame x t Compared with the previous reconstructed frame The movement between blocks is used to obtain the motion vector v corresponding to each block. t Step 2, Motion Compensation: This is achieved by applying the motion vector v defined in Step 1. t The corresponding pixels from the previous reconstructed frame are copied to the current frame to obtain the predicted frame. Then calculate the original frame x. t With the predicted frame The residual r between t ,Right now Step 3, Transformation and Quantization: Transform the residual r obtained in Step 2 into a quantized form. t Quantified as Before quantization, a linear transformation (e.g., Discrete Cosine Transform (DCT)) is used to obtain better compression performance. Step 4, Inverse Transform: Utilizing the quantization results from Step 3. Perform an inverse transformation to obtain the reconstructed residuals. Step 5, Entropy Coding: Use the entropy coding method to encode the motion vector v from Step 1. t And the quantification results in step 3 Encode into a bitstream and send to the decoder. Step 6, Frame Reconstruction: Reconstruct the frame from step 2 by... and the reconstruction residual in step 4 Adding them together yields the reconstructed frame. Right now The reconstructed frame will be used for motion estimation of the (t+1)th frame in step 1. For the decoder, based on the bitstream provided by the encoder in step 5, motion compensation in step 2 and inverse quantization in step 4 are performed, followed by frame reconstruction in step 6 to obtain the reconstructed frame.

[0036] Another related technology prior to this disclosure is end-to-end deep learning-based video compression. With the rapid development of deep learning, many deep learning-based algorithms have been introduced to replace or enhance video coding tools, including intra / inter-frame prediction, entropy coding, and intra-loop filtering. End-to-end image / video compression algorithms have been proposed that jointly optimize the entire image / video compression framework, rather than designing individual modules. For example, the end-to-end video coding scheme Deep Video Compression (DVC) jointly optimizes all components of video compression. Furthermore, content adaptation and error propagation awareness are considered, and an online encoder update scheme has been developed to improve video compression performance. Moreover, Feature-based Video Compression (FVC) has been implemented by developing all the main modules of the end-to-end compression framework within the feature space. Based on recurrent probabilistic models and weighted recurrent quality enhancement networks, Recurrent Learning Video Compression (RLVC) and Hierarchical Learned Video Compression (HLVC) have been proposed to fully utilize the temporal correlation between video frames. And four effective modules are introduced in Multiple Frames Prediction for Learned Video Compression (M-LVC). However, these traditional or learning-based methods are aimed at general natural scenes without specifically considering human content such as faces, bodies, or other parts. Figure 2 shows the basic framework of the first end-to-end video compression deep model, which jointly optimizes all components of video compression, such as motion estimation, motion compression, and residual compression. Specifically, it uses learning-based optical flow estimation to obtain motion information and reconstruct the current frame, and then uses two autoencoder-style neural networks to compress the corresponding motion and residual information. All modules learn together through a single loss function, and they cooperate with each other to trade off between reducing the number of compressed bits and improving the quality of the decoded video. There is a one-to-one correspondence between the novel end-to-end deep framework shown in Figure 2 and the traditional video compression framework shown in Figure 1. These relationships and their differences are briefly summarized as follows: Step 1, Motion Estimation and Compression: A Convolutional Neural Network (CNN) model is used to estimate optical flow, which is regarded as motion information v t Instead of directly encoding the raw optical flow values, a motion vector (MV) encoder-decoder network is used to compress and decode the optical flow values, where the quantized motion representation is denoted as... Then, the corresponding reconstructed motion information can be decoded using the MV decoder network. Step 2, Motion Compensation: A motion compensation network was designed to obtain the predicted frame based on the optical flow obtained in Step 1. Step 3, Transformation, Quantization, and Inverse Transformation: A highly nonlinear residual encoder-decoder network is used instead of a linear transform, and the residual... Nonlinear mapping to representation y t , then y t Quantified as Quantization was used to construct the end-to-end training scheme. The quantized representation... It is fed into the residual decoder network to obtain the reconstructed residual. Step 4, Entropy Encoding: During the testing phase, the quantized motion representation from Step 1 is... and the residual representation in step 3 The bits are encoded and sent to the decoder. During the training phase, a CNN is used to estimate the cost per bit. and The probability distribution of each symbol in the image. Step 5, frame reconstruction: This is the same as the traditional method and will not be described in detail here.

[0037] Another related technology prior to this disclosure is generative video coding. With the emergence of deep generative models such as Variational Auto-Encoding (VAE) and Generative Adversarial Networks (GANs), exciting progress has been made in the performance of facial video compression. X2Face has been designed to control face generation using images, audio, and pose codes. Furthermore, a realistic neural conversational head model implemented through few-shot adversarial learning has been proposed. Face-vidtovid has been proposed for video-to-video synthesis tasks. A novel scheme has also been proposed to drive the generative model to render target frames by utilizing compact 3D keypoint representations. Additionally, a mobile device video chat system based on a First Order Motion Model (FOMM) has been designed. VSBNet has been proposed to reconstruct original frames from landmarks using adversarial learning. Furthermore, an end-to-end conversational head video compression framework (CFTE) based on compact feature learning has been proposed, cleverly designed for efficient conversational face video compression, suitable for ultra-low bandwidth scenarios. The CFTE scheme utilizes compact feature representations to compensate for temporal evolution and reconstructs target face video frames in an end-to-end manner. Furthermore, it can be incorporated into video coding frameworks under rate-distortion target supervision. Although these algorithms have successfully achieved frame reconstruction using a small number of facial parameters thanks to the powerful rendering capabilities of deep generative models, some head pose movements and facial expression changes still fail to be accurately rendered compared to the original moving video. Figure 3 illustrates the basic framework of a deep learning-based video generation and compression scheme, which is based on a First Order Motion Model (FOMM). FOMM follows the motion of the driving video by deforming a reference source frame. While this method is applicable to various types of videos (such as Tai Chi and cartoons), some research teams have focused on its application to facial animation. FOMM follows an encoder-decoder architecture that includes motion transfer components: 1. A keypoint extractor learns using equivariant loss, without explicit labels. Through this keypoint extractor, two sets of ten learned keypoints are calculated for the source and driving frames. The learned keypoints are derived from a feature map of size 64×64 channels using a Gaussian map function, so each corresponding keypoint can represent feature information from different channels. It's worth mentioning that each keypoint is a (x, y) coordinate, which can represent the most important information in the feature map. 2. The dense motion network uses landmarks and source frames to generate a dense motion field and an occlusion map.3. The encoder encodes the source frame using traditional image / video compression methods such as High Efficiency Video Coding (HEVC) / VVC or Joint Photographic Experts Group (JPEG) / BPG. Here, VVC is used to compress the source frame. 4. The generated feature map is warped using a dense motion field (through differentiable grid sampling operations) and then multiplied with the occlusion map. 5. The decoder generates an image from the warped map.

[0038] Another related technology prior to this disclosure is the basic framework of another deep learning video generation compression scheme based on compact feature representation, as shown in Figure 4, which follows an encoder-decoder architecture. At the encoding end, the compression framework consists of three modules: an encoder for compressing keyframes, a feature extractor for extracting compact human features from subsequent frames, and a feature encoding module for compressing inter-frame prediction residuals of the compact human features. First, keyframes representing human textures are compressed using a VVC encoder. Each subsequent frame is represented by a compact feature matrix of size 1×4×4 using the compact feature extractor. It is important to note that the size of the compact feature matrix is ​​not fixed; the number of feature parameters can be increased or decreased depending on the specific bit rate consumption requirements. Subsequently, these extracted features undergo inter-frame prediction and quantization processing, and finally, the residuals are entropy-encoded to generate the final bitstream. At the decoding end, the compression framework also includes three main modules: a decoding module for reconstructing keyframes, a module for reconstructing compact features through entropy decoding and compensation, and a module for generating the final video using the reconstructed features and the decoded keyframes. More specifically, in the process of generating the final video, keyframes decoded from the VVC bitstream can be further represented in feature form through compact feature extraction. Subsequently, based on the features of the keyframes and subsequent frames, the relevant sparse motion field is calculated, thereby generating pixel-level dense motion maps and occlusion maps. Finally, based on a deep generative model, using the decoded keyframes, pixel-level dense motion maps, and occlusion maps represented by implicit motion fields, a final video with accurate appearance, pose, and expression is generated.

[0039] While generative video compression schemes can achieve promising rate-distortion (RD) performance, they still face some drawbacks and challenges that limit further performance improvements and practical applications. The data distribution modeling of the generated reconstruction may be inaccurate compared to the original distribution due to the generative training process. This inaccurate data distribution modeling can lead to unpleasant subjective distortions such as color bias in the final reconstructed video. For example, we observed the first subsequent frame reconstruction using the CFTE model at 256×256 and 512×512 resolutions. As shown in Figures 5a–5b, the solid and dashed lines represent the data distribution of the channel pixel values ​​in the original and reconstructed frames, respectively. It can be observed that the distribution of the reconstructed frames deviates significantly in the R, G, and B channels.

[0040] Specifically, for the original frame shown in Figure 6a, the values ​​of all channels are biased towards smaller values ​​for the 256×256 pixel reconstruction. However, for the 512×512 pixel reconstruction, the red channel is biased towards larger values, while the blue channel is biased towards smaller values. Therefore, the 256×256 pixel reconstruction appears slightly darker, as shown in Figure 6b, while the 512×512 pixel reconstruction shows a distinctly warm tone, as shown in Figure 6c.

[0041] In view of this, embodiments of this disclosure provide a video decoding method to solve all or part of the above-mentioned technical problems. This method can be applied to generative video compression frameworks, and is also applicable to all scenarios involving color shifts caused by the use of generative models. In other words, the video decoding method provided in embodiments of this disclosure can be applied to any video compression encoding-decoding related technologies.

[0042] As shown in Figure 7, the above video decoding method may include:

[0043] S702, after receiving the bitstream of the video and decoding it to obtain the reconstructed video, calculate the distribution parameters of each color channel of the key reference frame in the reconstructed video, wherein the distribution parameters are used to characterize the color features of the corresponding color channel.

[0044] It should be noted that the reconstructed video mentioned above includes key reference frames and subsequent frames. The key reference frames typically contain complete image information, while subsequent frames generally store differences from the reference frames. The method for obtaining this reconstructed video is shown in Figure 8. Figure 8 illustrates the general encoding-decoding process of the GVC algorithm, taking generative face video coding as an example. At the encoder end, the key reference frames of the video are compressed using traditional encoding techniques, while subsequent frames are represented using compact transmission symbols and encoded into the output bitstream. At the decoder end, the decoded key reference frames and the compact face information are input together into the synthesis model to reconstruct the video, thus obtaining the reconstructed video.

[0045] It should be noted that the color channels mentioned above, in addition to RGB (red, green, blue) channels, can also include channels from other color spaces or additional channels for specific application scenarios. For example, YCbCr (Y: luminance channel, representing the image's brightness information; Cb: blue chrominance channel, representing the difference between blue and luminance; Cr: red chrominance channel, representing the difference between red and luminance) and YUV (Y: luminance channel, representing the image's brightness information; U: chrominance U channel, representing blue chrominance information; V: chrominance V channel, representing red chrominance information). Depending on the specific application scenario, color spaces such as YPbPr (Y: luminance channel, representing the image's brightness information; Pb: blue chrominance channel; Pr: red chrominance channel), HSL / HSV (H: hue, representing color; S: saturation, representing color purity; L / V: brightness / value, representing color lightness / darkness), and Lab (L: luminance; a: representing color information from green to magenta; b: representing color information from blue to yellow) can also be used.

[0046] Additionally, it should be noted that the types of videos mentioned above include, but are not limited to: facial videos, motion videos, conference and speech videos, educational videos, entertainment and media videos, game videos, user-generated content, surveillance videos, real-time communication videos, virtual reality and augmented reality videos, etc., which can be expanded according to specific application scenarios and needs. Each type of video has its own specific visual and dynamic characteristics.

[0047] Optionally, the above distribution parameters include, but are not limited to, the mean and standard deviation, and may also be other characteristics of the color value distribution, such as kurtosis, skewness, frequency domain characteristics, etc., which are not limited here.

[0048] S704, for the same color channel, use the distribution parameters of the key reference frame to calibrate the color distribution of subsequent frames in the reconstructed video to obtain calibrated subsequent frames.

[0049] Optionally, in this embodiment of the disclosure, a channel-by-channel approach is adopted, using the distribution parameters of the key reference frame to calibrate the color distribution of subsequent frames in the reconstructed video. For example, assuming the color channels are RGB channels, for the R channel, the distribution parameters of the R channel of the key reference frame are used to calibrate the color distribution of the R channel in subsequent frames of the reconstructed video. For the B channel, the distribution parameters of the B channel of the key reference frame are used to calibrate the color distribution of the B channel in subsequent frames of the reconstructed video.

[0050] S706, Generate the target video based on the key reference frame and the calibrated subsequent frames.

[0051] Optionally, in this embodiment of the disclosure, the above calibration step is a post-processing step of GVC decoding, which can improve the subjective quality of the generated video without increasing the bitrate.

[0052] Through the above steps S702-S706, after receiving and decoding the bitstream of the video to obtain the reconstructed video, the distribution parameters of each color channel in the key reference frame of the reconstructed video are calculated, wherein the distribution parameters are used to characterize the color features of the corresponding color channel; for the same color channel, the color distribution of subsequent frames in the reconstructed video is calibrated using the distribution parameters of each color channel of the key reference frame to obtain calibrated subsequent frames; and a target video is generated based on the key reference frame and the calibrated subsequent frames. In other words, this embodiment of the present disclosure uses a channel-by-channel calibration method, using the distribution parameters of each color channel of the key reference frame to calibrate the color distribution of subsequent frames corresponding to the color channel, thereby solving the technical problem of color shift in reconstructed videos in related technologies and achieving the technical effect of improving the quality of reconstructed videos.

[0053] Since the combination of multiple color channels can represent various colors—for example, the RGB channel, through the combination of red, green, and blue channels, can represent various colors, becoming the most basic way of presenting an image—in one possible implementation disclosed herein, for the same color channel, the color distribution of subsequent frames in the video is calibrated and reconstructed using the distribution parameters of a key reference frame to obtain calibrated subsequent frames, including:

[0054] S11, obtain the first mean and first standard deviation of the first color channel of the key reference frame.

[0055] It should be noted that when the color channel is an RGB channel, the first color channel mentioned above can be an R channel, a G channel, or a B channel.

[0056] S12, using the first mean and the first standard deviation, calibrate the color values ​​of each pixel in the second color channel of the subsequent frame to obtain the calibrated color distribution of the second color channel, wherein the first color channel and the second color channel are the same color channel.

[0057] That is, after obtaining the first mean and the first standard deviation, the color values ​​of each pixel in the same color channel of subsequent frames are calibrated using the first mean and the first standard deviation. For example, taking the R channel and 4×4 pixels as an example, after obtaining the mean μ and standard deviation σ of the key reference frame, the color values ​​of 16 pixels are calibrated using the mean μ and the standard deviation σ respectively to obtain the adjusted color values ​​of the 16 pixels.

[0058] Since the final color of an image frame is determined by merging multiple color channels, this disclosure also proposes the following embodiments:

[0059] S13, combine the color distributions of multiple color channels after calibration to obtain the calibrated subsequent frame.

[0060] By performing channel-by-channel calibration in steps S11 to S13, it can be ensured that the compressed video maintains the same color performance as the original video.

[0061] Optionally, the color values ​​of each pixel in the second color channel of the subsequent frame are calibrated using the first mean and the first standard deviation to obtain the calibrated color distribution of the second color channel, including:

[0062] S121, obtain the initial color value of the target pixel from the second color channel of the subsequent frame, wherein the initial color value is used to characterize the color intensity of the target pixel in the second color channel.

[0063] The target pixel mentioned above can be any pixel among all pixels in the second color channel of subsequent frames, or it can be any pixel among only a portion of the pixels. That is, for a frame of image, only the color distribution of the region of interest to the user is calibrated. For example, in an image that includes a background and a face, only the color distribution of the face is calibrated, while the color distribution of the background is not calibrated, further realizing the flexibility and interactivity of the video color calibration process.

[0064] Optionally, the color value of a pixel can be determined based on the number of bits in the color channel. For example, for an 8-bit color channel, the color value of a pixel can be from 0 to 255; for a 16-bit color channel, the color value of a pixel can be from 0 to 65535. The value range for a 32-bit floating-point color channel can be from 0.0 to 1.0.

[0065] S122, determine the product of the initial color value and the first standard deviation to obtain the first result.

[0066] S123, sum the first result and the first mean to obtain the second result.

[0067] S124, set the second result as the color value of the target pixel after calibration.

[0068] Through the above steps S121 to S124, the standard deviation of the key reference frame is multiplied by the initial color value of the target pixel to be calibrated, and then the product is summed with the mean of the key reference frame to calibrate the initial color of the target pixel. This can further ensure that the compressed video is consistent with the original video in terms of color performance.

[0069] To improve algorithm efficiency and accuracy, enhance data comparability, and reduce the impact of outliers, this disclosure also proposes that, before determining the product of the initial color value and the first standard deviation to obtain the first result, the above method further includes:

[0070] S125, obtain the second mean and second standard deviation of the second color channel of the subsequent frame;

[0071] S126, using the second mean and the second standard deviation, the color values ​​of each pixel in the second color channel of the subsequent frame are standardized.

[0072] Optionally, standardizing the color values ​​of each pixel in the second color channel of subsequent frames using the second mean and the second standard deviation can specifically be done by: taking the difference between the color value of the pixel in the second color channel and the second mean, and then comparing the difference with the second standard deviation to obtain the standardized result of the color value of each pixel.

[0073] Optionally, in this embodiment of the disclosure, the mean and standard deviation corresponding to the color channels are also proposed to be determined in the following manner:

[0074] S21, Determine the average value based on the color value of each pixel in the color channel and the total number of pixels;

[0075] S22, Determine the standard deviation based on the color value of each pixel in the color channel and the mean value.

[0076] For example, taking the color channels as RGB channels, assuming the RGB channels are arranged in order, the mean and standard deviation of each RGB channel in the key reference frame are calculated as shown in the following formulas (1) and (2):

[0077] in, The key reference frame for reconstruction is represented by H and W, which represent the height and width of the frame, respectively. The index [c, i, j] represents the value at the c-th channel, the i-th position in the height direction, and the j-th position in the width direction.

[0078] For example, for a 4×4 frame, its color plane and corresponding pixel values ​​are shown in Figure 9. The mean value of the red channel can be calculated as:

[0079] Similarly, each subsequent frame of the reconstruction The channel-wise distribution is calculated as μ i,r / g / b and σ i,r / g / b Then, for each Channel calibration calculation is shown in Formula 3:

[0080] In this process, each channel of a subsequent frame is first standardized using its own distribution parameters, and then adjusted based on the distribution parameters provided by the key reference frame for reconstruction.

[0081] In summary, the video decoding method of this disclosure provides a channel-by-channel video color calibration method as a post-processing step in GVC decoding. As shown in Figure 10, after decoding the input video, an output video is obtained, which includes a key reference frame and reconstructed subsequent frames. The standard deviation and mean of the reconstructed key reference frame are calculated, and then the mean and standard deviation of the reconstructed key reference frame are used to provide the distribution parameters required for subsequent frame calibration. The specific calibration method is described in Formula 3 above. Finally, the calibrated subsequent frames are output. Through the above method of this disclosure, the subjective quality of the generated video can be improved without increasing the additional bitrate cost.

[0082] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0083] The technical solutions of this disclosure and how they solve the aforementioned technical problems are described in detail below with specific embodiments. The listed specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this disclosure will be described in detail below with reference to the accompanying drawings.

[0084] Corresponding to the application scenarios and methods provided in the embodiments of this disclosure, the embodiments of this disclosure also provide a video decoding device applied to a generative video compression framework. Figure 11 shows a structural block diagram of a video decoding device according to an embodiment of this disclosure, which may include:

[0085] The calculation module 1102 is used to receive the bit stream of the video and decode it to obtain the reconstructed video, and then calculate the distribution parameters of each color channel of the key reference frame in the reconstructed video, wherein the distribution parameters are used to characterize the color features of the corresponding color channel.

[0086] It should be noted that the reconstructed video mentioned above includes key reference frames and subsequent frames. The key reference frames typically contain complete image information, while subsequent frames generally store differences from the reference frames. The method for obtaining this reconstructed video is shown in Figure 8. Figure 8 illustrates the general encoding-decoding process of the GVC algorithm, taking generative face video coding as an example. At the encoder end, the key reference frames of the video are compressed using traditional encoding techniques, while subsequent frames are represented using compact transmission symbols and encoded into the output bitstream. At the decoder end, the decoded key reference frames and the compact face information are input together into the synthesis model to reconstruct the video, thus obtaining the reconstructed video.

[0087] It should be noted that the color channels mentioned above, in addition to RGB (red, green, blue) channels, can also include channels from other color spaces or additional channels for specific application scenarios. For example, YCbCr (Y: luminance channel, representing the image's brightness information; Cb: blue chrominance channel, representing the difference between blue and luminance; Cr: red chrominance channel, representing the difference between red and luminance) and YUV (Y: luminance channel, representing the image's brightness information; U: chrominance U channel, representing blue chrominance information; V: chrominance V channel, representing red chrominance information). Depending on the specific application scenario, color spaces such as YPbPr (Y: luminance channel, representing the image's brightness information; Pb: blue chrominance channel; Pr: red chrominance channel), HSL / HSV (H: hue, representing color; S: saturation, representing color purity; L / V: brightness / value, representing color lightness / darkness), and Lab (L: luminance; a: representing color information from green to magenta; b: representing color information from blue to yellow) can also be used.

[0088] Additionally, it should be noted that the types of videos mentioned above include, but are not limited to: facial videos, motion videos, conference and speech videos, educational videos, entertainment and media videos, game videos, user-generated content, surveillance videos, real-time communication videos, virtual reality and augmented reality videos, etc., which can be expanded according to specific application scenarios and needs. Each type of video has its own specific visual and dynamic characteristics.

[0089] Optionally, the above distribution parameters include, but are not limited to, the mean and standard deviation, and may also be other characteristics of the color value distribution, such as kurtosis, skewness, frequency domain characteristics, etc., which are not limited here.

[0090] The first calibration module 1104 is used to calibrate the color distribution of subsequent frames in the reconstructed video using the distribution parameters of the key reference frame for the same color channel, so as to obtain the calibrated subsequent frames.

[0091] Optionally, in this embodiment of the disclosure, a channel-by-channel approach is adopted, using the distribution parameters of the key reference frame to calibrate the color distribution of subsequent frames in the reconstructed video. For example, assuming the color channels are RGB channels, for the R channel, the distribution parameters of the R channel of the key reference frame are used to calibrate the color distribution of the R channel in subsequent frames of the reconstructed video. For the B channel, the distribution parameters of the B channel of the key reference frame are used to calibrate the color distribution of the B channel in subsequent frames of the reconstructed video.

[0092] The first generation module 1106 is used to generate the target video based on the key reference frame and the calibrated subsequent frames.

[0093] Optionally, in this embodiment of the disclosure, the above calibration step is a post-processing step of GVC decoding, which can improve the subjective quality of the generated video without increasing the bitrate.

[0094] Using the apparatus shown in Figure 11, after receiving and decoding the bitstream of the video to obtain the reconstructed video, the distribution parameters of each color channel in the key reference frame of the reconstructed video are calculated. These distribution parameters characterize the color features of the corresponding color channel. For the same color channel, the distribution parameters of the key reference frame are used to calibrate the color distribution of subsequent frames in the reconstructed video, resulting in calibrated subsequent frames. The target video is then generated based on the key reference frame and the calibrated subsequent frames. In other words, this embodiment of the present disclosure uses a channel-by-channel calibration method, employing the distribution parameters of each color channel in the key reference frame to calibrate the color distribution of subsequent frames corresponding to that color channel. This solves the technical problem of color shift in reconstructed videos in related technologies and achieves the technical effect of improving the quality of the reconstructed video.

[0095] The aforementioned first calibration module 1104 includes: an acquisition unit, configured to acquire the first mean and first standard deviation of the first color channel of the key reference frame; a calibration unit, configured to use the first mean and the first standard deviation to calibrate the color values ​​of each pixel of the second color channel of the subsequent frame to obtain the calibrated color distribution of the second color channel, wherein the first color channel and the second color channel are the same color channel; and a combination unit, configured to combine the calibrated color distributions of multiple color channels to obtain the calibrated subsequent frame.

[0096] The calibration unit includes: a first acquisition subunit, configured to acquire an initial color value of a target pixel from the second color channel of the subsequent frame, wherein the initial color value is used to characterize the color intensity of the target pixel in the second color channel; a second acquisition subunit, configured to determine the product of the initial color value and the first standard deviation to obtain a first result; a third acquisition subunit, configured to sum the first result and the first mean to obtain a second result; and a setting subunit, configured to set the second result as the calibrated color value of the target pixel.

[0097] Optionally, the first calibration module 1104 further includes: a first processing unit, configured to obtain a second mean and a second standard deviation of the second color channel of the subsequent frame before determining the product of the initial color value and the first standard deviation to obtain a first result; and a second processing unit, configured to use the second mean and the second standard deviation to standardize the color value of each pixel of the second color channel of the subsequent frame.

[0098] The aforementioned device is also used to determine the mean and standard deviation corresponding to the color channels in the following ways:

[0099] The mean is determined based on the color values ​​of each pixel in the color channel and the total number of pixels; the standard deviation is determined based on the color values ​​of each pixel in the color channel and the mean.

[0100] The functions of each module in the apparatus of this embodiment can be found in the corresponding description in the above method, and they have corresponding beneficial effects, which will not be repeated here.

[0101] Corresponding to the application scenarios and methods provided in the embodiments of this disclosure, the embodiments of this disclosure also provide another video decoding method. Figure 12 shows a video decoding method according to an embodiment of this disclosure, applied to a generative video compression framework, including:

[0102] S1202, receive the bit stream of video and indication information, wherein the indication information includes: first indication information, second indication information and third indication information, the first indication information is used to indicate the key reference frame, the second indication information is used to indicate the subsequent frame to be calibrated, and the third indication information is used to indicate the distribution parameters of each color channel of the key reference frame in the original video.

[0103] Optionally, in this embodiment of the disclosure, the above-mentioned indication information may also include at least one of the above-mentioned indication information. For example, for the reconstructed video, there may be some distortion, so the distribution parameters of the color channels of the key reference frame in the reconstructed video may have some error. In this case, the distribution parameters of each color channel of the key reference frame corresponding to the original video, i.e., the above-mentioned third indication information, can be received, and then the color distribution calibration of the subsequent frame to be calibrated can be performed based on the determined key reference frame of the reconstructed video and the subsequent frame to be calibrated.

[0104] Since the key reference frame may be in the first or last frame of the video, to further improve video decoding efficiency, the position of the key reference frame can be indicated by the indication information sent by the encoder (i.e., the first indication information mentioned above). Then, based on the distribution parameters of each color channel of the key reference frame of the reconstructed video and the subsequent frames to be calibrated, color distribution calibration is performed on the subsequent frames to be calibrated. Not every subsequent frame needs color distribution calibration; in this case, the subsequent frames to be calibrated can be indicated by the indication information sent by the encoder (i.e., the second indication information mentioned above). That is, in this embodiment of the disclosure, by combining the above-mentioned indication information sent by the encoder, the decoding efficiency and decoding accuracy of the decoder can be further improved.

[0105] Optionally, in this embodiment of the disclosure, the above-mentioned indication information may be encoded in the bitstream of the video or sent separately. Here, there is no specific limitation on the method of sending the above-mentioned indication information.

[0106] S1204 After decoding the bitstream of the video to obtain the reconstructed video, the key reference frame and the subsequent frame to be calibrated in the reconstructed video are determined based on the first indication information and the second indication information.

[0107] It should be noted that the reconstructed video mentioned above includes key reference frames and subsequent frames. The key reference frames typically contain complete image information, while subsequent frames generally store differences from the reference frames. The method for obtaining this reconstructed video is shown in Figure 8. Figure 8 illustrates the general encoding-decoding process of the GVC algorithm, taking generative face video coding as an example. At the encoder end, the key reference frames of the video are compressed using traditional encoding techniques, while subsequent frames are represented using compact transmission symbols and encoded into the output bitstream. At the decoder end, the decoded key reference frames and the compact face information are input together into the synthesis model to reconstruct the video, thus obtaining the reconstructed video.

[0108] It should be noted that the color channels mentioned above, in addition to RGB (red, green, blue) channels, can also include channels from other color spaces or additional channels for specific application scenarios. For example, YCbCr (Y: luminance channel, representing the image's brightness information; Cb: blue chrominance channel, representing the difference between blue and luminance; Cr: red chrominance channel, representing the difference between red and luminance) and YUV (Y: luminance channel, representing the image's brightness information; U: chrominance U channel, representing blue chrominance information; V: chrominance V channel, representing red chrominance information). Depending on the specific application scenario, color spaces such as YPbPr (Y: luminance channel, representing the image's brightness information; Pb: blue chrominance channel; Pr: red chrominance channel), HSL / HSV (H: hue, representing color; S: saturation, representing color purity; L / V: brightness / value, representing color lightness / darkness), and Lab (L: luminance; a: representing color information from green to magenta; b: representing color information from blue to yellow) can also be used.

[0109] Additionally, it should be noted that the types of videos mentioned above include, but are not limited to: facial videos, motion videos, conference and speech videos, educational videos, entertainment and media videos, game videos, user-generated content, surveillance videos, real-time communication videos, virtual reality and augmented reality videos, etc., which can be expanded according to specific application scenarios and needs. Each type of video has its own specific visual and dynamic characteristics.

[0110] Optionally, the above distribution parameters include, but are not limited to, the mean and standard deviation, and may also be other characteristics of the color value distribution, such as kurtosis, skewness, frequency domain characteristics, etc., which are not limited here.

[0111] S1206, For the same color channel, use the third indication information to calibrate the color distribution of the subsequent frames to be calibrated in the reconstructed video, and obtain the calibrated subsequent frames.

[0112] Optionally, in this embodiment, a channel-by-channel approach is used to calibrate the color distribution of subsequent frames in the reconstructed video using the distribution parameters of the original video key reference frame. For example, assuming the color channels are RGB channels, for the R channel, the distribution parameters of the R channel of the original video key reference frame are used to calibrate the color distribution of the R channel in subsequent frames of the reconstructed video. For the B channel, the distribution parameters of the B channel of the original video key reference frame are used to calibrate the color distribution of the B channel in subsequent frames of the reconstructed video. For specific calibration methods, please refer to the calibration method in the first video decoding method described above, which will not be described in detail here.

[0113] S1208, generate the target video based on the key reference frame in the reconstructed video and the calibrated subsequent frame.

[0114] Optionally, in this embodiment of the disclosure, the above calibration step is a post-processing step of GVC decoding, which can improve the subjective quality of the generated video without increasing the bitrate.

[0115] Through the aforementioned steps S1202-S1208, the video bitstream and indication information are received. The indication information includes: first indication information indicating a key reference frame, second indication information indicating a subsequent frame to be calibrated, and third indication information indicating the distribution parameters of each color channel in the key reference frame of the original video. After decoding the video bitstream to obtain a reconstructed video, the key reference frame and the subsequent frame to be calibrated in the reconstructed video are determined based on the first and second indication information. For the same color channel, the third indication information is used to calibrate the color distribution of the subsequent frame to be calibrated in the reconstructed video, resulting in a calibrated subsequent frame. A target video is generated based on the key reference frame in the reconstructed video and the calibrated subsequent frame. In other words, this embodiment of the present disclosure, during the video decoding process, uses the video bitstream and indication information sent by the encoder to calibrate the color distribution of the corresponding color channel in the subsequent frame by means of the distribution parameters of each color channel in the indication information, based on the channel-by-channel calibration method. This solves the technical problem of color shift in reconstructed video in related technologies and achieves the technical effect of improving the quality of the reconstructed video.

[0116] Corresponding to the application scenarios and methods provided in the embodiments of this disclosure, the embodiments of this disclosure also provide another video decoding device. Figure 13 shows a structural block diagram of a video decoding device according to an embodiment of this disclosure. The device may include:

[0117] The receiving module 1302 is used to receive the bit stream of video and indication information, wherein the indication information includes: first indication information, second indication information and third indication information, the first indication information is used to indicate the key reference frame, the second indication information is used to indicate the subsequent frame to be calibrated, and the third indication information is used to indicate the distribution parameters of each color channel of the key reference frame in the original video.

[0118] The first determining module 1304 is used to determine the key reference frame and the subsequent frame to be calibrated in the reconstructed video based on the first indication information and the second indication information after decoding the bit stream of the video to obtain the reconstructed video.

[0119] The second calibration module 1306 is used to calibrate the color distribution of subsequent frames to be calibrated in the reconstructed video for the same color channel using the third indication information, so as to obtain the calibrated subsequent frames.

[0120] The second generation module 1308 generates a target video based on the key reference frame in the reconstructed video and the calibrated subsequent frames.

[0121] The apparatus shown in Figure 13 receives a video bitstream and indication information, wherein the indication information includes: first indication information indicating a key reference frame, second indication information indicating a subsequent frame to be calibrated, and third indication information indicating the distribution parameters of each color channel of the key reference frame in the original video. After decoding the video bitstream to obtain a reconstructed video, the key reference frame and the subsequent frame to be calibrated in the reconstructed video are determined based on the first and second indication information. For the same color channel, the color distribution of the subsequent frame to be calibrated in the reconstructed video is calibrated using the third indication information to obtain the calibrated subsequent frame. A target video is generated based on the key reference frame in the reconstructed video and the calibrated subsequent frame. In other words, in the video decoding process, this embodiment of the present disclosure, based on the video bitstream and indication information sent by the encoder, calibrates the color distribution of the corresponding color channel in the subsequent frame by means of the distribution parameters of each color channel in the indication information in a channel-by-channel calibration manner, thereby solving the technical problem of color offset in reconstructed video in related technologies and achieving the technical effect of improving the quality of reconstructed video.

[0122] The functions of each module in the apparatus of this embodiment can be found in the corresponding description in the above method, and they have corresponding beneficial effects, which will not be repeated here.

[0123] Corresponding to the application scenarios and methods provided in the embodiments of this disclosure, the embodiments of this disclosure also provide a video encoding method applied to a generative video compression framework. Figure 14 shows a flowchart of a video encoding method according to an embodiment of this disclosure. The method may include:

[0124] S1402, the key reference frames in the original video are encoded to obtain the first encoded bitstream.

[0125] S1404, the subsequent frames in the original video are represented using compact transmission symbols, and the representation results are encoded to obtain the second coded bitstream.

[0126] It should be noted that the specific methods for determining the first and second encoded bit streams mentioned above can be found in Figure 8, and will not be described in detail here.

[0127] S1406, the first encoded bitstream and the second encoded bitstream are determined as the bitstream of the video corresponding to the original video, and the bitstream of the video is sent so that the decoder, after receiving the bitstream of the video, uses one of the video decoding methods described above to perform video decoding and obtain the target video.

[0128] Through steps S1402 to S1406, the key reference frames in the original video are encoded to obtain a first encoded bitstream; subsequent frames in the original video are represented using compact transmission symbols, and the representation results are encoded to obtain a second encoded bitstream; the first encoded bitstream and the second encoded bitstream are determined as the bitstream of the video corresponding to the original video, and the bitstream of the video is sent so that the decoder, after receiving the bitstream of the video, uses one of the video decoding methods described above to decode the video and obtain the target video. That is, by sending the bitstream of the encoded video from the original video to the decoder, the decoder can calibrate the color distribution of the corresponding color channels of subsequent frames in a channel-by-channel calibration manner by using the distribution parameters of each color channel of the key reference frames. This solves the technical problem of color shift in reconstructed video in related technologies and achieves the technical effect of improving the quality of the reconstructed video.

[0129] Corresponding to the application scenarios and methods provided in the embodiments of this disclosure, the embodiments of this disclosure also provide a video encoding apparatus. Figure 15 shows a structural block diagram of a video encoding apparatus according to an embodiment of this disclosure. The apparatus may include:

[0130] The first encoding module 1502 is used to encode the key reference frames in the original video to obtain the first encoded bit stream;

[0131] The second encoding module 1504 is used to represent subsequent frames in the original video using compact transmission symbols and to encode the representation results to obtain a second encoded bitstream.

[0132] The first sending module 1506 is used to determine the first encoded bit stream and the second encoded bit stream as the bit stream of the video corresponding to the original video, and send the bit stream of the video so that the decoder, after receiving the bit stream of the video, uses the video decoding method described above to perform video decoding to obtain the target video.

[0133] Using the apparatus shown in Figure 15, key reference frames in the original video are encoded to obtain a first encoded bitstream. Subsequent frames in the original video are represented using compact transmission symbols, and the representation results are encoded to obtain a second encoded bitstream. The first and second encoded bitstreams are determined as the bitstreams of the video corresponding to the original video, and this video bitstream is sent. Upon receiving the video bitstream, the decoder uses one of the video decoding methods described above to decode the video and obtain the target video. In other words, by sending the bitstream of the encoded video to the decoder, the decoder can calibrate the color distribution of subsequent frames corresponding to the color channels using the distribution parameters of each color channel of the key reference frames in a channel-by-channel calibration manner. This solves the technical problem of color shift in reconstructed video in related technologies and achieves the technical effect of improving the quality of the reconstructed video.

[0134] The functions of each module in the apparatus of this embodiment can be found in the corresponding description in the above method, and they have corresponding beneficial effects, which will not be repeated here.

[0135] Corresponding to the application scenarios and methods provided in the embodiments of this disclosure, the embodiments of this disclosure also provide another video coding method applied to a generative video compression framework. Figure 16 shows a flowchart of another embodiment of the video coding method of this disclosure, which may include:

[0136] S1602, the key reference frames in the original video are encoded to obtain the first encoded bitstream.

[0137] S1604, the subsequent frames in the original video are represented using compact transmission symbols, and the representation results are encoded to obtain the second coded bitstream.

[0138] It should be noted that the specific methods for determining the first and second encoded bit streams mentioned above can be found in Figure 8, and will not be described in detail here.

[0139] S1606, the first encoded bitstream and the second encoded bitstream are determined as the bitstreams of the video corresponding to the original video.

[0140] S1608, determine the indication information and send the indication information and the bit stream of the video, so that after receiving the bit stream of the video and the indication information, the decoder uses the other video decoding method described above to decode the video and obtain the target video. The indication information includes: first indication information, second indication information and third indication information. The first indication information is used to indicate the key reference frame, the second indication information is used to indicate the subsequent frame to be calibrated, and the third indication information is used to indicate the distribution parameters of each color channel of the key reference frame in the original video.

[0141] Through steps S1602 to S1608, the key reference frames in the original video are encoded to obtain a first coded bitstream; subsequent frames in the original video are represented using compact transmission symbols, and the representation result is encoded to obtain a second coded bitstream; the first coded bitstream and the second coded bitstream are determined as the bitstream of the video corresponding to the original video; indication information is determined, and the indication information and the video bitstream are sent, so that the decoder, after receiving the video bitstream and the indication information, uses the other video decoding method described above to decode the video and obtain the target video. That is, by sending the bitstream of the encoded video and the indication information to the decoder, the decoder can combine the indication information and use a channel-by-channel calibration method to calibrate the color distribution of the corresponding color channels in subsequent frames by using the distribution parameters of each color channel of the key reference frames in the original video. This solves the technical problem of color shift in reconstructed video in related technologies and achieves the technical effect of improving the quality of reconstructed video.

[0142] Corresponding to the application scenarios and methods provided in the embodiments of this disclosure, the embodiments of this disclosure also provide another video encoding apparatus. Figure 17 shows a structural block diagram of a video encoding apparatus according to an embodiment of this disclosure. The apparatus may include:

[0143] The third encoding module 1702 is used to encode the key reference frames in the original video to obtain the first encoded bitstream;

[0144] The fourth encoding module 1704 is used to represent subsequent frames in the original video using compact transmission symbols and to encode the representation results to obtain the second encoded bitstream;

[0145] The second determining module 1704 is used to determine the first encoded bitstream and the second encoded bitstream as the bitstream of the video corresponding to the original video;

[0146] The second sending module 1706 is used to determine the indication information and send the indication information and the bit stream of the video, so that after receiving the bit stream of the video and the indication information, the decoder uses the other video decoding method described above to decode the video and obtain the target video. The indication information includes: first indication information, second indication information and third indication information. The first indication information is used to indicate the key reference frame, the second indication information is used to indicate the subsequent frame to be calibrated, and the third indication information is used to indicate the distribution parameters of each color channel of the key reference frame in the original video.

[0147] Using the apparatus described in Figure 17, key reference frames in the original video are encoded to obtain a first encoded bitstream; subsequent frames in the original video are represented using compact transmission symbols, and the representation results are encoded to obtain a second encoded bitstream; the first and second encoded bitstreams are determined as the bitstreams of the video corresponding to the original video; indication information is determined, and the indication information and the video bitstream are sent, so that the decoder, upon receiving the video bitstream and the indication information, uses the aforementioned alternative video decoding method to decode the video and obtain the target video. That is, by sending the encoded video bitstream and indication information to the decoder, the decoder can combine the indication information with a channel-by-channel calibration method, using the distribution parameters of each color channel of the key reference frame in the original video to calibrate the color distribution of the corresponding color channel in subsequent frames, thereby solving the technical problem of color shift in reconstructed video in related technologies and achieving the technical effect of improving the quality of reconstructed video.

[0148] The functions of each module in the apparatus of this embodiment can be found in the corresponding description in the above method, and they have corresponding beneficial effects, which will not be repeated here.

[0149] This disclosure provides a non-transitory computer-readable storage medium storing a bitstream of video, the non-transitory computer-readable storage medium being part of a computing device configured to execute a set of instructions to decode the bitstream according to the operations provided in this disclosure, or the computing device being configured to execute a set of instructions to encode the bitstream of video according to the operations provided in this disclosure.

[0150] This disclosure also provides a chip including a processor for calling and executing instructions stored in a memory, causing a communication device on which the chip is installed to perform the methods provided in this disclosure.

[0151] This disclosure also provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, output interface, processor, and memory are connected through an internal connection path. The processor is used to execute code in the memory. When the code is executed, the processor is used to execute the method provided in the application embodiment.

[0152] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting Advanced Reduced Instruction Set Machines (ARM) architecture.

[0153] Further, optionally, the aforementioned memory may include read-only memory and random access memory. The memory may be volatile memory or non-volatile memory, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available. Examples include Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0154] In the above embodiments, implementation can be achieved, in whole or in part, by software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to this disclosure are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.

[0155] In the description of this disclosure, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this disclosure, as well as the features of those different embodiments or examples.

[0156] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means two or more, unless otherwise explicitly specified.

[0157] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process. Furthermore, the scope of the preferred embodiments of this disclosure includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functionality involved.

[0158] The logic and / or steps described in the flowchart or otherwise herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).

[0159] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. All or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware, the program being stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiments.

[0160] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. This storage medium can be a read-only memory, a disk, or an optical disk, etc.

[0161] The above description is merely an exemplary embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope described in this disclosure, and these should all be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.

Claims

1. A video decoding method, applied to a generative video compression framework, comprising: After receiving the bitstream of the video and decoding it to obtain the reconstructed video, the distribution parameters of each color channel in the key reference frame of the reconstructed video are calculated, wherein the distribution parameters are used to characterize the color features of the corresponding color channel; For the same color channel, the color distribution of subsequent frames in the reconstructed video is calibrated using the distribution parameters of the key reference frame to obtain calibrated subsequent frames; The target video is generated based on the key reference frame and the calibrated subsequent frames.

2. The method according to claim 1, wherein, The distribution parameters include at least the mean and standard deviation of each color channel.

3. The method according to claim 2, wherein, For the same color channel, the color distribution of subsequent frames in the reconstructed video is calibrated using the distribution parameters of the key reference frame to obtain calibrated subsequent frames, including: Obtain the first mean and first standard deviation of the first color channel of the key reference frame; The color values ​​of each pixel in the second color channel of the subsequent frame are calibrated using the first mean and the first standard deviation to obtain the calibrated color distribution of the second color channel, wherein the first color channel and the second color channel are the same color channel; The color distributions of multiple color channels after calibration are combined to obtain the calibrated subsequent frames.

4. The method according to claim 3, wherein, The color values ​​of each pixel in the second color channel of the subsequent frames are calibrated using the first mean and the first standard deviation to obtain the calibrated color distribution of the second color channel, including: The initial color value of the target pixel is obtained from the second color channel of the subsequent frame, wherein the initial color value is used to characterize the color intensity of the target pixel in the second color channel; The product of the initial color value and the first standard deviation is determined to obtain the first result; Summing the first result and the first mean yields the second result; Set the second result as the color value of the target pixel after calibration.

5. The method according to claim 4, wherein, Before determining the product of the initial color value and the first standard deviation to obtain the first result, the method further includes: Obtain the second mean and second standard deviation of the second color channel in the subsequent frames; The color values ​​of each pixel in the second color channel of the subsequent frame are standardized using the second mean and the second standard deviation.

6. The method according to any one of claims 2 to 5, wherein, The mean and standard deviation of each color channel are determined as follows: The mean value is determined based on the color value of each pixel in the color channel and the total number of pixels. The standard deviation is determined based on the color value of each pixel in the color channel and the mean value.

7. The method according to any one of claims 1 to 6, wherein, The color channels include at least the red, green, and blue (RGB) channels.

8. The method according to any one of claims 1 to 6, wherein, The color channels also include at least one of the following: YCbCr channel, YUV channel, YPbPr channel, HSL channel, HSV channel, and Lab channel.

9. The method according to claim 8, wherein, The YCbCr channel includes a luminance Y channel, a blue chromaticity Cb channel, and a red chromaticity Cr channel; the YUV channel includes a luminance Y channel, a chromaticity U channel, and a chromaticity V channel.

10. The method according to any one of claims 1 to 9, wherein, The distribution parameters also include at least one of kurtosis, skewness, and frequency domain characteristics.

11. The method according to any one of claims 1 to 10, wherein, The video types include at least one of the following: facial videos, motion videos, conference and speech videos, educational videos, entertainment and media videos, gaming videos, user-generated content, surveillance videos, real-time communication videos, virtual reality videos, and augmented reality videos.

12. The method according to any one of claims 4 to 11, wherein, The target pixel is the pixel in the user's region of interest in the second color channel of the subsequent frame, and the method only performs color value calibration on the pixels in the user's region of interest.

13. The method according to any one of claims 4 to 12, wherein, The range of the initial color value is determined by the number of bits in the color channel; the initial color value range for an 8-bit color channel is 0 to 255, the initial color value range for a 16-bit color channel is 0 to 65535, and the initial color value range for a 32-bit floating-point color channel is 0.0 to 1.

0.

14. The method according to any one of claims 5 to 13, wherein, The color values ​​of each pixel in the second color channel of the subsequent frame are standardized using the second mean and the second standard deviation. Specifically, this includes: calculating the difference between the color value of the pixel and the second mean, and performing a ratio operation between the difference and the second standard deviation to obtain the standardized result of the pixel color value.

15. The method according to any one of claims 1 to 14, wherein, The calibration step is a post-processing step in generative video compression decoding, and the calibration process does not increase the bitrate.

16. The method according to any one of claims 1 to 15, wherein, The reconstructed video is obtained through a generative video compression framework, in which key reference frames are compressed using traditional coding methods, and subsequent frames are represented by compact symbols and reconstructed by a synthesis model and the decoded key reference frames.

17. A video coding method applied to a generative video compression framework, comprising: Encode the key reference frames in the original video to obtain the first encoded bitstream; Subsequent frames in the original video are represented using compact transmission symbols, and the representation results are encoded to obtain a second coded bitstream; The first encoded bitstream and the second encoded bitstream are determined as the bitstream of the video corresponding to the original video, and the bitstream of the video is sent.

18. A video decoding apparatus, applied to a generative video compression framework, comprising: The calculation module is used to receive the bit stream of the video and decode it to obtain the reconstructed video, and then calculate the distribution parameters of each color channel of the key reference frame in the reconstructed video, wherein the distribution parameters are used to characterize the color features of the corresponding color channel. The first calibration module is used to calibrate the color distribution of subsequent frames in the reconstructed video using the distribution parameters of the key reference frame for the same color channel, so as to obtain calibrated subsequent frames. The first generation module is used to generate a target video based on the key reference frame and the calibrated subsequent frames.

19. A video encoding apparatus, applied to a generative video compression framework, comprising: The first encoding module is used to encode the key reference frames in the original video to obtain the first encoded bitstream; The second encoding module is used to represent subsequent frames in the original video using compact transmission symbols and to encode the representation results to obtain a second encoded bitstream. The first sending module is used to determine the first encoded bitstream and the second encoded bitstream as the bitstream of the video corresponding to the original video, and send the bitstream of the video.

20. A non-transitory computer-readable storage medium storing a bitstream of original video, said non-transitory computer-readable storage medium being part of a computing device configured to execute a set of instructions to cause the computing device to process the original video according to operations including: The key reference frames in the original video are encoded to obtain a first coded bitstream; subsequent frames in the original video are represented using compact transmission symbols, and the representation results are encoded to obtain a second coded bitstream; the first coded bitstream and the second coded bitstream are determined as the bitstream of the video corresponding to the original video, and the bitstream of the video is sent.