Ai-based video conference using robust face repair with adaptive quality control
The video conferencing framework addresses instability in face reenactment by using a discrete codebook and adaptive learning to encode face features, achieving stable and high-quality face reconstruction at ultra-low bitrates.
Patent Information
- Application Number
- CN202380083317.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-02
- Filing Date
- 2023-11-30
- Publication Date
- 2025-07-15
AI Technical Summary
The existing video conferencing system is unstable due to changes in lighting, posture and expression during the replay of faces, and it is difficult to efficiently compress face details, resulting in artifacts and high computational complexity.
A discrete codebook-based representation method is adopted, combining general face priors and data dependence detail recovery, through the combination of general branches and adaptive branches, the combination weight is adjusted using the online adaptive learning mechanism to achieve high-quality face repair.
Achieve robust high-quality face reconstruction at extremely low bit rates, reducing artifacts, reducing computational complexity, providing flexible quality control and higher visual effects.
Smart Images

Figure CN120323017A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the benefit of priority to U.S. Application No. 63 / 429,574, filed on December 2, 2022, the entire content of which is incorporated herein by reference. Technical field
[0003] At least one embodiment of the present application generally relates to a method or apparatus for compressing and decompressing images and videos of video conferencing using human - centered video content. Background art
[0004] In recent years, video conferencing has developed rapidly and has become a daily communication means in work and life. Generally speaking, standard video codecs for compressing natural image / video data have been developed, such as AVC (which has been widely used in such video conferencing applications), HEVC, and VVC. In recent years, end - to - end learning image coding (LIC) or neural network (NN) - based video coding has also been developed. Summary of the invention
[0005] At least one of the embodiments of the present invention generally relates to a method or apparatus in the context of a video conferencing framework based on face restoration. Instead of using information from different source frames and driving frames, information such as pose, expression, identity, appearance, and texture from, for example, the current target frame is used, thereby avoiding the instability of previous video conferencing solutions based on face re - enactment. To provide a similar ultra - low bit rate, the proposed system uses a discrete codebook - based representation.
[0006] According to a first aspect, a method is provided. The method includes the following steps: determining at least one embedded feature of a video image; obtaining a codebook - based representation of the at least one embedded feature based on a codebook; resampling the video image to obtain a low - quality resampled video image; compressing the low - quality resampled video image to obtain a low - quality latent representation; and transmitting the codebook - based representation, the low - quality latent representation, and a combination weight.
[0007] According to a second aspect, a method is provided. The method includes the following steps: receiving a codebook - based representation, a low - quality latent representation of a video image, and a combination weight; extracting codewords corresponding to the codebook - based representation to form decoded embedded features; decoding the low - quality latent representation; calculating low - quality embedded features based on the decoded low - quality input; and reconstructing the video image based on the decoded embedded features, the low - quality embedded features, and the combination weight.
[0008] According to another aspect, an apparatus is provided. The apparatus includes a processor. The processor can be configured to implement the general aspect by executing any of the above - mentioned methods.
[0009] In another general aspect according to at least one embodiment, there is provided an apparatus comprising: means according to any of the decoding embodiments; and at least one of the following: (i) an antenna configured to receive a signal including video blocks; (ii) a band limiter configured to limit the received signal within a band including the video blocks; and (iii) a display configured to display an output representing the video blocks.
[0010] In another general aspect according to at least one embodiment, there is provided a non-transitory computer-readable medium containing data content generated according to any one of the encoding embodiments and variations.
[0011] In another general aspect according to at least one embodiment, there is provided a signal including video data generated according to any one of the encoding embodiments and variations.
[0012] In another general aspect according to at least one embodiment, a bitstream is formatted to include data content generated according to any one of the encoding embodiments and variations.
[0013] In another general aspect according to at least one embodiment, there is provided a computer program product including instructions that, when executed by a computer, cause the computer to perform any one of the decoding embodiments and variations.
[0014] These and other aspects, features, and advantages will become apparent by reading the detailed description of the exemplary embodiments in conjunction with the accompanying drawings.
[0015] In another general aspect according to at least one embodiment, there is provided a non-transitory computer-readable medium containing data content including instructions for performing any one of the encoding or decoding methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 An example of the general workflow of an AI (Artificial Intelligence)-based video conferencing system is shown.
[0017] Figure 2 A preferred embodiment of the workflow of the proposed AI-based encoding module is shown.
[0018] Figure 3 A preferred embodiment of the workflow of the proposed AI-based decoding module is shown.
[0019] Figure 4 A preferred embodiment of the workflow of the proposed online adaptive learning by simultaneously adjusting low quality and combined weights is shown.
[0020] Figure 5 Shows a preferred embodiment of the workflow of the proposed online adaptive learning by combining weights only.
[0021] Figure 6 Shows exemplary results of the original face image and the reconstructed face image.
[0022] Figure 7 Shows an exemplary application of the proposed method to a video conferencing scenario: (a) fewer bits and less realistic, but visually pleasing; (b) moderate bit transmission and somewhat close to a real face; (c) real face details at the maximum number of bits.
[0023] Figure 8 Shows an embodiment of a method for encoding a video using the said embodiment.
[0024] Figure 9 Shows an embodiment of a method for decoding a video using the said embodiment.
[0025] Figure 10 Shows an embodiment of an apparatus for encoding or decoding using the said embodiment.
[0026] Figure 11 Shows a standard general video compression scheme.
[0027] Figure 12 Shows a standard general video decompression scheme.
[0028] Figure 13 Shows a processor-based system for encoding / decoding based on the said general aspects. Detailed Description
[0029] Video codec tools in existing video codecs are designed to improve the codec efficiency for general image and video content, and some tools are specifically designed for screen content. They are not optimized for video conferencing scenarios.
[0030] In most cases, the human face becomes the main content of a video conference (e.g., one or more people are conversing at the center of the video frame). Since, from a structural perspective, face attributes have extensive commonalities among people, a general representation can be used to efficiently encode these features, and the number of bits required to transmit this general representation is much less than using an off-the-shelf codec to compress the original pixels. This enables the encoding framework to compress the human face at an extremely low bitrate and reconstruct the face with good quality. Recently, NVIDIA's Maxine solution applies the face reenactment algorithm to the video conference scenario. Briefly, a high-quality (HQ) source frame is encoded and transmitted to the decoder together with the facial keypoint representation in order to synthesize consecutive video frames at the receiving end. To reduce severe artifacts in practical applications, an enhancement algorithm in a previous method has been studied, which uses multiple source images and only reenacts the segmented face pixels.
[0031] Due to the significant differences between the source frame and the target frame in practical applications, the prior art of face reenactment-based methods is inherently unstable. The transmission of the facial keypoint representation is highly compact but only contains pose and expression information and cannot provide rich details for high-quality face generation. Transferring identity and appearance from the source frame to the target frame is inevitably prone to artifacts, especially when non-negligible changes occur in lighting, pose, expression, etc. Using multiple source frames can alleviate this problem, but at the cost of maintaining a large number of source frames and performing multiple reenactment processes in the decoder.
[0032] For a general video conference, given an input video frame set I1...I N , the encoder generates a compressed representation L i for each video frame I i , and this representation requires fewer bits to be sent to the decoder compared to the original input video frame I i . The decoder restores the output video frame according to the received compressed representation L i and the previously received L1...L i-1 . The goal is to minimize both the distortion (e.g., MSE or SSIM) and the bitrate R(L ). i
[0033] Figure 1 shows an example of a general video conference workflow based on AI. Each input frame I i is fed into the face detection module, and the face is detected Each face is I iA cropping region, which is defined by a bounding box that contains the detected face at the center and some expansion areas. For example, the region is centered on the center of the detected face, and the width and height of the bounding box are a times and b times the width and height of the face respectively (a ≥ 1, b ≥ 1). This scheme has no restrictions on the face detection method or how to crop the bounding box of the face region. In addition, it can also be decided to only consider some of the detected faces (for example, the largest face or the face at the center of the video frame). This scheme also has no restrictions on how many faces or which faces to consider.
[0034] Let B i represent the remaining background pixels in frame I i that are not included in any of the faces considered in the decision-making. The video conferencing system can process B i in different ways. For example, the optional encoding and decoding module can aggressively compress B i using traditional HEVC / VVC, LIC, or video encoding, and then send it to the decoder, where the decoded can be obtained. In some cases, such as when using a predefined virtual background, B i can be directly discarded. How to process the background pixels B i is beyond the scope of discussion of this scheme. Therefore, the optional processing flow for B i is marked with a dashed line.
[0035] For each face to be considered On the encoder side, the AI-based encoder calculates the corresponding latent representation The number of bits required to transmit this representation through the transmission module is usually small. On the decoder side, the recovered latent representation is also calculated Usually, the latent representation is further compressed in the transmission module before transmission, for example, by lossless arithmetic coding, and a corresponding decoding process is required in the transmission module to recover This scheme has no restrictions on the possible further compression and decoding methods of the latent representation. Based on the recovered latent representation The AI-based decoder reconstructs the output face In the case of providing the decoded background the output face is then merged with to generate the final reconstructed frame This scheme has no restrictions on how to merge with
[0036] Previous video conferencing solutions are based on the concept of face reenactment, i.e., transferring the facial movements of a driving face image to another source face image. Given video frames I1... I N , traditional HEVC / VVC, LIC, or video coding methods are used to send the faces in the first M (1 ≤ M < N) frames to the decoder at a high bitrate to ensure the quality of the decoded faces. These faces are called source features, which carry the appearance and texture information of the people in the conference session. For example, in one existing method M = 1, while in another scenario M > 1. The faces in the remaining frames are called driving faces. Facial landmark key points (such as the left and right eyes, nose, eyebrows, lips, etc.) are extracted from the source frames and the driving frames, which carry the pose and expression information of the people. Usually, some additional information, such as 3D head pose, is also calculated from the source frames and the driving frames. Then, for the face in the driving frame I l using the corresponding face in the source frame I i Based on the calculated 3D head pose and key points, a transfer function can be learned to transfer the pose and expression of the driving frame face to the source frame face and use a reenactment neural network to generate an output reenacted face Then, multiple reenacted faces using multiple source frame faces are combined by interpolation to obtain the final output face
[0037] Previous solutions exhibit serious flaws when applied to real faces in natural scenarios. First, artifacts are often inevitable because it is very difficult to generate objects such as real hair, teeth, accessories, etc. that cannot be described by facial key points. Applying the reenactment process only to tightly cropped or segmented face regions can reduce but not eliminate artifacts, and it will bring additional computational and transmission overhead. In addition, previous solutions are inherently unstable because the reenacted face depends on the appearance and texture information of the source frame and the pose and expression information from another driving frame. Due to large differences between the source frame and the target frame caused by changes in lighting, pose, expression, etc., the performance will be affected. By maintaining a large pool of candidate source frames and only selecting the frames that are most similar to the current target driving frame, this problem can be alleviated but not eliminated, at the cost of greatly increasing the decoding complexity because a large source frame pool needs to be maintained and the reenactment process needs to be executed multiple times in the decoder.
[0038] The described embodiment presents a novel video conferencing framework based on face restoration. All information (pose, expression, identity, appearance, and texture) is from the current target frame, rather than using information from different source frames and driving frames, thus avoiding the instability of previous face reenactment-based video conferencing solutions. To provide a similar ultra-low bitrate, the proposed system adopts a discrete codebook-based representation. The core idea is to combine general face prior learning with data-dependent detail recovery to achieve robust high-quality face restoration at an extremely low bitrate (e.g., peak signal-to-noise ratio (PSNR) of 33 dB and a bitrate of only 0.03 bpp). Data-dependent detail recovery also avoids the difficulties encountered in facial keypoint-based generation for hair, teeth, accessories, etc.
[0039] The proposed framework consists of two branches. The general branch generates and transmits an integer vector indicating the codeword index, and the decoder extracts rich high-quality codebook features from the integer vector based on the same codebook shared with the encoder. The baseline HQ face can be robustly restored using the HQ codebook-based features. The adaptive branch optionally uses a low-quality (LQ) low bitrate face input obtained by resizing the input and further aggressively compressing it through LIC to provide additional detail fidelity and expressive features. The LQ features and HQ features are weighted and combined for the final reconstruction, thus flexibly balancing the bitrate and recovery quality. For ultra-low bitrates, the system relies more on the HQ features by assigning lower weights to the LQ features to ensure an HQ face with less detail. At higher bitrates, better LQ features can be obtained, and larger weights can provide more detail and fidelity. We will describe resizing from the perspective of downsampling, but any resizing, including upsampling, can be applied to any embodiment.
[0040] Furthermore, the described embodiment also presents an online adaptive learning mechanism for adjusting the LQ input and combination weights for the adaptive branch on the encoder side during testing. Since video conferencing is a learning task targeting ground truth (GT) during the test phase, effective adaptation can be achieved by online adjusting the network input and combination weights through the direct stochastic gradient descent (SGD) algorithm, thus better reconstructing for each specific data without any overhead in transmission or decoding calculations.
[0041] Proposed baseline solution
[0042] Figure 2 Shows an embodiment of the workflow of an AI-based encoder. First, in the general branch, an input frame of size is provided to the system and k inThey are the height, width, and number of channels respectively. For example, for an RGB color image, k in = 3, for a grayscale image, k in = 1, for an RGB + depth image, k in = 4, and so on. The embedding module calculates the embedding features of size The embedding module is usually a neural network (NN), which consists of multiple computational layers, such as convolution, (non)-linear activation, normalization, attention, skip connections, resizing, etc. The embedding features The height and width w of the embedding features j i depend on the size of the input image and the network structure of the embedding module, and the number of feature channels k depends on the network structure of the embedding module. The encoder has a learnable codebook containing m codewords . Each codeword c l is represented as a k-dimensional feature vector. Then, the code generation module calculates the codebook-based representation based on the embedding features and the codebook Specifically, each element in is also a k-dimensional feature vector, which is mapped to the optimal codeword c that is closest to idx(u,v) (u, v):
[0043]
[0044] where is the distance between l and c can be approximated by the codeword index idx(u, v), and the embedding features can be represented by the approximate integer codebook-based representation which contains codeword indices. Compared with the original input , the number of bits required for transmitting this integer codebook-based representation is very small.
[0045] In the adaptive branch, the input is downsampled in the downsampling module at a ratio of s (e.g., 4 times along both the height and width) to obtain a low-quality For example, downsampling can be performed using a bicubic / bilinear filter, and this scheme has no restrictions on the downsampling method. Then, the encoding module encodes the low-quality For example, downsampling can be performed using a bicubic / bilinear filter, and this scheme has no restrictions on the downsampling method. Then, the encoding module encodes the low-quality Perform aggressive compression to compute a low-quality latent representation for transmission The encoding module can use various methods to compress the low quality For example, an NN-based LIC method can be used. Additionally, traditional video encoding tools such as H.265 / H.266 can also be used. In a preferred embodiment, the compression ratio is high, so the number of bits required for the low-quality latent representation is very small. This scheme has no restrictions on the specific method or compression settings used to compress the low quality
[0046] Finally, the codebook-based representation and the low-quality latent representation together constitute Figure 1 the latent representation in and send it to the decoder. At the same time, the combined weight W j i is also sent to the decoder to guide the decoding process.
[0047] Figure 3 Shows an embodiment of the workflow of an AI-based decoder. First, in the general branch, after receiving the codebook-based representation the feature extraction module extracts the corresponding codeword c (u,v) for each index idx(u,v) based on the same codebook as in the encoder to form the decoded embedded feature of size idx(u,v) In the adaptive branch, after receiving the low-quality latent representation the decoding module uses a decoding method corresponding to the encoding method used in the encoding module to decode the decoded low-quality input For example, an NN-based LIC method can be used. Additionally, any traditional image or video codec such as HEVC, VVC, etc. can also be used. Then, the LQ embedding module computes the low-quality embedded feature of size based on the decoded low-quality input The LQ embedding network is similar to the embedding module in the encoder and is typically a neural network (NN) containing layers such as convolution, non-linear activation, normalization, attention, skip connections, resizing, etc. This scheme has no restrictions on the network architecture of the LQ embedding module.
[0048] Given the decoded embedded feature the low-quality embedded feature and the combined weight Wj received from the encoder i , the reconstruction module computes the reconstruction output The reconstruction module can be composed of multiple computational layers, such as convolution, (non)-linear activation, normalization, attention, skip connection, resizing, etc. There are also various ways to combine the decoded embedded features and the low-quality embedded features (e.g., by concatenation, modulation, etc.). The combination weight W j i determines the importance of the low-quality embedded features when combined with . This scheme has no restrictions on the network architecture of the reconstruction module or the way of combination and . The combination weight W j i is sent from the encoder to the decoder. The encoder can determine the combination weight W j i in various ways. For example, in one embodiment, the best-performing W can be selected from a preset set of weights based on a target performance metric (e.g., rate-distortion trade-off) j i . W can be selected individually for each video frame j i , or the system can determine W based on the average performance metric of some video frames (e.g., the first few frames of a video conference session) j i , and then fix the selected weight for the remaining frames.
[0049] The proposed online solution
[0050] In a preferred embodiment of this scheme, an online adaptive learning mechanism is further proposed for automatically determining the combination weight W j i , and providing additional flexibility to dynamically improve video conference performance according to target requirements. The proposed online adaptive learning mechanism adjusts the combination weight W according to the target online loss during the inference process j i as well as the optional low-quality Online adaptation learning occurs on the encoder side, and the encoder side sends the online-adjusted combination weight W j i to the decoder. The decoding process is the same as Figure 1 and Figure 3 because the determination of the combination weight W j i in the encoder does not change the processing flow in the decoder. Figure 4 and Figure 5 give two preferred embodiments of the workflow of online adaptive learning, where Figure 4 the combination weight W is adjusted simultaneouslyj i and low quality while Figure 5 only adjust the combination weight W j i .
[0051] Specifically, during the online adaptive learning process, the system first performs, based on the input and the initial combination weight W j i the encoding and decoding processes described by Figure 1 , Figure 2 and Figure 3 to obtain the decoded embedded features low quality low quality latent representation and the reconstruction output The system keeps the decoded embedded features unchanged. Then, the loss module calculates the online loss based on the reconstruction output the original input and the low quality latent representation calculate the online loss For example, the rate - distortion trade - off loss can be used:
[0052]
[0053] where represents and the distortion between (e.g., MSE, SSIM, perceptual loss like LIPIPS, or a weighted combination of these losses). is the rate - distortion loss, representing the bit consumption of the low quality latent representation (e.g., the entropy - likelihood estimated by the first prior method). This loss is differentiable, and the online SGD module calculates the gradients of the online loss with respect to the weight W j i and the gradients of the online loss with respect to the low quality The gradients are back - propagated to update the combination weight and the low quality
[0054]
[0055] where t is the index of the current iteration. If there are a total of T iterations, then t = 1,..., T. α and β are step - sizes for online adaptation, which can be preset as hyper - parameters empirically, or determined dynamically by searching some different settings, similar to the initial combination weight Wj i (0). There are no restrictions on how to set the hyperparameters in this scheme.
[0056] Finally, after T times of online update iterations, the updated is used to recalculate the low-quality latent representation and combine it with Figure 2 the updated W in j i (T) and the codebook-based representation and send them together to the decoder.
[0057] Note that in order to make the online loss differentiable with respect to the low-quality input for calculating the gradient Figure 2 and Figure 3 the method used in the encoding and decoding modules to compress the low-quality input is the NN-based LIC method. In contrast, Figure 5 describes the preferred online adaptive learning workflow, where only the combined weight W j i is adjusted. In this case, the encoding and decoding modules can use non-differentiable traditional video codecs such as HEVC / VVC.
[0058] Similar to the case of Figure 4 in the online adaptive learning process, Figure 5 the system in first performs the encoding and decoding processes as shown in j i such as Figure 1 , Figure 2 and Figure 3 based on the input low-quality embedded features and the reconstructed output The system keeps the decoded embedded features and the low-quality embedded features unchanged. Then, the loss module calculates the online loss based on the reconstructed output and the original input For example, represents and the distortion between (e.g., MSE, SSIM, perceptual loss similar to LIPIPS, or a weighted combination of these losses). This loss is differentiable, and the online stochastic gradient descent (SGD) module calculates the online loss with respect to the weight W j i, the gradient of and backpropagate it to update the combination weights:
[0059]
[0060] where t is the index of the current iteration. If a total of T iterations are performed, then t = 1,..., T. α is the step size for online adaptation, which can be preset as a hyperparameter based on experience or dynamically determined by searching for some different settings, similar to the initial combination weight W j i (0). This scheme has no restrictions on the setting of hyperparameters.
[0061] Finally, after T online update iterations, as Figure 5 shown, the updated W j i (T) and Figure 2 the low-quality latent representation described in and the codebook-based representation are sent to the decoder together.
[0062] It is worth mentioning that, in some embodiments, the entire adaptation branch can be selectively skipped, where the combination weight W j i is set to W j i = 0, and the reconstruction module reconstructs the output only based on the decoded embedded features to
[0063] Exemplary training process
[0064] The training process learns the learnable codebook the embedding network parameters and the reconstruction network parameters. In addition, when the encoding module and the decoding module use NN-based LIC, or the downsampling module uses an NN-based method, the corresponding network parameters are also learned during the training process. In a preferred embodiment, different network modules are trained in several different stages. For example, in the first stage, high-quality face inputs are used to train the learnable codebook the embedding network parameters and the reconstruction network parameters from the general branch in an end-to-end manner, where the training objective is to minimize the reconstruction distortion between the reconstructed output and the input . Multiple distortion loss functions can be used, such as MSE, MSSSIM, perceptual LPIPS, etc., or a weighted combination of different loss functions. A generative adversarial network (GAN) training strategy can be used to improve the quality of the learned codebook to achieve a visually pleasing reconstruction effect.
[0065] Then, in the second stage, the encoding and decoding modules in the adaptive branch are trained in an end-to-end manner. For example, first, a general image dataset with various image qualities is used, and then low-quality face images are used to fine-tune the learned parameters. The training objective is to minimize the rate-distortion trade-off loss of the reconstructed output and the rate-distortion loss of the latent representation , similar to Equation (2). The training method described in the first existing solution can be used here.
[0066] Then, in the third stage, a training dataset similar to the real video conferencing test data is used to train the LQ embedding module and fine-tune the reconstruction module in an end-to-end manner, where all other learned network parameters and the learned codebook are fixed.
[0067] In other embodiments, other training strategies can be adopted. For example, other training stages can be used, and in each stage, different modules can be trained or fine-tuned based on different sets of losses. Alternatively, the entire network can be trained end-to-end in one stage. This solution has no restrictions on the training process.
[0068] Some differences from the existing solutions
[0069] Video conferencing solution based on robust face restoration with ultra-low bitrate and excellent visual quality
[0070] The proposed pipeline that combines the general branch and the adaptive branch for efficient video conferencing based on face restoration is novel. The general branch uses an efficient discrete codebook representation to ensure baseline high-quality face reenactment. The adaptive branch provides additional fidelity and expressive details by transmitting low-quality, low-bitrate face images.
[0071] Flexible quality control for video conferencing
[0072] The proposed solution implements a flexible quality control function for video conferencing. The HQ features of the general branch and the LQ features of the adaptive branch are combined with weights, where the combination weights can be adjusted during testing to balance the bitrate and the reconstruction quality. The combination weights can be set manually or automatically.
[0073] Flexible online adaptive quality control for video conferencing
[0074] The proposed solution provides a mechanism that automatically adjusts the LQ face image and its corresponding combination weights for each video frame according to actual needs. This realizes a flexible online adaptive quality control function, where users can adjust the LQ face image and the combination weights according to different quality metrics and different bitrate and quality trade-offs.
[0075] Advantages over existing solutions
[0076] Since the learned high-quality codebook contains the learned high-quality face prior, the reconstructed face is even more visually pleasing than the original input. An example is shown as Figure 6 shown.
[0077] Quality control flexibly meets various requirements during testing.
[0078] Compared with previous AI-based video conferencing solutions based on face reenactment, it provides a more robust and higher-quality (HQ) video conferencing experience.
[0079] The framework flexibly adopts different network architectures for each network module component.
[0080] Flexibly adapts to various encoding and decoding methods in the adaptive branch, including NN-based or traditional codecs.
[0081] Video conferencing has become a major tool for communication in people's daily work and life. This application is crucial for all companies involved in cloud services and terminal devices, such as Apple, Amazon, Google, Tencent, Alibaba, OPPO, Huawei, Zoom, Microsoft, Nvidia, etc.
[0082] Figure 7 Shows an exemplary application with multiple scenarios. Device 700 captures the face region and compresses it using our solution. The captured real input image can be displayed on the display device 720 of the sender. Any type of quality controllable interface 740 can control the bits used for encoding the face to some extent, or control the fidelity of the face to be transmitted at the receiving device 710. The quality control mechanism can vary, but as a simple example, user 760 can use the human-machine interface panel 740 on device 700 to control the quality of the face to be displayed at the receiving display device 730. The less realistic the face, the fewer bits required for the proposed compression method. Figure 7 (a) Shows the case of selecting an option for face encoding with lower fidelity but lower bit rate. In this case, the general branch can only activate the encoding of the face region by mapping the embedded features to the optimized codewords. Then, at the receiver 710, the general branch decoder remaps the codewords to the embedded features and generates a reenacted face image. In this case, the reenacted face image may be more visually pleasing, but has lower fidelity to the real input face because the transmitted information does not provide any details of the face texture. Therefore, the overall facial expression of the synthesized face may depend on the training dataset used to construct the codewords.
[0083] Figure 7(b) shows the scenario 860 where the user selects medium fidelity and medium transmitted bits. In this case, some real face textures can be transmitted to the receiving device 810 by activating the proposed adaptive branch. In addition to the minimum information required to present the face through the general branch, the proposed adaptive branch further compresses the missing details at a lower resolution. Depending on the downsampling rate or quantization level used to encode the compressed representation of the input, the facial quality reproduced on the receiving side may vary.
[0084] Finally, in Figure 7 (c), when the user selects the highest fidelity 960, the reconstructed face on the receiving side 910 looks closer to a real face compared to the previous scenarios. To make the face more realistic, the proposed adaptive branch compresses the input with a lower quantization step to reconstruct the details of the face texture after decoding.
[0085] The proposed solution includes an encoder and a decoder and implements new functions that cannot be achieved by the prior art. Its detectability is obvious. The bitstream also requires relevant syntax information to implement the proposed adaptive quality control.
[0086] Figure 8 An embodiment of a method 800 for encoding / decoding video data is shown. The method starts at start box 801 and then proceeds to box 810 to determine at least one embedded feature of the video image. Control proceeds from box 810 to box 820 to determine a codebook-based representation of obtaining at least one embedded feature based on the codebook. The control flow proceeds from box 820 to box 830 to resample the video image to obtain a low-quality resampled video image. The control flow proceeds from box 830 to box 840 to compress the low-quality resampled video image to obtain a low-quality latent representation. The control flow proceeds from box 840 to box 850 to transmit the codebook-based representation, the low-quality latent representation, and the combined weights.
[0087] Any mention of resampling in this specification includes downsampling, upsampling, or sampling remaining unchanged.
[0088] Figure 9One embodiment of a method 900 for decoding video data is shown. The method starts at a start block 901 and then proceeds to block 910 to receive a codebook-based representation, a low-quality potential representation of a video image, and a combination weight. Control proceeds from block 910 to block 920 to extract a codeword corresponding to the codebook-based representation to form a decoded embedded feature. Control proceeds from block 920 to block 930 to decode the low-quality potential representation. Control proceeds from block 930 to block 940 to calculate a low-quality embedded feature based on the decoded low-quality input. Control proceeds from block 940 to block 950 to reconstruct a video image based on the decoded embedded feature, the low-quality embedded feature, and the combination weight.
[0089] Figure 10 An embodiment of an apparatus 1000 for compressing, encoding or decoding video using the above method is shown. The apparatus includes a processor 1010 and can be interconnected with a memory 1020 through at least one port. The processor 1010 and the memory 1020 can also have one or more additional interconnections with external connections.
[0090] The processor 1010 is also configured to insert or receive information in a bit stream and perform compression, encoding or decoding using the above-mentioned methods.
[0091] The embodiments described herein cover a variety of aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are specific and are usually described in a manner that may sound restrictive, at least to show the various features. However, this is for the purpose of describing clearly and does not limit the application or scope of these aspects. In fact, all different aspects can be combined and interchanged to provide further aspects. In addition, these aspects can also be combined and interchanged with aspects described in previous application documents.
[0092] The aspects described and contemplated in this application can be implemented in many different forms. Figure 11 , Figure 12 and Figure 13 Some embodiments are provided, but other embodiments are contemplated, and Figure 11 , Figure 12 and Figure 13 The discussion does not limit the breadth of implementation. At least one aspect generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects can be implemented as methods, apparatus, computer-readable storage media storing instructions for encoding or decoding video data according to any of the methods described, and / or computer-readable storage media storing a bitstream generated according to any of the methods described.
[0093] In this application, the terms "reconstruction" and "decoding" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image", "picture", and "frame" may be used interchangeably. Generally, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" or "reconstruction" is used on the decoder side.
[0094] This document describes various methods, and each method includes one or more steps or actions for implementing the method. Unless the correct operation of the method requires steps or actions in a specific order, the order and / or use of specific steps and / or actions can be modified or combined. Additionally, terms such as "first", "second", etc. may be used in different embodiments to modify elements, components, steps, operations, etc., such as "first decoding" and "second decoding". Unless specifically required, the use of such terms does not imply an order for the modified operations. Thus, in this example, the first decoding does not have to be performed before the second decoding, and it can occur, for example, before, during, or within a time period overlapping with the second decoding.
[0095] The various methods and other aspects described in this application can be used to modify the modules of the video encoder 100 and video decoder 200 as shown in Figure 11 and Figure 12 such as the intra prediction, entropy encoding, and / or decoding modules (160, 360, 145, 330). Additionally, aspects of this application are not limited to VVC or HEVC, and can be applied, for example, to other standards and recommendations (whether existing or future developed) and extensions of any such standards and recommendations (including VVC and HEVC). Unless otherwise stated or technically excluded, the aspects described in this application can be used alone or in combination.
[0096] A variety of numerical values are used in this application. The specific values are for illustrative purposes only, and the aspects are not limited to these specific values.
[0097] Figure 11 An encoder 100 is shown. Various variants of the encoder 100 are envisioned, but for clarity, only the encoder 100 is described below, and not all expected variants.
[0098] Before encoding, the video sequence may undergo pre - encoding processing (101), for example, applying a color transformation to the input color picture (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0), or remapping the input picture components to make the signal distribution more compression - resistant (e.g., performing histogram equalization on one of the color components). Metadata may be associated with the pre - processing and appended to the bitstream.
[0099] In encoder 100, pictures are encoded by encoder elements as described below. The image to be encoded is segmented (102) and processed, for example, in units of coding units (CUs). Each unit is encoded, for example, using an intra-frame or inter-frame mode. When a unit is encoded in the intra-frame mode, it performs intra-frame prediction (160). In the inter-frame mode, motion estimation (175) and motion compensation (170) are performed. The encoder decides (105) which of the intra-frame mode or inter-frame mode to use for encoding the unit and indicates the intra-frame / inter-frame decision, for example, by a prediction mode flag. The prediction residual is calculated, for example, by subtracting (110) the prediction block from the original image block.
[0100] The prediction residual is then transformed (125) and quantized (130). The quantized transform coefficients, motion vectors, and other syntax elements are entropy encoded (145) to output a bitstream. The encoder may skip the transformation and directly quantize the untransformed residual signal. The encoder may also bypass the transformation and quantization, that is, directly encode the residual without applying the transformation or quantization process.
[0101] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transform coefficients are dequantized (140) and inverse-transformed (150) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (155) to reconstruct the image block. A loop filter (165) is applied to the reconstructed image, for example, performing deblocking / sample adaptive offset (SAO) filtering, to reduce encoding artifacts. The filtered image is stored in a reference picture buffer (180).
[0102] Figure 12 A block diagram of video decoder 200 is shown. In decoder 200, the bitstream is decoded by decoder elements as described below. Video decoder 200 generally performs a decoding process opposite to the encoding process described in Figure 11 . Encoder 100 generally also performs video decoding as part of video data encoding.
[0103] Specifically, the input to the decoder includes a video bitstream, which may be generated by video encoder 100. First, the bitstream is entropy decoded (230) to obtain transform coefficients, motion vectors, and other encoded information. The picture partitioning information indicates how the picture is partitioned. Thus, the decoder can partition the picture (235) according to the decoded picture partitioning information. The transform coefficients are dequantized (240) and inverse-transformed (250) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (255) to reconstruct the image block. The prediction block can be obtained (270) by intra-frame prediction (260) or motion compensation prediction (i.e., inter-frame prediction) (275). A loop filter (265) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (280).
[0104] The decoded image can be further post-processed (285) after decoding, such as inverse color transformation (e.g., from YcbCr 4:2:0 to RGB 4:4:4) or inverse remapping, which performs the inverse process of the remapping process performed in the precoding process (101). The post-processing after decoding can use the metadata obtained in the precoding process and signaled in the bitstream.
[0105] Figure 13 A block diagram illustrating an example of a system implementing various aspects and embodiments is shown. System 1000 can be embodied as a device including various components described below and configured to perform one or more aspects described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptops, smartphones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, networked household appliances, and servers. The elements of System 1000 can be embodied in a single integrated circuit (IC), multiple ICs, and / or discrete components, either individually or in combination. For example, in at least one embodiment, the processing and encoder / decoder elements of System 1000 are distributed across multiple ICs and / or discrete components. In various embodiments, System 1000 is communicatively coupled to one or more other systems or other electronic devices, e.g., via a communication bus or through dedicated input and / or output ports. In various embodiments, System 1000 is configured to implement one or more aspects described herein.
[0106] System 1000 includes at least one processor 1010 configured to execute instructions loaded therein to implement various aspects such as those described in the present application. Processor 1010 can include embedded memory, input / output interfaces, and various other circuits known in the art. System 1000 includes at least one memory 1020 (e.g., volatile memory devices and / or non-volatile memory devices). System 1000 includes a storage device 1040, which can include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 1040 can include internal storage devices, additional storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.
[0107] System 1000 includes an encoder / decoder module 1030, which is configured, for example, to process data to provide encoded video or decoded video, and the encoder / decoder module 1030 may include its own processor and memory. The encoder / decoder module 1030 represents a module that may be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of an encoding and a decoding module. Additionally, the encoder / decoder module 1030 may be implemented as a separate element of system 1000 or may be incorporated into the processor 1010 as a combination of hardware and software known to those skilled in the art.
[0108] Program code loaded onto the processor 1010 or the encoder / decoder 1030 to perform the various aspects described herein may be stored in the storage device 1040 and subsequently loaded onto the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, the memory 1020, the storage device 1040, and the encoder / decoder module 1030 may store one or more of a variety of items during the execution of the processes described herein. These stored items may include, but are not limited to, input video, decoded video or a portion of decoded video, bitstreams, matrices, variables, and intermediate or final results from equations, formulas, operations, and operation logic processing.
[0109] In some embodiments, the memory internal to the processor 1010 and / or the encoder / decoder module 1030 is used to store instructions and provide the processing working memory required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory may be the memory 1020 and / or the storage device 1040, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, fast external dynamic volatile memory (e.g., RAM) is used as the working memory for video encoding and decoding operations (e.g., MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, 13818-1 is also known as H.222, 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard being developed by the Joint Video Exploration Team (JVET))).
[0110] Inputs to the components of system 1000 may be provided through a variety of input devices as shown in block 1130. Such input devices include, but are not limited to: (i) a radio frequency (RF) section that receives, for example, RF signals transmitted by a broadcaster over the air; (ii) component (COMP) input terminals (or a set of COMP input terminals); (iii) universal serial bus (USB) input terminals; and / or (iv) high-definition multimedia interface (HDMI) input terminals. Figure 13 Other examples not shown include composite video.
[0111] In various embodiments, the input devices of block 1130 have associated input processing elements known in the art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or limiting the signal bandwidth to a frequency band), (ii) downconverting the selected signal, (iii) again limiting the bandwidth to a narrower frequency band to select, for example, a signal band (which may be referred to as a channel in some embodiments), (iv) demodulating the downconverted and bandwidth-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements to perform these functions, such as including a frequency selector, a signal selector, a bandwidth limiter, a channel selector, filters, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs multiple of these functions, such as downconverting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted through a wired (e.g., cable) medium and perform frequency selection by filtering, downconverting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above (and other) elements, remove some of them, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting an amplifier and an analog / digital converter. In various embodiments, the RF section includes an antenna.
[0112] In addition, the USB and / or HDMI terminals may include respective interface processors for connecting the system 1000 to other electronic devices via the USB and / or HDMI connections. It should be understood that various aspects of input processing (such as Reed-Solomon error correction) may be implemented within a separate input processing IC or within the processor 1010 as needed. Similarly, aspects of USB or HDMI interface processing may be implemented within a separate interface IC or within the processor 1010 as needed. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements, such as including the processor 1010 and the encoder / decoder 1030, which operate in conjunction with memory and storage elements to process the data stream as needed for presentation on an output device.
[0113] Various elements of the system 1000 may be provided within an integrated housing. Within the integrated housing, the various elements may be interconnected using suitable connection means and data may be transferred between them, such connection means being, for example, internal buses (including inter-integrated circuit (I2C) buses), wiring, and printed circuit boards known in the art.
[0114] The system 1000 includes a communication interface 1050 that permits communication with other devices via a communication channel 1060. The communication interface 1050 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 1060. The communication interface 1050 may include, but is not limited to, a modem or a network card, and the communication channel 1060 may be implemented, for example, within a wired and / or wireless medium.
[0115] In various embodiments, data is streamed or otherwise provided to the system 1000 using a wireless network such as a Wi-Fi network (e.g., IEEE 802.11, where IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signals of these embodiments are received via the communication channel 1060 and the communication interface 1050 adapted for Wi-Fi communication. The communication channel 1060 of these embodiments is typically connected to an access point or router that provides access to an external network including the Internet to permit streaming applications and other over-the-top communications. Other embodiments use a set-top box to provide streaming data to the system 1000, and the set-top box transfers data via the HDMI connection of the input box 1130. Other embodiments use the RF connection of the input box 1130 to provide streaming data to the system 1000. As described above, various embodiments provide data in a non-streaming manner. In addition, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0116] System 1000 can provide output signals to a variety of output devices, including display 1100, speaker 1110, and other peripheral devices 1120. Display 1100 of various embodiments includes one or more of a touchscreen display, an organic light emitting diode (OLED) display, a curved display, and / or a foldable display. Display 1100 can be used for a television, a tablet computer, a laptop computer, a mobile phone (cellular phone), or another device. Display 1100 can also be integrated with other components (e.g., as in a smartphone) or be standalone (e.g., an external display for a laptop computer). In an example of various embodiments, other peripheral devices 1120 include one or more of a standalone digital video disc (or digital versatile disc) (DVR, for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 1120 that provide functions based on the output of system 1000. For example, a disc player performs the function of playing the output of system 1000.
[0117] In various embodiments, control signals are communicated between system 1000 and display 1100, speaker 1110, or other peripheral devices 1120 using signals such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that allow device-to-device control with or without user intervention. The output devices can be communicatively coupled to system 1000 via dedicated connections through their respective interfaces 1070, 1080, and 1090. Alternatively, the output devices can be connected to system 1000 using communication channel 1060 through communication interface 1050. Display 1100 and speaker 1110 can be integrated into a single unit in an electronic device (e.g., a television) with other components of system 1000. In various embodiments, display interface 1070 includes a display driver, such as a timing controller (T Con) chip.
[0118] Display 1100 and speaker 1110 can alternatively be separated from one or more other components, e.g., if the RF portion of input 1130 is part of a separate set-top box. In various embodiments where display 1100 and speaker 1110 are external components, the output signals can be provided through a dedicated output connection, such as including an HDMI port, a USB port, or a COMP output.
[0119] Embodiments may be implemented by computer software executed by a processor 1010, or by hardware, or by a combination of hardware and software. As a non-limiting example, embodiments may be implemented by one or more integrated circuits. The memory 1020 may be of any type suitable for the technical environment and may be implemented using any suitable data storage technology, such as, by way of non-limiting example, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 1010 may be of any type suitable for the technical environment and may include one or more microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures (as non-limiting examples).
[0120] Multiple implementations involve decoding. "Decoding" as used in this application may cover, for example, all or part of the process of performing on a received coded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such a process also includes or alternatively includes processes performed by the decoders of the various implementations described in this application.
[0121] As a further example, in an embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Based on the context of the specific description, whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally refer to a broader decoding process will be apparent, and it is believed that those skilled in the art can well understand it.
[0122] Multiple implementations involve encoding. Similar to the above discussion regarding "decoding", "encoding" as used in this application may cover, for example, all or part of the process of performing on an input video sequence to produce a coded bitstream. In various embodiments, such processes include one or more processes typically performed by an encoder, such as partitioning, differential encoding, transform, quantization, and entropy encoding. In various embodiments, such processes also include or alternatively include processes performed by the encoders of the various implementations described in this application.
[0123] As a further example, in an embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally refer to a broader encoding process will be clear based on the context of the specific description, and it is believed that those skilled in the art can well understand it.
[0124] It should be noted that the grammatical elements used herein are descriptive terms. Therefore, they do not exclude the use of other grammatical element names.
[0125] When a figure is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding device. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.
[0126] Multiple embodiments may involve parametric models or rate-distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is typically considered, usually in view of limitations in computational complexity. It can be measured by a rate-distortion optimization (RDO) metric, minimum mean square (LMS), mean absolute error (MAE), or other similar measures. Rate-distortion optimization is typically formulated as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. There are different ways to solve the rate-distortion optimization problem. For example, these methods can be based on extensive testing of all coding options, including all considered modes or coding parameter values, and a complete evaluation of their coding costs and the associated distortion of the reconstructed signal after encoding and decoding. Faster methods can also be used to save coding complexity, especially by calculating an approximate distortion based on the predicted or prediction residual signal rather than the reconstructed signal. These two methods can also be used in combination, for example, using approximate distortion only for certain possible coding options and full distortion for other coding options. Other methods only evaluate a subset of the possible coding options. More generally, many methods employ any of a variety of techniques to perform the optimization, but the optimization does not necessarily involve a complete evaluation of coding costs and associated distortion.
[0127] The embodiments and aspects described herein can be implemented as, for example, a method or process, apparatus, software program, data stream, or signal. Even if discussed only in the context of a single form of implementation (e.g., only as a method), the implementation of the features discussed can also be implemented in other forms (e.g., a device or program). The apparatus can be implemented as appropriate hardware, software, and firmware. A method can be implemented, for example, as a processor, which generally refers to a processing device, such as a computer, microprocessor, integrated circuit, or programmable logic device. The processor also includes communication devices, such as computers, mobile phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.
[0128] Reference to “one embodiment” or “an embodiment” or “one implementation” or “an implementation” and other variations thereof means that the specific features, structures, characteristics, etc. associated with that embodiment are included in at least one embodiment. Thus, the phrases “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation” and any other variations thereof that appear in multiple places in this application do not necessarily all refer to the same embodiment.
[0129] In addition, the present application may be related to "determining" various information. Determining information may include one or more of the following, such as estimated information, calculated information, predicted information, or retrieving information from a memory.
[0130] In addition, the present application may be related to "accessing" various information. Accessing information may include one or more of the following, such as receiving information, retrieving information (e.g., retrieving from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0131] In addition, the present application may be related to "receiving" various information. Similar to "accessing", receiving is a broad term. Receiving information may include one or more of the following, such as accessing information or retrieving information (e.g., from a memory). In addition, "receiving" generally involves, in some way, during operations such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0132] It should be understood that the use of any of the following, " / ", "and / or", and "at least one", for example, in the cases of "A / B", "A and / or B", and "at least one of A and B", is intended to cover only selecting the first-listed option (A), or only selecting the second-listed option (B), or selecting both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to cover only selecting the first-listed option (A), or only selecting the second-listed option (B), or only selecting the third-listed option (C), or only selecting the first and second-listed options (A and B), or only selecting the first and third-listed options (A and C), or only selecting the second and third-listed options (B and C), or selecting all three options (A and B and C). For those of ordinary skill in the art and related fields, this can be extended to as many items as are listed.
[0133] In addition, as used herein, the term "signal" refers, among other things, to indicating something to a corresponding decoder. For example, in some embodiments, the encoder signals a particular one of a plurality of transforms, codec modes, or flags. Thus, in an embodiment, the same transform, parameters, or mode are used on the encoder side and the decoder side. For example, the encoder can transmit (explicitly signal) a particular parameter to the decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has the particular parameter as well as other parameters, signaling (implicitly signaling) can be done without transmission to simply allow the decoder to know and select the particular parameter. By avoiding the transmission of any actual functionality, bit savings are achieved in a variety of embodiments. It should be understood that signaling can be done in a variety of ways. For example, in a variety of embodiments, one or more syntax elements, flags, etc. are used to signal information to the corresponding decoder. Although the foregoing relates to the verb form of the word "signal", the word "signal" can also be used as a noun herein.
[0134] As will be appreciated by one of ordinary skill in the art, implementations can produce a variety of formatted signals to carry information, for example, that can be stored or transmitted. The information can include, for example, instructions for performing a method or data produced by one of the implementations. For example, a signal can be formatted to carry an encoded video stream and SEI messages of the embodiments. Such a signal can, for example, be formatted as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. Formatting can include, for example, encoding the video stream and modulating a carrier with the encoded video stream. The information carried by the signal can be, for example, analog or digital information. As is well known, signals can be transmitted over a variety of different wired or wireless links. Signals can be stored on a processor-readable medium.
[0135] The foregoing sections describe a number of embodiments that cover various claim categories and types. The features of these embodiments can be provided individually or in any combination. In addition, embodiments can individually or in any combination include one or more of the following features, devices, or aspects that cover various claim categories and types:
[0136] We have described a number of embodiments that cover various claim categories and types. The features of these embodiments can be provided individually or in any combination. In addition, embodiments can individually or in any combination cover various claim categories and types, including one or more of the following features, devices, or aspects, including a first device that includes a user equipment and a second device that includes a network.
[0137] A video conferencing framework based on face restoration, where all information comes from the current target frame.
[0138] In one embodiment, there is an adaptive online learning mechanism that adjusts low-quality inputs and / or the combination weights used to combine low-quality and high-quality portions.
[0139] The present disclosure contemplates creating and / or transmitting and / or receiving and / or decoding a bitstream or signal that includes one or more of the syntax elements or variants thereof.
[0140] In one embodiment, a television, set-top box, mobile phone, tablet computer, or other electronic device performs the transformation method according to any of the embodiments.
[0141] In one embodiment, a television, set-top box, mobile phone, tablet computer, or other electronic device performs the transformation method determined according to any of the embodiments and displays (e.g., using a display, screen, or other type of display) the resulting image.
[0142] In one embodiment, a television, set-top box, mobile phone, tablet computer, or other electronic device selects a channel, limits the bandwidth, or tunes (e.g., using a tuner) to receive a signal including an encoded image and performs the transformation method according to any of the embodiments.
[0143] In one embodiment, a television, set-top box, mobile phone, tablet computer, or other electronic device wirelessly receives (e.g., using an antenna) a signal including an encoded image and performs the transformation method.
Claims
1. A method, comprising: Determining at least one embedded feature of a video image; Obtaining a codebook-based representation of the at least one embedded feature based on a codebook; Resampling the video image to obtain a low-quality resampled video image; Compressing the low-quality resampled video image to obtain a low-quality latent representation; And Transmitting the codebook-based representation, the low-quality latent representation, and a combination weight.
2. An apparatus, comprising a memory and a processor, the processor configured to execute: Determining at least one embedded feature of a video image; Obtaining a codebook-based representation of the at least one embedded feature based on a codebook; Resampling the video image to obtain a low-quality resampled video image; Compressing the low-quality resampled video image to obtain a low-quality latent representation; And Transmitting the codebook-based representation, the low-quality latent representation, and a combination weight.
3. A method, comprising: Receiving a codebook-based representation, a low-quality latent representation of a video image, and a combination weight; Extracting codewords corresponding to the codebook-based representation to form decoded embedded features; Decoding the low-quality latent representation; Calculating low-quality embedded features based on the decoded low-quality input; And Reconstructing the video image based on the decoded embedded features, the low-quality embedded features, and the combination weight.
4. An apparatus, comprising a memory and a processor, the processor configured to execute: Receiving a codebook-based representation, a low-quality latent representation of a video image, and a combination weight; Extracting codewords corresponding to the codebook-based representation to form decoded embedded features; Decoding the low-quality latent representation; Calculating low-quality embedded features based on the decoded low-quality input; And Reconstructing the video image based on the decoded embedded features, the low-quality embedded features, and the combination weight.
5. The method according to claim 1, or the apparatus according to claim 2, wherein the combination weight is adjusted during the inference process.
6. The method according to claim 1, or the apparatus according to claim 2, wherein the low-quality latent representation is adjusted during the inference process.
7. The method according to claim 1, or the apparatus according to claim 2, wherein the combination weight and the low-quality latent representation are updated by the gradient of the online loss with respect to the combination weight through backpropagation.
8. The method according to claim 1, or the apparatus according to claim 2, wherein the update of the low-quality input is used to recalculate the low-quality latent representation and is sent to the decoder.
9. The method according to claim 1, or the apparatus according to claim 2, wherein the step size for online adaptation is set by hyperparameters or determined dynamically.
10. The method according to claim 1, or the apparatus according to claim 2, wherein the combination weight is determined by a preset set of weights.
11. A device, comprising: The apparatus according to claim 4; And At least one of the following: (i) an antenna configured to receive a signal including a video block; (ii) a band limiter configured to limit the received signal to a band including the video block; and (iii) a display configured to display an output representative of the video block.
12. A non-transitory computer-readable medium comprising data content generated by or according to the method of any one of claims 1 and 5 to 10 or by the apparatus of any one of claims 2 and 5 to 10 for playback using a processor.
13. A signal comprising video data generated by or according to the method of any one of claims 1 and 5 to 10 or by the apparatus of any one of claims 2 and 5 to 10 for playback using a processor.
14. A computer program product comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1, 3, and 5 to 10.
15. A non-transitory computer-readable medium comprising data content including instructions for performing the method of any one of claims 1, 3, and 5 to 10.