Method and apparatus for talking face video compression
The method and apparatus enhance video compression by utilizing 3D facial representations to improve compression efficiency and support speech-to-face communication at ultra-low bit rates.
Patent Information
- Application Number
- JP2025516005
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-10
- Filing Date
- 2023-10-17
- Publication Date
- 2025-11-12
Smart Images

Figure 2025536871000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This disclosure claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 379,779, filed October 17, 2022, and claims the benefit of U.S. Patent Application No. 18 / 484,099, entitled "METHOD AND APPARATUS FOR TALKING FACE VIDEO COMPRESSION," filed October 10, 2023. Both of these applications are incorporated herein by reference in their entirety.
[0002] Technical Field FIELD OF THE DISCLOSURE
[0002] The present disclosure relates generally to video processing, and more particularly to a method and apparatus for talking face video compression. [Background technology]
[0003] background
[0003] A video is a set of static pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video may be compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, most commonly based on prediction, transform, quantization, entropy coding, and in-loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC) (HEVC / H.265) standard, the Versatile Video Coding (VVC) (VVC / H.266) standard, and the AVS standard, which define specific video coding formats, are developed by standardization organizations. As increasingly advanced video coding techniques are adopted in video standards, the coding efficiency of new video coding standards becomes increasingly higher. Summary of the Invention
[0004] Summary of the Invention An embodiment of the present disclosure provides a method for decoding video data, the method including: decompressing compressed frames to generate key frames representing a face; generating a first set of parameters for the key frames, the first set of parameters being associated with a three-dimensional (3D) facial representation of the face; reconstructing a second set of parameters associated with the 3D facial representation of the face for each of one or more inter frames according to a compressed inter prediction residual of the second set of parameters; and generating a video including the face based on the key frames, the first set of parameters, and the second set of parameters.
[0005]
[0005] An embodiment of the present disclosure provides a method for encoding video data, the method including: compressing keyframes representing a face; extracting a set of parameters related to a three-dimensional (3D) face expression from one or more inter-frames; and compressing an inter-prediction residual of the set of parameters.
[0006] An embodiment of the present disclosure provides an apparatus for decoding video data, including: a decoder configured to decompress compressed frames to generate key frames representing a face; an extractor configured to generate a first set of parameters related to a three-dimensional (3D) facial representation of the face for the key frames and to reconstruct a second set of parameters related to the 3D facial representation of the face for each of one or more inter frames according to a compressed inter prediction residual of the second set of parameters; and a generator configured to generate a video including the face based on the key frames, the first set of parameters, and the second set of parameters.
[0007] An embodiment of the present disclosure provides an apparatus for encoding video data, the apparatus including: an encoder configured to compress keyframes representing a face; an extractor configured to extract a set of parameters related to a three-dimensional (3D) face expression from one or more inter frames; and an encoding module configured to compress an inter prediction residual of the set of parameters.
[0008] An embodiment of the present disclosure provides an apparatus for decoding video data, the apparatus including: a memory configured to store instructions; and one or more processors configured to execute the instructions to cause the apparatus to: decompress compressed frames to generate keyframes representing a face; generate a first set of parameters for the keyframes, the first set of parameters being associated with a three-dimensional (3D) facial representation of the face; reconstruct, for each of one or more interframes, a second set of parameters being associated with the 3D facial representation of the face according to a compressed inter prediction residual of the second set of parameters; and generate a video including the face based on the keyframes, the first set of parameters, and the second set of parameters.
[0009]
[0009] An embodiment of the present disclosure provides an apparatus for encoding video data, the apparatus including: a memory configured to store instructions; and one or more processors configured to execute the instructions to cause the apparatus to: compress keyframes representing a face; extract a set of parameters related to a three-dimensional (3D) facial expression from one or more inter-frames; and compress an inter-prediction residual of the set of parameters.
[0010]
[0010] An embodiment of the present disclosure provides an apparatus for decoding video data, the apparatus including: a memory configured to store instructions; and one or more processors configured to execute the instructions to cause the apparatus to: receive a bitstream including three-dimensional (3D) parameters; reconstruct a 3D mesh according to the 3D parameters; learn optical flow for synthesis; and reconstruct a video.
[0011]
[0011] An embodiment of the present disclosure provides a non-transitory computer-readable storage medium for storing a bitstream of video. The bitstream includes a set of parameters related to a three-dimensional (3D) facial representation. When the set of parameters is decoded by a decoder, the set of parameters causes the decoder to perform a method including: decompressing compressed frames to generate keyframes representing a face; generating a first set of parameters for the keyframes, the first set of parameters related to a three-dimensional (3D) facial representation of the face; reconstructing a second set of parameters related to the 3D facial representation of the face for each of one or more interframes according to a compressed inter prediction residual of the second set of parameters; and generating a video including the face based on the keyframes, the first set of parameters, and the second set of parameters.
[0012]
[0012] An embodiment of the present disclosure provides a non-transitory computer-readable storage medium for storing a video bitstream. The bitstream includes a set of parameters related to a three-dimensional (3D) facial expression. When the set of parameters is decoded by a decoder, the set of parameters causes the decoder to perform a method including: compressing keyframes representing a face; extracting the set of parameters related to the three-dimensional (3D) facial expression from one or more inter frames; and compressing an inter prediction residual of the set of parameters.
[0013]
[0013] An embodiment of the present disclosure provides a computer program product including computer program instructions: the computer program instructions enable a computer to perform a method for decoding video data in accordance with the above method embodiment for decoding video data.
[0014]
[0014] An embodiment of the present disclosure provides a computer program product including computer program instructions: the computer program instructions enable a computer to perform a method for encoding video data in accordance with the above method embodiment for encoding video data.
[0015]
[0015] An embodiment of the present disclosure provides a computer program that enables a computer to perform a method for decoding video data according to the above method embodiment for decoding video data.
[0016]
[0016] An embodiment of the present disclosure provides a computer program that enables a computer to perform a method for encoding video data according to the above method embodiment for encoding video data.
[0017] BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Embodiments and various aspects of the present disclosure are illustrated in the following detailed description and the accompanying drawings, in which various features are not drawn to scale. [Brief explanation of the drawings]
[0018] [Figure 1]
[0018] FIG. 1 is a schematic diagram illustrating an exemplary system for encoding image data according to some embodiments of the present disclosure. [Figure 2A]
[0019] FIG. 1 is a schematic diagram illustrating an exemplary encoding process of a hybrid video encoding system according to an embodiment of the present disclosure. [Figure 2B]
[0020] FIG. 2 is a schematic diagram illustrating another exemplary encoding process of a hybrid video encoding system according to an embodiment of the present disclosure. [Figure 3A]
[0021] FIG. 2 is a schematic diagram illustrating an exemplary decoding process of a hybrid video coding system according to an embodiment of the present disclosure. [Figure 3B]
[0022] FIG. 2 is a schematic diagram illustrating another exemplary decoding process of a hybrid video coding system according to an embodiment of the present disclosure. [Figure 4]
[0023] 1 is a block diagram of an exemplary apparatus for encoding image data according to some embodiments of the present disclosure. [Figure 5]
[0024] 1 is a schematic diagram illustrating the architecture of a traditional video compression framework. [Figure 6]
[0025] FIG. 1 is a schematic diagram illustrating an example architecture of an end-to-end deep-based video compression framework according to some embodiments of the present disclosure. [Figure 7]
[0026] FIG. 1 is a schematic diagram illustrating another exemplary architecture of an end-to-end deep-based video generation and compression framework. [Figure 8]
[0027] FIG. 1 is a schematic diagram illustrating an example encoder-decoder coding framework with a 1x4x4 compact feature size for talking face video according to some embodiments of the present disclosure. [Figure 9]
[0028] FIG. 1 is a schematic diagram illustrating a general encoder-decoder generated compression framework for 3DMM-assisted talking face video. [Figure 10]
[0029] FIG. 1 is a schematic diagram illustrating an exemplary encoder-decoder generated compression framework for talking face video according to some embodiments of the present disclosure. [Figure 11]
[0030] FIG. 1 is a schematic diagram illustrating an exemplary method for encoding video data according to some embodiments of the present disclosure. [Figure 12]
[0031] FIG. 1 is a schematic diagram illustrating an exemplary method for encoding video data according to some embodiments of the present disclosure. [Figure 13]
[0032] FIG. 2 is a schematic diagram illustrating an exemplary method for decoding video data according to some embodiments of the present disclosure. [Figure 14]
[0033] FIG. 2 is a schematic diagram illustrating an exemplary method for decoding video data according to some embodiments of the present disclosure. [Figure 15]
[0034] FIG. 2 is a schematic diagram illustrating an exemplary method for decoding video data according to some embodiments of the present disclosure. [Figure 16]
[0035] FIG. 2 is a schematic diagram illustrating an exemplary method for decoding video data according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0019] Detailed Description
[0036] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which like numbers in the various drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following description of exemplary embodiments do not represent all implementations consistent with the present disclosure. Instead, these implementations are merely examples of apparatus and methods consistent with aspects related to the present disclosure, as recited in the appended claims. Particular aspects of the present disclosure are described in more detail below. In the event of a conflict with a term or definition incorporated by reference, the term and definition provided herein shall control.
[0020]
[0037] Various image / video coding standards have been developed for image / video compression, such as JPEG, JPEG2000, H.264 / MPEG4 part 10, Audio Video coding Standard (AVS), and H.265 / HEVC. Currently, a new Universal Video Coding (VVC) standard is under development to further improve video coding efficiency. VVC is based on the same hybrid video coding system that has been used in recent video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc. All of the above standards use a hybrid coding framework including intra / inter prediction, transform, quantization, and entropy coding, which are used to exploit spatial / temporal, visual, and statistical redundancies in images / video.
[0021]
[0038] Traditional compression algorithms compress video frames by using block-based motion estimation, discrete cosine transform (DCT), etc. In addition, traditional or learning-based end-to-end video compression methods assume a universal natural scene without special consideration of human motion information. For these existing 2D generative compression algorithms, there is no support for including semantic information for controlling head movement posture in the compressed code stream. However, there is an increasing need for speech-to-face communication at ultra-low bit rates.
[0022]
[0039] 1 is a block diagram illustrating a system 100 for encoding image data according to some disclosed embodiments. Image data may include an image (also called a "picture" or "frame"), multiple images, or a video. An image is a static picture. Multiple images may or may not be related to each other (either spatially or temporally). A video is a set of images arranged in time sequence.
[0023]
[0040] 1, system 100 includes a source device 120 that provides encoded video data that is subsequently decoded by a destination device 140. Consistent with disclosed embodiments, source device 120 and destination device 140 may each include any of a wide range of devices, including a desktop computer, a notebook (e.g., laptop) computer, a server, a tablet computer, a set-top box, a mobile phone, a vehicle, a camera, an image sensor, a robot, a television, a camera, a wearable device (e.g., a smartwatch or wearable camera), a display device, a digital media player, a video game console, a video streaming device, etc. Source device 120 and destination device 140 may be capable of wireless or wired communication.
[0024]
[0041] 1, source device 120 may include image / video encoder 124 and output interface 126. Destination device 140 may include input interface 142 and image / video decoder 144. Image / video encoder 124 encodes an input bitstream and outputs encoded bitstream 162 via output interface 126. Encoded bitstream 162 is transmitted over communication medium 160 and received by input interface 142. Image / video decoder 144 then decodes encoded bitstream 162 to generate decoded data.
[0025]
[0042] Specifically, source device 120 may further include various devices (not shown) for providing source image data to be processed by image / video encoder 124. Devices for providing source image data may include image / video capture devices (cameras, image / video archives or storage devices containing previously captured images / video, image / video supply interfaces that receive images / video from image / video content providers, etc.).
[0026]
[0043] Each of the image / video encoder 124 and the image / video decoder 144 may be implemented as any of a wide variety of suitable encoder or decoder circuitry, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. If the encoding or decoding is implemented partially in software, the image / video encoder 124 or the image / video decoder 144 may store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware by using one or more processors to perform techniques according to this disclosure. Each of the image / video encoder 124 or the image / video decoder 144 may be included within one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC) within the respective device.
[0027]
[0044] Image / video encoder 124 and image / video decoder 144 may operate according to any video coding standard, such as Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), AOMedia Video1 (AV1), Joint Photographic Experts Group (JPEG), Moving Picture Experts Group (MPEG), etc. Alternatively, image / video encoder 124 and image / video decoder 144 may be custom devices that do not conform to an existing standard. Although not shown in FIG. 1 , in some embodiments, image / video encoder 124 and image / video decoder 144 may each be integrated with an audio encoder and decoder and may include appropriate MUX-DEMUX units or other hardware and software to handle the encoding of both audio and video in a common data stream or separate data streams.
[0028]
[0045] Output interface 126 may include any type of medium or device capable of transmitting encoded bitstream 162 from source device 120 to destination device 140. For example, output interface 126 may include a transmitter or transceiver configured to transmit encoded bitstream 162 in real time directly from source device 120 to destination device 140. Encoded bitstream 162 may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device 140.
[0029]
[0046] Communication medium 160 may include a transitory medium, such as a wireless broadcast or a wired network transmission. For example, communication medium 160 may include a radio frequency (RF) spectrum or one or more physical transmission paths (e.g., cables). Communication medium 160 may form part of a packet-based network (such as a local area network, a wide area network, or a global network such as the Internet, etc.). In some embodiments, communication medium 160 may include routers, switches, base stations, or any other equipment that may be useful for facilitating communication from source device 120 to destination device 140. For example, a network server (not shown) may receive encoded bitstream 162 from source device 120 and provide encoded bitstream 162 to destination device 140 (e.g., via network transmission).
[0030]
[0047] Communication medium 160 may also be in the form of a storage medium (e.g., a non-transitory storage medium), such as a hard disk, a flash drive, a compact disc, a digital video disc, a Blu-ray disc, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded image data. In some embodiments, a computing device at a media production facility, such as a disc stamping facility, may receive the encoded image data from source device 120 and produce a disc containing the encoded video data.
[0031]
[0048] Input interface 142 may include any type of medium or device capable of receiving information from communication medium 160. The received information includes encoded bitstream 162. For example, input interface 142 may include a receiver or transceiver configured to receive encoded bitstream 162 in real time.
[0032]
[0049] Exemplary image data encoding and decoding techniques (such as those utilized by image / video encoder 124 and image / video decoder 144) will now be described with reference to FIGS. 2A-2B and 3A-3B.
[0033]
[0050] FIG. 2A shows a schematic diagram of an exemplary encoding process 200A according to an embodiment of the present disclosure. For example, encoding process 200A may be performed by an encoder, such as image / video encoder 124 of FIG. 1. As shown in FIG. 2A, the encoder may encode a video sequence 202 into a video bitstream 228 according to process 200A. The video sequence 202 may include a set of pictures (referred to as "original pictures") arranged in temporal order. Each original picture of the video sequence 202 may be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder may perform process 200A at the level of basic processing units for each original picture of the video sequence 202. For example, the encoder may perform process 200A in an iterative manner, in which case the encoder may encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder may perform process 200A in parallel for a region of each original picture of video sequence 202.
[0034]
[0051] In FIG. 2A , an encoder may provide a basic processing unit (referred to as an “original BPU”) of an original picture of a video sequence 202 to a prediction stage 204 to generate prediction data 206 and a prediction BPU 208. The encoder may subtract the prediction BPU 208 from the original BPU to generate a residual BPU 210. The encoder may provide the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may provide the prediction data 206 and the quantized transform coefficients 216 to a binary coding stage 226 to generate a video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as a “forward path.” During process 200A, after quantization stage 214, the encoder may provide quantized transform coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 to generate reconstructed residual BPU 222. The encoder may add reconstructed residual BPU 222 to prediction BPU 208 to generate prediction reference 224 used in prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as the "reconstruction path." The reconstruction path may be used to ensure that both the encoder and decoder use the same reference data for prediction.
[0035]
[0052] The encoder may perform a process 200A that iterates between encoding each original BPU of an original picture (in the forward path) and generating (in the reconstruction path) a prediction reference 224 for encoding the next original BPU of the original picture. After encoding all original BPUs of an original picture, the encoder may proceed to encode the next picture in the video sequence 202.
[0036]
[0053] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to any act of receiving, inputting, obtaining, retrieving, acquiring, reading, accessing, or in any manner for inputting data.
[0037]
[0054] In the prediction stage 204, in the current iteration, the encoder may receive the original BPU and a prediction reference 224, perform a prediction operation, and generate predicted data 206 and a predicted BPU 208. The prediction reference 224 may be generated from a reconstruction path of a previous iteration of the process 200A. The purpose of the prediction stage 204 is to reduce information redundancy by extracting predicted data 206, which may be used to reconstruct the original BPU as a predicted BPU 208 from the prediction data 206 and the prediction reference 224.
[0038]
[0055] Ideally, predicted BPU 208 would be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, predicted BPU 208 typically differs slightly from the original BPU. To record such differences, after generating predicted BPU 208, the encoder may subtract predicted BPU 208 from the original BPU to generate residual BPU 210. For example, the encoder may subtract pixel values (e.g., grayscale or RGB values) of predicted BPU 208 from corresponding pixel values of the original BPU. Each pixel of residual BPU 210 may have a residual value resulting from such a subtraction between corresponding pixels of the original BPU and predicted BPU 208. Compared to the original BPU, predicted data 206 and residual BPU 210 may have fewer bits, which can be used to reconstruct the original BPU without significant quality degradation. Thus, the original BPU is compressed.
[0039]
[0056] To further compress the residual BPU 210, in the transform stage 212, the encoder may reduce spatial redundancy in the residual BPU 210 by decomposing the residual BPU 210 into a set of two-dimensional "basic patterns," each associated with a "transform coefficient." The base patterns may have the same size (e.g., the size of the residual BPU 210). Each base pattern may represent a variation frequency (e.g., frequency of luminance variation) component of the residual BPU 210. None of the base patterns can be reproduced from any combination (e.g., a linear combination) of any other base patterns. In other words, this decomposition may decompose the variation of the residual BPU 210 into the frequency domain. Such a decomposition is similar to the discrete Fourier transform of a function, with the base patterns similar to the basis functions (e.g., trigonometric functions) of the discrete Fourier transform and the transform coefficients similar to the coefficients associated with the basis functions.
[0040]
[0057] Different transform algorithms may use different base patterns. Different transform algorithms (e.g., discrete cosine transform, discrete sine transform, or the like) may be used in transform stage 212. The transform in transform stage 212 is invertible. That is, the encoder may reconstruct residual BPU 210 by inversely operating the transform (referred to as the "inverse transform"). For example, to reconstruct pixels of residual BPU 210, the inverse transform may generate a weighted sum by multiplying the values of corresponding pixels of the base pattern by their associated coefficients and adding these products. For video coding standards, both the encoder and decoder may use the same transform algorithm (and therefore the same base pattern). Thus, the encoder may record only the transform coefficients, and the decoder may reconstruct residual BPU 210 from the transform coefficients without receiving the base pattern from the encoder. Compared to residual BPU 210, the transform coefficients may have fewer bits, but can be used to reconstruct residual BPU 210 without significant quality degradation. In this way, the residual BPU 210 is further compressed.
[0041]
[0058] The encoder may further compress the transform coefficients in the quantization stage 214. In the transform process, different base patterns may represent different fluctuation frequencies (e.g., luminance fluctuation frequencies). Because the human eye is generally better at perceiving low-frequency fluctuations, the encoder may ignore high-frequency fluctuation information without causing significant quality degradation during decoding. For example, in the quantization stage 214, the encoder may generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (referred to as a "quantization parameter") and rounding the quotient to the nearest integer. After such an operation, some transform coefficients of high-frequency base patterns may be transformed to zero, and transform coefficients of low-frequency base patterns may be transformed to smaller integers. The encoder may ignore zero-valued quantized transform coefficients 216, thereby further compressing the transform coefficients. The quantization process is also invertible, and the quantized transform coefficients 216 may be reconstructed into transform coefficients by the inverse operation of quantization (referred to as "dequantization").
[0042]
[0059] Because the encoder ignores the remainder of such division in rounding operations, quantization stage 214 may be lossy. Typically, quantization stage 214 may contribute the greatest information loss in process 200A. The greater the information loss, the fewer bits the quantized transform coefficients 216 require. To obtain different levels of information loss, the encoder may use different values of the quantization parameter or any other parameter of the quantization process.
[0043]
[0060] In binary encoding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 by using a binary encoding technique (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossless compression algorithm, etc.). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in binary encoding stage 226 (e.g., a prediction mode used in prediction stage 204, parameters of the prediction operation, a transform type in transform stage 212, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), or the like). The encoder may use output data of binary encoding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 may be further packetized for network transmission.
[0044]
[0061] Referring to the reconstruction path of process 200A, in inverse quantization stage 218, the encoder may perform inverse quantization on quantized transform coefficients 216 to generate reconstructed transform coefficients. In inverse transform stage 220, the encoder may generate reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add reconstructed residual BPU 222 to prediction BPU 208 to generate prediction reference 224, which will be used in the next iteration of process 200A.
[0045]
[0062] It should be noted that other variations of process 200A may be used to encode video sequence 202. In some embodiments, the stages of process 200A may be performed in a different order by the encoder. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be split into multiple stages. For example, transform stage 212 and quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may omit one or more stages of FIG. 2A.
[0046]
[0063] 2B shows a schematic diagram of another exemplary encoding process 200B according to an embodiment of the present disclosure. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder that complies with a hybrid video coding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B additionally includes a mode decision stage 230 and divides prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.
[0047]
[0064] Generally, prediction techniques can be categorized into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-prediction") may use pixels from one or more already-encoded neighboring BPUs within the same picture to predict a current BPU. That is, the prediction reference 224 in spatial prediction may include neighboring BPUs. Spatial prediction may reduce inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter-prediction") may use regions from one or more already-encoded pictures to predict a current BPU. That is, the prediction reference 224 in temporal prediction may include an encoded picture. Temporal prediction may reduce inherent temporal redundancy of a picture.
[0048]
[0065] Referring to process 200B, in the forward path, the encoder performs prediction operations in a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra prediction. For an original BPU of a picture being coded, the prediction reference 224 may include one or more neighboring BPUs coded (in the forward path) and reconstructed (in the reconstruction path) within the same picture. The encoder may generate the predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, or the like. In some embodiments, the encoder may perform extrapolation at the pixel level, such as by extrapolating the value of a corresponding pixel for each pixel of the predicted BPU 208. The neighboring BPUs used for extrapolation may be located relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., below-left, below-right, above-left, or above-right of the original BPU), or any direction defined in the video coding standard used. For intra prediction, the prediction data 206 may include, for example, the locations (e.g., coordinates) of the neighboring BPUs used, the sizes of the neighboring BPUs used, parameters of the extrapolation, the orientation of the neighboring BPUs used relative to the original BPU, or the like.
[0049]
[0066] For another example, in the temporal prediction stage 2044, the encoder may perform inter-prediction. For an original BPU of a current picture, the prediction reference 224 may include one or more pictures (referred to as "reference pictures") that have been coded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures may be coded and reconstructed for each BPU. For example, the encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a reconstructed BPU. Once all reconstructed BPUs of the same picture have been generated, the encoder may generate the reconstructed picture as the reference picture. The encoder may perform a "motion estimation" operation to search for a matching region within a range (referred to as a "search window") of the reference picture. The location of the search window in the reference picture may be determined based on the location of the original BPU of the current picture. For example, the search window may be centered at a location in the current picture that has the same coordinates in the reference picture as the coordinates of the original BPU, and may extend outward over a predetermined distance. When the encoder identifies a region similar to the original BPU within the search window (e.g., by using a pel-recursive algorithm, a block matching algorithm, or the like), the encoder may determine such a region as a matching region. The matching region may have different dimensions than the original BPU (e.g., smaller than, equal to, larger than, or a different shape than the original BPU). Because the reference picture and the current picture are temporally separated in the timeline, the matching region may be thought of as "moving" to the location of the original BPU over time. The encoder may record the direction and distance of such movement as a "motion vector." If multiple reference pictures are used, the encoder may search for the matching region for each reference picture and determine its associated motion vector. In some embodiments, the encoder may assign weights to the pixel values of the matching region in each matching reference picture.
[0050]
[0067] Motion estimation may be used to identify various types of motion, such as, for example, translation, rotation, zooming, or the like. For inter prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, a motion vector associated with the matching region, the number of reference pictures, weights associated with the reference pictures, or the like.
[0051]
[0068] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., a motion vector) and the prediction reference 224. For example, the encoder may shift the matching region of the reference picture according to the motion vector, in which case the encoder may predict the original BPU of the current picture. If multiple reference pictures are used, the encoder may shift the matching region of the reference picture according to each motion vector and average the pixel values of the matching region. In some embodiments, if the encoder assigns weights to the pixel values of the matching region of each matching reference picture, the encoder may add a weighted sum of the pixel values of the shifted matching region.
[0052]
[0069] In some embodiments, inter prediction may be unidirectional or bidirectional. Unidirectional inter prediction may use one or more reference pictures in the same temporal direction relative to the current picture. Unidirectional inter prediction uses a reference picture preceding the current picture. Bidirectional inter prediction may use one or more reference pictures in both temporal directions relative to the current picture.
[0053]
[0070] Still referring to the forward path of process 200B, after spatial prediction stage 2042 and temporal prediction stage 2044, in mode decision stage 230, the encoder may select a prediction mode (e.g., one of intra-prediction or inter-prediction) for the current iteration of process 200B. For example, the encoder may perform a rate-distortion optimization technique, where the encoder may select a prediction mode to minimize the value of a cost function depending on the bitrates of the candidate prediction modes and the distortion of the reconstructed reference picture under the candidate prediction modes. Depending on the selected prediction mode, the encoder may generate a corresponding predicted BPU 208 and predicted data 206.
[0054]
[0071] In the reconstruction path of process 200B, if an intra-prediction mode is selected in the forward path, after generating prediction reference 224 (e.g., the current BPU coded and reconstructed in the current picture), the encoder may directly provide prediction reference 224 to spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If an inter-prediction mode is selected in the forward path, after generating prediction reference 224 (e.g., the current picture coded and reconstructed in all BPUs), the encoder may provide prediction reference 224 to loop filter stage 232, where the encoder may apply a loop filter to prediction reference 224 to reduce or remove distortions (e.g., blocking artifacts) introduced by inter prediction. The encoder may apply various loop filter techniques in loop filter stage 232, such as, for example, deblocking, sample adaptive offset, adaptive loop filter, or the like. The loop-filtered reference pictures may be stored in a buffer 234 (or a "decoded picture buffer") for later use (e.g., to be used as inter-predicted reference pictures for future pictures in the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) along with the quantized transform coefficients 216, the prediction data 206, and other information in the binary encoding stage 226.
[0055]
[0072] FIG. 3A shows a schematic diagram of an exemplary decoding process 300A according to an embodiment of the present disclosure. Process 300A may be a decompression process corresponding to compression process 200A of FIG. 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder (e.g., image / video decoder 144 of FIG. 1) may decode video bitstream 228 into video stream 304 according to process 300A. Video stream 304 may be very similar to video sequence 202. However, due to information loss in the compression and decompression processes (e.g., quantization stage 214 in FIGS. 2A-2B), video stream 304 is generally not identical to video sequence 202. Similar to processes 200A and 200B in FIGS. 2A-2B, the decoder may perform process 300A at the level of a basic processing unit (BPU) for each picture encoded in video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, in which case the decoder may decode a basic processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for several regions of each picture encoded in video bitstream 228.
[0056]
[0073] In FIG. 3A , a decoder may provide a portion of a video bitstream 228 associated with a basic processing unit (referred to as a “coded BPU”) of a coded picture to a binary decoding stage 302. In the binary decoding stage 302, the decoder may decode this portion into prediction data 206 and quantized transform coefficients 216. The decoder may provide the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may provide the prediction data 206 to a prediction stage 204 to generate a predicted BPU 208. The decoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder may provide the prediction reference 224 to the prediction stage 204 for performing a prediction operation in the next iteration of the process 300A.
[0057]
[0074] The decoder may perform process 300A to iteratively decode each coded BPU of a coded picture and generate a prediction reference 224 for coding the next coded BPU of the coded picture. After decoding all coded BPUs of a coded picture, the decoder may output the picture to the video stream 304 for display and proceed to decode the next coded picture in the video bitstream 228.
[0058]
[0075] In binary decoding stage 302, the decoder may perform the inverse operation of the binary coding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to prediction data 206 and quantized transform coefficients 216, the decoder may decode other information in binary decoding stage 302, such as, for example, a prediction mode, parameters of the prediction operation, a transform type, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), or the like. In some embodiments, if video bitstream 228 is transmitted over a network in packets, the decoder may depacketize video bitstream 228 before providing it to binary decoding stage 302.
[0059]
[0076] 3B shows a schematic diagram of another exemplary decoding process 300B according to an embodiment of the present disclosure. Process 300B may be modified from process 300A. For example, process 300B may be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 300A, process 300B additionally divides prediction stage 204 into spatial prediction stage 2042 and temporal prediction stage 2044, and additionally includes loop filter stage 232 and buffer 234.
[0060]
[0077] In process 300B, prediction data 206 decoded by the decoder from binary decoding stage 302 for a coded elementary processing unit (referred to as a "current BPU") of a coded picture being decoded (referred to as a "current picture") may include various types of data, depending on what prediction mode was used by the encoder to code the current BPU. For example, if intra-prediction is used by the encoder to code the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra-prediction, parameters of the intra-prediction operation, or the like. Parameters of the intra-prediction operation may include, for example, locations (e.g., coordinates) of one or more neighboring BPUs used as references, sizes of the neighboring BPUs, parameters of extrapolation, orientations of the neighboring BPUs relative to the original BPU, or the like. For another example, if inter-prediction is used by the encoder to code the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter-prediction, parameters of the inter-prediction operation, or the like. Parameters for the inter prediction operation may include, for example, the number of reference pictures currently associated with the BPU, weights associated with each of the reference pictures, locations (e.g., coordinates) of one or more matching regions within each reference picture, one or more motion vectors associated with each of the matching regions, or the like.
[0061]
[0078] Based on the prediction mode indicator, the decoder may determine whether to perform spatial prediction (e.g., intra prediction) in spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in temporal prediction stage 2044. Details of performing such spatial or temporal prediction are described in FIG. 2B and will not be repeated hereafter. After performing such spatial or temporal prediction, the decoder may generate a predicted BPU 208. The decoder may add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, as described in FIG. 3A.
[0062]
[0079] In process 300B, the decoder may provide prediction reference 224 to spatial prediction stage 2042 or temporal prediction stage 2044 to perform a prediction operation in the next iteration of process 300B. For example, if the current BPU is decoded by using intra prediction in spatial prediction stage 2042, after generating prediction reference 224 (e.g., the decoded current BPU), the decoder may provide prediction reference 224 directly to spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If the current BPU is decoded by using inter prediction in temporal prediction stage 2044, after generating prediction reference 224 (e.g., the reference picture from which all BPUs are decoded), the encoder may provide prediction reference 224 to loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may apply a loop filter to prediction reference 224 in the manner described in FIG. 2B . The loop-filtered reference picture may be stored in a buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., to be used as an inter-prediction reference picture for a future encoded picture of the video bitstream 228). The decoder may store one or more reference pictures used in the temporal prediction stage 2044 in the buffer 234. In some embodiments, if the prediction mode indicator in the prediction data 206 indicates that inter-prediction was used to encode the current BPU, the prediction data may further include loop filter parameters (e.g., loop filter strength).
[0063]
[0080] Referring back to FIG. 1 , each of the image / video encoder 124 and the image / video decoder 144 may be implemented as any suitable hardware, software, or combination thereof. FIG. 4 is a block diagram of an example apparatus 400 for processing image data according to an embodiment of the present disclosure. For example, the apparatus 400 may be an encoder or a decoder. As shown in FIG. 4 , the apparatus 400 may include a processor 402. When the processor 402 executes the instructions described herein, the apparatus 400 may be a dedicated machine for encoding or decoding image data. The processor 402 may be any type of circuitry capable of manipulating or processing information. For example, processor 402 may include any number or combination of a central processing unit (or "CPU"), a graphics processing unit (or "GPU"), a neural processing unit ("NPU"), a microcontroller unit ("MCU"), an optical processor, a programmable logic controller, a microcontroller, a microprocessor, a digital signal processor, an intellectual property (IP) core, a programmable logic array (PLA), a programmable array logic (PAL), a generic array logic (GAL), a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a system on a chip (SoC), an application-specific integrated circuit (ASIC), or the like. In some embodiments, processor 402 may also be a set of processors grouped as a single logic element.For example, as shown in FIG. 4, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0064]
[0081] The device 400 may also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, or the like). For example, as shown in FIG. 4, the stored data may include program instructions (e.g., program instructions for implementing a stage in process 200A, 200B, 300A, or 300B), data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 may access the program instructions and data for processing (e.g., via bus 410) and execute the program instructions to perform operations or manipulations on the data for processing. The memory 404 may include a high-speed random-access storage device or a non-volatile storage device. In some embodiments, memory 404 may include any number or combination of random-access memory (RAM), read-only memory (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, CompactFlash (CF) cards, or the like. Memory 404 may also be a collection of memories (not shown in FIG. 4) grouped as a single logical element.
[0065]
[0082] Bus 410 may be a communication device that transfers data between elements internal to apparatus 400, such as an internal bus (e.g., a CPU-memory bus), an external bus (e.g., a universal serial bus port, a peripheral component interconnect express port), or the like.
[0066]
[0083] For ease of explanation without causing ambiguity, the processor 402 and other data processing circuitry will be collectively referred to in this disclosure as "data processing circuitry." The data processing circuitry may be implemented entirely as hardware or as a combination of software, hardware, or firmware. In addition, the data processing circuitry may be a single, stand-alone module or may be fully or partially integrated into any other element of the device 400.
[0067]
[0084] Device 400 may further include a network interface 406 for providing wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, or the like). In some embodiments, network interface 406 may include any number or combination of a network interface controller (NIC), a radio frequency (RF) module, a transponder, a transceiver, a modem, a router, a gateway, a wired network adapter, a wireless network adapter, a Bluetooth® adapter, an infrared adapter, a near-field communication ("NFC") adapter, a cellular network chip, or the like.
[0068]
[0085] In some embodiments, apparatus 400 may optionally further include a peripheral interface 408 for providing connection to one or more peripheral devices. As shown in Figure 4, the peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, touchpad, or touchscreen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera, or an input interface coupled to a video archive), etc.
[0069]
[0086] It should be noted that a video codec (e.g., a codec performing process 200A, 200B, 300A, or 300B) may be implemented as any combination of software or hardware modules within apparatus 400. For example, some or all stages of process 200A, 200B, 300A, or 300B may be implemented as one or more software modules of apparatus 400, such as program instructions that may be loaded into memory 404. As another example, some or all stages of process 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of apparatus 400, such as dedicated data processing circuitry (e.g., FPGA, ASIC, NPU, or the like).
[0070]
[0087] In accordance with the disclosed embodiments, deep learning can be used in image and video compression to achieve competitive performance compared to traditional compression techniques. For example, end-to-end image compression algorithms, such as JPEG, JPEG2000, and HEVC, exhibit better rate-distortion (RD) performance due to end-to-end training and nonlinear transformations. Furthermore, video compression algorithms based on deep neural networks (DNNs), such as the deep video compression model (DVC), can achieve promising RD performance. These methods can work without a priori knowledge of the video content. For video conferencing / telephony applications, deep generative models such as First Order Motion Model (FOMM) and Face_vid2vid (Face Video-to-Video Synthesis) can achieve promising performance at very low bitrates. In particular, these models exploit the fact that variations in these videos typically reside in human motion information, which provides strong priors that can be used in frame synthesis. These features are described by variations in human structure, such as landmarks or keypoints, and are further propagated to animate reference frames and generate moving human footage.
[0071]
[0088] Despite their realization, the above-mentioned deep learning techniques rely on keypoints or landmarks with explicit representations in terms of 2D, rather than 3D, facial representation, thereby limiting their performance in rendering high-quality video and subsequent applications. Specifically, with the development of the Metaverse, along with the widespread use of digital human characters in numerous applications, real-world video communication systems are being applied to the Metaverse market, thus requiring corresponding control functions (e.g., pose control, expression control, etc.). However, these 2D generative compression algorithms cannot support the needs of the Metaverse due to the limitations of facial representation. For better rate-distortion performance and a more promising application market for 3D human communication, 3D facial a priori information provided by 3D Morphable Models (3DMMs) (e.g., HeadGAN, Face2Faceρ, MeshGAN, FACEGAN) has been explored over the past few years. These 3DMM-based generative algorithms can not only achieve ultra-low bitrate talking-face video compression, but also control pose and expression for interesting applications.
[0072]
[0089] Traditional video compression standards such as Advanced Video Coding (AVC), HEVC, and VVC have been developed to achieve superior compression performance. In all of these standards, a block-based hybrid video coding framework is used to exploit spatial redundancy, temporal redundancy, and information entropy redundancy in video.
[0073]
[0090] Generally, a video compression encoder generates a bitstream based on an input current frame, and a decoder reconstructs a video frame based on the received bitstream. Figure 5 is a schematic diagram showing the architecture of a traditional video compression framework. Figure 5 shows that the classical framework of video compression follows a predictive transformation architecture.
[0074]
[0091] Specifically, for input frame xt is divided into a set of blocks (eg, square regions) of equal size (eg, 8x8). The encoding procedure of a traditional video compression algorithm on the encoder 500 side includes the following steps:
[0075]
[0092] Motion estimation by the block-based motion estimation module 501 of the encoder 500: The motion estimation module 501 estimates the current frame x t and the previous reconstructed frame
number
[0076]
[0093] Motion compensation by the motion compensation module 502 of the encoder 500: Predicted frame
number
number
number
[0077]
[0094] The transform module 503 and the Q module 504 of the encoder 500 transform and quantize the residual r t is provided by the Q module 504
number
[0078]
[0095] Inverse transformation by the inverse transformation module 505 of the encoder 500: quantization result
number
number
[0079]
[0096] Entropy coding by the entropy coding module 506 of the encoder 500: the motion vector v t and the quantization result
number
[0080]
[0097] Frame reconstruction by the reconstruction module 507: Reconstructed frame
number
number
number
number
[0081]
[0098] For a decoder (not shown), based on the bits provided by the entropy coding module 506 of the encoder 500, motion compensation, inverse quantization, and then frame reconstruction are performed to generate a reconstructed frame.
number
[0082]
[0099] As mentioned above, deep learning-based algorithms can be introduced to replace or enhance traditional video coding tools, including intra / inter prediction, entropy coding, and in-loop filtering. End-to-end image / video compression algorithms can be used for joint optimization of the entire image / video compression framework, rather than designing one specific module. For example, the DVC approach, an end-to-end video coding scheme that jointly optimizes all elements for video compression, can be used. Furthermore, to address content adaptive and error propagation aware problems, an online coder update scheme can be used to improve video compression performance. Additionally, FVC can be used by exploiting all key modules of the end-to-end compression framework in the feature space. Based on recurrent probability models and weighted recurrent quality enhancement networks, "Recurrent Learning for Video Compression (RLVC)" and "HLVC" can be used to exploit temporal correlation between video frames. Four effective modules in "Multiple Frames Prediction for Learned Video Compression (M-LVC)" can be used. However, like traditional video coding tools, these learning-based video compression methods assume generic natural scenes without specifically considering human content such as faces, bodies, or other parts.
[0083]
[0100] FIG. 6 is a schematic diagram illustrating an example architecture of an end-to-end deep-based video compression framework according to some embodiments of the present disclosure. FIG. 6 illustrates the basic framework of a first end-to-end video compression deep model, which jointly optimizes all elements for video compression, such as motion estimation, motion compression, and residual compression. Specifically, learning-based optical flow estimation is utilized to obtain motion information and reconstruct the current frame. Then, two autoencoder-style neural networks are employed to compress the corresponding motion and residual information. All modules are jointly trained via a single loss function (here, they cooperate with each other by considering the trade-off between reducing the number of compression bits and improving the quality of the decoded video). There is a one-to-one correspondence between the traditional video compression framework shown in FIG. 5 and the novel end-to-end deep-based framework shown in FIG. 6. An overview of their relationships and differences is provided below. The procedure of the encoder 600 may include the following steps.
[0084]
[0101] Motion estimation and compression: In the optical flow network module 601, a CNN (Convolutional Neural Network) model estimates motion information v t Instead of directly encoding the raw optical flow values, the MV encoder-decoder network compresses and decodes the optical flow values. First, the MV encoder net module 602 encodes the motion information v t can be used to encode the motion information v t The coded motion representation of
number
number
[0085]
[0102] Motion Compensation: A motion compensation network, denoted as the motion compensation net module 605, calculates a predicted frame based on the obtained optical flow.
number
number
number
[0086]
[0103] Transform, quantization, and inverse transform: Linear transforms are replaced by using a highly nonlinear residual encoder-decoder network, such as the residual encoder net module 606 shown in FIG. 6, and the residual r t is the expression y t Then, y t is achieved by Q Module 607
number
number
number
[0087]
[0104] Entropy coding: In the experimental stage, quantized motion representation
number
number
number
[0088]
[0105] Furthermore, the loss of the encoder 600 can be determined according to the original frame, the reconstructed frame, and the encoded frame, and the determined loss can also be used to refine the network in the encoder 600 to achieve better performance.
[0089]
[0106] Frame reconstruction (not shown): This is the same as the traditional method.
[0090]
[0107] With the emergence of deep generative models, including variational auto-encoding (VAE) and generative adversarial networks (GAN), facial video compression can achieve promising performance improvements. For example, X2Face can be used to control face generation via image, audio, and pose coding. Furthermore, a realistic neural talking head model can be used via few-shot adversarial learning. Face-vidtovid can be used for video-to-video synthesis tasks. Furthermore, methods that leverage compact 3D keypoint representations to drive generative models for rendering target frames can also be used. Furthermore, a mobile-compatible video chat system based on FOMM can be used. VSBNet, which utilizes few-shot adversarial learning to reconstruct source frames from landmarks, can also be used. Additionally, an end-to-end talking-head video compression framework based on compact feature learning (CFTE), designed for highly efficient talking-head video compression for ultra-low bandwidth scenarios, can be used. The CFTE scheme exploits compact feature representations to compensate for temporal evolution and reconstruct target face video frames in an end-to-end manner. Furthermore, the CFTE scheme can be incorporated into a video coding framework under rate-distortion supervision. Although these algorithms achieve frame reconstruction with some facial parameters through the powerful rendering capabilities of deep generative models, some head pose and facial expression motions are still not accurately rendered compared to the original video.
[0091]
[0108] 7 is a schematic diagram illustrating another exemplary architecture of a deep-based video generative compression framework according to some embodiments of the present disclosure. Figure 7 provides the basic framework of a deep-based video generative compression scheme based on a first-order motion model (FOMM). FOMM deforms a reference source frame to follow the motion of driving video. While this method can be applied to various types of video (e.g., Tai Chi, cartoons), this method typically focuses on facial animation applications. FOMM follows an encoder-decoder architecture with a motion transfer element that includes the following steps:
[0092]
[0109] First, a keypoint extractor (also called a motion module) is trained by using equivalent loss without explicit labels. Two sets of 10 training keypoints are calculated for the source frame and the driving frame by this keypoint extractor. The training keypoints are transformed from a feature map with a size of channel × 64 × 64 through a Gaussian map function, so every corresponding keypoint can represent different channel feature information. It should be mentioned that every keypoint is a point (x, y) that can represent the most important information of the feature map.
[0093]
[0110] Second, a dense motion network uses the landmarks and source frames to generate a dense motion field and occlusion map.
[0094]
[0111] Then, the encoder 710 encodes the source frames via traditional image / video compression methods such as HEVC / VVC or JPEG / BPG, where VVC is used to compress the source frames.
[0095]
[0112] In a later stage, the resulting feature map is warped using a dense motion field (using differentiable grid-sample operations) and then multiplied with the occlusion map.
[0096]
[0113] Finally, the decoder 720 generates the image from the warped map.
[0097]
[0114] 8 is a schematic diagram illustrating an exemplary encoder-decoder coding framework with a 1x4x4 compact feature size for talking face video according to some embodiments of the present disclosure. Figure 8 illustrates another basic framework of a deep-based video generative compression technique based on compact feature representation (i.e., CFTE). The basic framework follows an encoder-decoder architecture that applies a context-based coding scheme.
[0098]
[0115] On the encoder 810 side, the compression framework includes three modules: an encoder (also called a VVC encoding module) for compressing key frames, a feature extractor for extracting compact human features of other inter frames, and a feature encoding module for compressing inter prediction residuals of the compact human features. First, key frames representing human texture are compressed by a VVC encoder. Through the compact feature extractor, each subsequent inter frame is represented by a compact feature matrix with a size of 1x4x4. It should be mentioned that the size of the compact feature matrix is not fixed, and the number of feature parameters can also be increased or decreased according to the specific requirements of bit consumption. Then, these extracted features are inter predicted and quantized, and the residuals are finally entropy coded as the final bitstream.
[0099]
[0116] On the decoder 820 side, this compression framework also includes three main modules, including: decoding to reconstruct key frames, reconstructing compact features through entropy decoding and compensation, and generating final videos by utilizing the reconstructed features and decoded key frames. More specifically, during the generation of final videos, decoded key frames from a VVC bitstream can be further represented in the form of "features via compact feature extraction." Then, given the features from key and inter frames, an appropriate coarse motion field is calculated to facilitate the generation of pixel-wise dense motion maps and occlusion maps. Finally, based on a deep generative model, the decoded key frames, pixel-wise dense motion maps, and occlusion maps with implicit motion field characterization are used to generate final videos with accurate appearance, pose, and expression.
[0100]
[0117] To further improve coding performance, numerous studies have focused on 3D faces. A 3D head model is adopted, and only the pose parameters are coded for the task of face-specific video compression. Subsequently, both eigenspace models and principal component analysis (PCA) models have been used in this task. However, based on these traditional 3D techniques, the visual quality of the reconstructed images is unacceptable. With the development of deep generative models, this 3D MMM-assisted face video generation task may provide promising results.
[0101]
[0118] 9 is a schematic diagram illustrating a general encoder-decoder generation compression framework for 3DMM-assisted speaking face video according to some embodiments of the present disclosure. Generally speaking, 3DMM-assisted face video generation involves:
number
number
number
[0102]
[0119] While traditional or learning-based end-to-end video compression methods can achieve relatively efficient compression performance in moving human videos, the direct application of common compression algorithms to ultra-low bitrate talking face video compression systems has several drawbacks.
[0103]
[0120] First, traditional compression algorithms compress every video frame by using block-based motion estimation, discrete cosine transform (DCT), etc., so it is still difficult to further reduce the coding bits. As a result, this type of algorithm is not suitable for ultra-low bitrate human video compression scenes.
[0104]
[0121] The second drawback is that these traditional or learning-based end-to-end video compression methods assume a universal natural scene without considering human motion information. In particular, features from talking faces or moving objects are described by the variation of feature structures (such as landmarks or keypoints) with strong priors, which can greatly help reconstruct higher quality videos.
[0105]
[0122] FOMM or Face vidtovid Although these generative compression algorithms such as [1] have fully realized frame reconstruction with several parameters through the powerful rendering ability of deep generative models, some head pose motions and facial expression motions still cannot be accurately rendered compared with the original talking face or moving body video. That is, most of the head pose motions and facial expression motions based on 2D face representation (i.e., 2D landmarks and 2D keypoints) either perform poorly in terms of photorealism, or do not satisfy the identity preservation problem, or do not fully convey the driving pose and expression.
[0106]
[0123] Furthermore, these existing 2D generative compression algorithms that include semantic information do not support controlling head movement pose in the compressed code stream. This drawback limits the application of human communication within the metaverse. On the other hand, for most 3D MMM-assisted generative models, the transmitted facial parameters are complex and therefore require more compression bits, making them inadequate for ultra-low bandwidth communication scenarios.
[0107]
[0124] To overcome these problems with a viable solution, this disclosure provides a framework for ultra-low bitrate speech-to-face communication. The disclosed framework uses techniques for reconstructing the 3D face of a digital human character and is based on the assumption of consistency and persistence of human appearance. That is, the disclosed 3DMM-assisted compression framework is based on the view that the temporal evolution of facial images can be well described by a 3D facial representation. In particular, such a latent, compact, and meaningful representation, which can be automatically learned, can efficiently remove redundancy and capture temporal variations. This is also consistent with the recent hypothesis in neuroscience research that the human visual system transforms the evolution of visual signals into a space that follows a more linear temporal trajectory.
[0108]
[0125] In particular, relevant 3D facial parameters are extracted, facilitating 3D face reconstruction with extremely low transmitted bitrates, and a 3D facial template is stored at the receiver side. Subsequently, after receiving the bitstream and decoding the 3DMM parameters, 3D meshes (i.e., source mesh and driving mesh) are reconstructed and guided to learn the optical flow required for face synthesis. Technically, the identity information of the driving face is explicitly excluded in the reconstructed driving mesh. In this way, the disclosed network can focus on source face motion estimation without the interference of the driving face shape. Finally, with the help of accurate motion estimation information and source appearance, a talking face video can be reconstructed. Therefore, the disclosed scheme enjoys the advantages of high flexibility, enhanced robustness, and semantic control due to digital character-level representation and 3D facial template-based rendering.
[0109]
[0126] In some embodiments, a 3DMM-assisted talking face video compression scheme is proposed.
[0110]
[0127] 10 is a schematic diagram illustrating an example encoder-decoder generated compression framework 1000 for talking face video according to some embodiments of the present disclosure. The encoder-decoder generated compression framework 1000 may include an encoder 1010 and a decoder 1020.
[0111]
[0128] As shown in FIG. 10, a 3DMM-based attitude-controlled generative compression framework 1000 is presented that can be used to realize ultra-low bitrate human video communication.
[0112]
[0129] The functionality of the encoder 1010 may be implemented in the source device 120 of Figure 1, the encoder 200A of Figure 2A, or the encoder 200B of Figure 2B. As shown in Figure 10, the encoder 1010 may include a keyframe encoding module 1011 (such as a VVC encoder using a VVC encoding technique) for compressing keyframes, a 3D facial parameter extraction module 1012 that applies a 3DMM pre-trained model (e.g., WM3DR) to extract appropriate 3D parameters of other inter-frames, and a compact facial parameter compression module 1013 (also referred to as a feature encoding module) for compressing inter-prediction residuals of 3D facial parameters.
[0113]
[0130] First, keyframes representing human textures are compressed by the keyframe encoding module 1011. Then, the compressed keyframes are embedded into a bitstream to be sent to a receiver. Through the 3D facial parameter extraction module 1012, which applies a 3DMM pre-trained model, each subsequent interframe is represented by 3DMM parameters (e.g., expression parameters and pose parameters, including translation parameters, angle parameters, etc.). It should be mentioned that the size of the compact 3D facial parameters is not fixed, and the number of expression parameters can also be increased or decreased according to the specific requirements of bit consumption. Next, these extracted features are inter-predicted and quantized by the compact facial parameter compression module 1013, and the residue is finally entropy-encoded by the compact facial parameter compression module 1013 and embedded into the final bitstream.
[0114]
[0131] The functionality of the decoder 1020 may be implemented in the destination device 140 of Figure 1, the decoder 300A of Figure 3A, or the decoder 300B of Figure 3B. As shown in Figure 10, the decoder 1020 may include a keyframe decoding (reconstruction) module 1021 for reconstructing keyframes, a compact facial parameter reconstruction module 1023 for reconstructing compact 3D facial parameters by entropy decoding and compensation, a mesh reconstruction module 1024 for applying a 3DMM template to reconstruct a 3D mesh based on the decoded 3D facial parameters, and a generation module 1027 for generating a final video by utilizing the reconstructed 3D mesh and the decoded keyframes. Additionally, the 3D facial parameter extraction module 1022 may extract 3D facial parameters (e.g., identifier (id), expression, translation, angle, etc.) from the reconstructed keyframes, while the 3D facial parameter extraction module 1025 may generate inter-frame 3D facial parameters by combining the reconstructed 3D facial parameters from the compact facial parameter reconstruction module 1023 with the identifier (id) information from the 3D facial parameter extraction module 1022.
[0115]
[0132] More specifically, during the generation of final videos, decoded keyframes from a VVC bitstream can be further represented in the form of a 3D face mesh through 3D facial parameter extraction and mesh reconstruction based on a 3DMM pre-trained model (e.g., the WM3DR model). Next, given a mesh reconstruction module 1024 that applies a 3D face template, a corresponding 3D face mesh can be reconstructed based on 3D facial features from keyframes and interframes. Thereafter, by using the reconstructed 3D meshes (i.e., the keyframe mesh and the interframe mesh) as guidance, a coarse-to-fine motion estimation module 1026 in the decoder 1020 can generate a pixel-wise dense motion map and an occlusion map. Finally, based on a generation module 1027, such as a deep generative model including a discriminator, the decoded keyframes, the pixel-wise dense motion map, and the occlusion map with implicit motion field characterization are used to generate output videos with accurate appearance, pose, and expression.
[0116]
[0133] In some embodiments, the encoder 1010 may include a 3D facial parameter extraction module 1012 based on a 3DMM.
[0117]
[0134] Generally speaking, a 3DMM has a geometry given by
number
number
number
number
number
[0118]
[0135] In some embodiments, the 3D face templates can be stored offline at the encoder 1010 and decoder 1020 side instead of being transmitted in each communication connection. The reason behind this is that facial appearance can be assumed to be an invariant within a certain period of time. Also, in some embodiments, since it is not the first time that the receiver (i.e., decoder 1020) draws a specific 3D face with the same identifier, it is feasible to store the face template models offline at the receiver (i.e., decoder 1020) side. It should be noted that these 3D face templates can be generated by a 3D face reconstruction algorithm or manual interaction. The input of the transmitter side (i.e., encoder 1010) is the talking face video {I} captured by a camera. t |t=0,1,2...N} and a 3D face template
number
[0119]
[0136] In this task, the 3D facial parameters P(representation β∈R) generated by the 3D facial parameter extraction module 1012 are 64 , translation l∈R 3 and angle θ∈R 3 Regarding the representation parameters, different dimensions can be selected for transmission according to the bandwidth occupancy. Indeed, if the representation dimension is R 64 From R 32 or lower representation dimensions (e.g., R 16 ,R 8 ,……), the reconstruction quality of the talking face video will be worse, but it is beneficial to reduce bit consumption.
[0120]
[0137] In some embodiments, the encoder 1010 may encode one or more parameters of inter-frames. To efficiently encode parameters of 3D compact features of inter-frames, a predicted feature compression framework of the compact facial parameter compression module 1013 is provided. In the predicted feature compression framework, previously coded features are used to predict the current feature, and only the predicted residual is coded. First, these 3D compact facial parameters P∈{β, l, θ} generated by the 3D facial parameter extraction module 1012 are quantized and inter-predicted from the corresponding decoded features from the previous frame to remove redundancies in the feature representation. This process is given by the following equation:
number
number
[0121]
[0138] After quantization and inter-prediction, the residual is entropy coded by a zero-order exponential Golomb function coded by the compact facial parameter compression module 1013. Then, these binary codes are compressed into a final bitstream by context-based arithmetic coding. It should be noted that the context model of our lossless coding method is based on "Prediction by Partial Matching" (PPM), where a set of previous symbols in the uncompressed symbol stream is used to predict the next symbol. It is understood that such a context model is an adaptive statistical data compression technique with a skewed probability distribution, which can further realize high-efficiency coding of the compact facial parameter residual. Finally, the coded bitstream compressed by the context-based entropy algorithm is transmitted to the decoder 1020 side via a communication network.
[0122]
[0139] In some embodiments, the decoder 1020 may provide motion estimation based on a 3D face mesh.
[0123]
[0140] At the receiver (i.e., decoder 1020) side, the bitstream is entropy decoded by a compact facial parameter reconstruction module 1023, so that the reconstructed 3D facial parameters are obtained by compensating with the reconstructed 3D facial parameters of the previous frame. After obtaining these 3D compact facial parameters P∈β,l,θ, 3D meshes (i.e., key-frame meshes and inter-frame meshes) are reconstructed by a mesh reconstruction module 1024 based on the 3DMM templates. The related process may refer to the process of compact facial parameter extraction based on 3DMM at the encoder 1010 side as mentioned above.
[0124]
[0141] Based on the reconstructed face mesh, a pixel-wise fine motion map and occlusion map learning scheme can be generated by the coarse-fine motion estimation module 1026. First, the coarse motion field M coarseis the reconstructed 3D face mesh of the keyframe and interframe in the decoder 1020 (i.e., Mesh key and Mesh inter ) is obtained by transforming the 3D mesh. Specifically, the reconstructed 3D face mesh can be projected to 2D mesh points, mapped to a 2D image plane, and then a difference operation can be performed to represent the motion trajectory. However, limited by the capabilities of the 3D MMM, this transformed coarse flow is not accurate enough. Specifically, this flow cannot represent the motion of other parts other than the face, such as the upper body and hair.
[0125]
[0142] To solve this problem, M coarse After obtaining M coarse is used to estimate the coarse-deformed frames of the fine motion field representation (i.e., F) along with the downsampled keyframes. cdf ) to help generate the coarsely deformed frame F cdf , original keyframe F key and gross motor field M coarse By concatenating the U-Net predictor, we also obtain a pixel-wise dense motion map (i.e., M dense ) and the occlusion map (i.e., M occlusion ) Such an operation can fully exploit the implicit motion field characterization from the compact feature representation and can be beneficial for estimating the final video. M dense =P1(f U-Net (concat(F cdf ,M coarse ,F key )))、 M occlusion =P2(f U-Net (concat(F cdf ,M coarse ,F key ))) Here, P1(·) and P2(·) indicate two different prediction outputs. Therefore, the coarse-to-fine motion estimation module 1026 calculates the fine motion map (M dense ) and occlusion map (M occlusion ) to the generation module 1027.
[0126]
[0143] As mentioned above, unlike existing algorithms in which a rendered face from a 3DMM template is directly input to a deep generative network, the face generation compression algorithm provided by the present disclosure via the dense motion estimation module 1026 provides a mechanism to leverage the motion between two 3D rendered face meshes (e.g., a key-frame mesh and an inter-frame mesh) for better 2D face generation.
[0127]
[0144] In some embodiments, the decoder 1020 may provide a generation module 1027. Deep neural networks, especially deep generative networks, have strong inference capabilities for reconstructing realistic images. The generation module 1027 here may apply a deep generative network. To achieve promising generation results, a feature warping strategy is used to dense is used to warp the reconstructed keyframes from the keyframe decoding (reconstruction) module 1021 according to the dense motion field M dense Compared with the real results
number
number
number
number
[0128]
[0145] In some embodiments, model supervision and loss functions may also be considered in the framework 1000.
[0129]
[0146] In the proposed framework 1000, perceptual loss, adversarial loss, identity-preserving loss, and reconstruction texture are adopted to manage the end-to-end training process. It should be noted that these related loss functions do not need to be used together and can be combined according to actual task requirements.
[0130]
[0147] Perceptual loss: To reconstruct more realistic images, a perceptual loss is used as the reconstruction loss to combine the pre-trained VGG-19 network. The perceptual loss is calculated by combining the coarsely deformed frame F with the downsampled original interframe I. cdf and the conversion result
number
number
number
[0131]
[0148] Adversarial Loss: To further improve the realism of our generated images, we use multiple discriminators (i.e., D i ) is operated for different image resolutions. The corresponding losses of the generator G and discriminator D are given by:
number
[0132]
[0149] Identifier-Preserving Loss: For identifier identification and preservation, a pre-trained face recognition model (i.e., ArcFace) is employed to capture the most salient facial features and make the reconstruction results have a small distance in the deep feature space. A specific implementation is similar to perceptual loss.
number
[0133]
[0150] Reconstruction Texture Loss: Detailed texture is an important consideration in face generation. To effectively capture facial texture information, a Gram matrix is introduced to calculate feature correlations in different layers from a pre-trained VGG-19 network.
number
[0134]
[0151] In summary, the overall end-to-end training loss is given by: L total =λ initial L per-initial +λ final L per-final +λ adv (L G +L D )+λ id L id +λ tex L tex where λ initial and λ final are both set to 10. adv , λ id and λ tex are equal to 1, 40, and 100, respectively. Note that these values of λ are set through empirical experiments. Other reasonable values may also be considered to obtain a better training model.
[0135]
[0152] In some embodiments, posture-controlled talking face video compression is further provided.
[0136]
[0153] The proposed 3DMM assist framework 1000 can control head movement pose with compressed code streams, which can be further applied to ultra-low bandwidth face video conferencing, human communication in the metaverse, virtual uploaders for live commerce, etc.
number
[0137]
[0154] The advantages of the posture-controlled talking face video compression method for different scenarios are presented below.
[0138]
[0155] Ultra-Low Bandwidth Facial Video Conferencing: Over the past three years, the world has experienced an unprecedented and prolonged COVID-19 epidemic, and the demand for video conferencing / chat has increased dramatically. The 3DMM-assisted generative compression network provided by this disclosure fully exploits the strong statistical regularity of facial images to realize end-user facial image reconstruction for ultra-low bitrates, and compresses only 3D compact parameters, thereby improving the efficiency of transmitting talking facial videos.
[0139]
[0156] Human Communication in the Metaverse: The Metaverse is a virtual world and digital living space (mapped or surpassed by the real world) constructed by humans using digital technology and can interact with the real world. Human-character communication is very important to this new social system. The 3DMM-assisted generative compression network provided by this disclosure can effectively realize human parameter transfer and human character reconstruction. For example, the extracted pose information and facial parameters in a 3DMM template can be used to draw corresponding facial meshes, and appropriate motion information can be learned from these meshes and directly transferred to a specific human character. Therefore, the communication bit rate remains stable, just like the communication bit rate in face video conferencing.
[0140]
[0157] Virtual Characters for Live Commerce: With the growing demand for domestic live broadcasting, virtual commerce characters seem to be a development shortcut and a new trend in live broadcast delivery. Virtual characters for live commerce can provide better freshness and attract more customers to live rooms, while retaining the interactivity of real people. With the rise of the younger generation, the two-dimensional world is also more attractive. The 3DMM-assisted generation compression network provided by the present disclosure can be applied to virtual characters for live commerce. The network bandwidth load when there are many customers can be reduced if these customers are in the live room at the same time.
[0141]
[0158] FIG. 11 is a schematic diagram illustrating an example method 1100 for encoding video data according to some embodiments of the present disclosure. For example, method 1100 may be performed by one or more processors associated with an encoder, such as image / video encoder 124 (FIG. 1), encoder 1010 (FIG. 10), etc. In some embodiments, image / video encoder 124 may be integrated into apparatus 400 shown in FIG. 4, such that method 1100 may be performed by apparatus 400. In other examples, method 1100 may be performed by an encoder simulated by a general-purpose processing unit and necessary auxiliary elements. In this situation, the encoder is implemented as a user application or program. As shown in FIG. 11, method 1100 includes the following steps 1110-1130. Additionally, optional steps 1140 and 1150 within dashed boxes are described in detail below according to some other embodiments.
[0142]
[0159] In step 1110, the encoder compresses the keyframes representing the faces.
[0143]
[0160] Specifically, the encoder may compress key frames according to Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or other video standards. Key frames may not be inter-frame coded. For example, key frames may be coded as I-frames (intra-frames), which are coded independently of other frames. On the decoding side, key frames may be decoded without reference to other frames. In a non-limiting example of VVC coding, consecutive key frames consisting of several key frames may be coded "intra-only" (or "all-intra," simplified as "AI"). In this disclosure, key frames contain texture information and can therefore be used to reconstruct a subject's face. Therefore, although not specifically detailed in some embodiments, key frames may be quantized and entropy coded to achieve acceptable quality under actual environmental regulations.
[0144]
[0161] In some embodiments, the key frame and one or more inter frames are generated sequentially by an image capture device or an image capture device array. In other words, the key frame and the inter frames of the key inter frames are organized along a time axis. The consecutive frames are captured by a capture device or a capture device array directed toward a subject (e.g., a human face). When generated by a capture device array, frames from various devices in the array can be merged to generate a merged frame as a key frame or an inter frame. The merging process here can be a concatenation process of two inter frames from various devices to obtain a fine picture of the target or other pixel fusion techniques.
[0145]
[0162] In step 1120, the encoder extracts a set of parameters related to a three-dimensional (3D) facial expression from one or more inter-frames that reference the keyframe for encoding. In some non-limiting examples of the present disclosure, inter-frames other than keyframes may be represented by a set of parameters rather than actual pictures to achieve very low bitrate human video communication. In some examples, the set of parameters may include facial expression, translation, or angle. In other examples, the set of parameters may also include texture, shape, and other parameters that can be used to represent a human face as a supplement. It should be noted that in the present disclosure, any frame other than a keyframe may exist as an inter-frame. Therefore, visual information within most frames can be represented by compact parameters since most frames are inter-frames.
[0146]
[0163] In a non-limiting example, the encoder may be used to construct a 3D morphable model (3DMM), and a set of parameters may be extracted by the 3DMM. In some examples, the 3DMM may be a weakly-supervised multi-face 3D reconstruction (WM3DR) model.
[0147]
[0164] As mentioned above, the 3DMM is given by
number
number
number
number
number
[0148]
[0165] As already explained above, the 3D face template can be stored offline at the encoder and decoder side instead of being transmitted in each communication connection. It should be noted that the 3D MMM extractor may extract one feature information or a combination of different feature information. In some embodiments, a 3D MMM face model other than the WM3DR model can be used.
[0149]
[0166] Therefore, the expression β∈R 64 , translation l∈R 3 and angle θ∈R 33D face parameters P including R can be transmitted to save bits. Regarding the representation parameters, different dimensions can be selected for transmission according to bandwidth occupancy. Indeed, when the representation dimension is R 64 From R 32 or lower representation dimensions (e.g., R 16 , R 8 , ...), the reconstruction quality of the talking face video will be worse, but it is beneficial to reduce bit consumption.
[0150]
[0167] In step 1130, the encoder compresses the inter prediction residual of the set of parameters. Like the inter frame itself, the generated parameters of the corresponding inter frame may also possess the residual. Thus, the generated set of parameters can be represented by fewer bits to obtain better coding efficiency.
[0151]
[0168] As shown in FIG. 12, the step 1130 of compressing the inter prediction residual of a set of parameters may be realized by the following steps 1210 and 1220.
[0152]
[0169] In step 1210, the encoder determines the difference between a set of parameters of two adjacent frames.
[0153]
[0170] As mentioned above, previously coded features can be used to predict the current feature, and only the residual after prediction is coded. First, the 3D compact face parameters (i.e., a set of parameters related to the 3D face representation) P∈{β, l, θ} can be quantized and inter-predicted from the corresponding decoded features from the previous frame to remove redundancy in the feature representation. This process is given by:
number
number
[0154]
[0171] In step 1220, the encoder encodes this difference as a compressed inter-prediction residual in a set of parameters.
[0155]
[0172] After quantization and inter-prediction, the residuals can be entropy coded by zero-order exponential Golomb coding. Then, these binary codes are compressed into a final bitstream by context-based arithmetic coding. It should be noted that our lossless coding method is based on prediction by context model "prediction by partial matching (PPM)", where a set of previous symbols in the uncompressed symbol stream is used to predict the next symbol. It is understood that such a context model is an adaptive statistical data compression technique with a skewed probability distribution, which can further realize high-efficiency coding of compact facial parameter residuals. Finally, the coded bitstream compressed by the context-based entropy algorithm is transmitted to the decoder side via a communication network.
[0156]
[0173] It will be appreciated that one key frame with a subsequent inter frame may be used to encode a small (e.g., 4 second) episode of video. For longer videos, more than one key frame may be used. That is, method 1100 may be utilized to encode video with several key frames and corresponding inter frames.
[0157]
[0174] 11 , method 1100 may further include step 1140 for generating a bitstream including compressed keyframes of a set of parameters and compressed inter-prediction residuals. The encoder may generate the bitstream for transmission or storage. When processed by a decoder, the generated bitstream will be decoded as keyframes and a set of parameters that can be used to reconstruct a subject's face.
[0158]
[0175] 11, method 1100 may further include transmitting 1150 the bitstream over a public or private network. Output interface 126 (FIG. 1) may be used to transmit the bitstream over communication medium 160 (FIG. 1) to destination device 140 (FIG. 1).
[0159]
[0176] It can be concluded here that "method 1100 can compress any keyframe representing human texture, for example, by a VVC encoder." Through the pre-trained WM3DR model, each subsequent interframe is represented by 3DMM parameters (i.e., expression parameters and pose parameters). The size of the compact 3D face parameters is not fixed, and the number of expression parameters can also be increased or decreased according to the specific requirements of bit consumption. In the following process, these extracted features can be inter-predicted and quantized, and the residual is finally entropy coded as the final bitstream.
[0160]
[0177] FIG. 13 is a schematic diagram illustrating an exemplary method for decoding video data according to some embodiments of the present disclosure. For example, method 1300 may be performed by one or more processors, such as image / video decoder 144 (FIG. 1), decoder 1020 (FIG. 10), etc. In some embodiments, image / video encoder 144 may be incorporated into device 400 shown in FIG. 4, such that method 1300 may be performed by device 400. In other examples, method 1300 may be performed by a decoder simulated by a general-purpose processing unit and necessary auxiliary elements. In this situation, the decoder is implemented as a user application or program. As shown in FIG. 13, method 1300 includes the following steps 1310-1340. Additionally, optional steps 1350 and 1360 within dashed boxes are described in detail below according to some other embodiments.
[0161]
[0178] In step 1310, the decoder decompresses / decodes the compressed frames to generate key frames representing faces. As mentioned above, key frames may be intra-coded and therefore may not depend on other frames. The decompressed / decoded key frames contain rich texture information of a human face (essential for reconstructing a human face). However, the key frames may also contain facial expression, translation, or angle information.
[0162]
[0179] It is understood that keyframes themselves can be combined to generate coarse video with a relatively low frame rate. However, keyframes can also be used to construct video when the communication environment is not ideal. Under these circumstances, the parameters generated from interframes may be delayed or missing, so keyframes may be the only candidate for constructing video as a supplementary scheme.
[0163]
[0180] In step 1320, the decoder generates a first set of parameters related to a three-dimensional (3D) facial expression of the face for the keyframe. It should be noted here that the set of parameters, the first set of parameters, and the second set of parameters may have the same components (such as expression, translation, or angle). The antecedents "first" and "second" merely facilitate identifying the subject for which the parameters are generated. As shown in FIG. 10, the reconstructed keyframe may be input to a pre-trained WM3DR model to extract driving descriptors including expression, translation, angle, and other parameters. In some embodiments, various 3D MMM face models other than the WM3DR model may be selected. The extraction procedure on the decoding side may be performed similarly to the encoding side, and the same descriptors may be extracted here.
[0164]
[0181] In step 1330, the decoder reconstructs, for each interframe of the one or more interframes, a second set of parameters related to a 3D facial representation of the face according to the compressed inter-prediction residual of the second set of parameters. Specifically, the compressed inter-prediction residual of the second set of parameters may be entropy decoded, such that the reconstructed 3D facial parameters (the second set of parameters) are obtained by compensating with the reconstructed 3D facial parameters of the previous frame.
[0165]
[0182] In step 1340, the decoder generates a video including a face based on the keyframes, the first set of parameters generated in step 1320, and the second set of parameters generated in step 1330. In other words, a face in the video may be generated according to the keyframes, the first set of parameters, and the second set of parameters. In another example, the talking face may also be used to combine with several other elements (e.g., background, clothing, props) to generate an avatar of the subject. For example, the video may be used in ultra-low bandwidth face video conferencing, human communication in the metaverse, virtual characters for live commerce, and other possible scenarios.
[0166]
[0183] In some embodiments, as shown in FIG. 14, step 1340 of generating a video including a face based on key frames, a first set of parameters, and a second set of parameters may be realized by the following steps 1410 to 1440.
[0167]
[0184] In step 1410, the decoder generates a keyframe mesh according to a first set of parameters, for example, by a 3D morphable model (3DMM). In step 1420, the decoder generates an interframe mesh according to a second set of parameters by the 3DMM for each interframe of one or more interframes. After obtaining the 3D compact facial parameters P∈β,l,θ, the 3D meshes (i.e., the keyframe mesh and the interframe mesh) are reconstructed based on the 3DMM template. The related process may refer to the process of compact facial parameter extraction based on the 3DMM on the encoder side as described above (which is not repeated here for brevity in this disclosure).
[0168]
[0185] In step 1430, the decoder determines a dense motion map and an occlusion map according to the key-frame mesh and the inter-frame mesh of each frame (i.e., key-frame and inter-frame). The dense motion map and the occlusion map will provide the optical flow for synthesizing a human face, and this disclosure does not limit how to generate the dense motion map and the occlusion map.
[0169]
[0186] In some embodiments, as shown in Figure 15, step 1430 may be realized by the following steps: in step 1510, the decoder obtains a coarse motion field of the face according to the key-frame mesh and the inter-frame mesh. In step 1520, the decoder generates a coarse deformation frame by inputting the key-frame and the coarse motion field into a deep neural network. The deep neural network may be built within the decoder. In step 1530, the decoder estimates a dense motion map and an occlusion map by concatenating the coarse motion frame, the key-frame, and the coarse motion field into the deep neural network.
[0170]
[0187] As mentioned above, the gross motor field M coarseis the reconstructed 3D face mesh (i.e., Mesh key and Mesh inter , also called key-frame mesh and inter-frame mesh). Specifically, the reconstructed 3D face mesh is projected to 2D mesh points, mapped into the 2D image plane, and then a difference operation is performed to represent the motion trajectory. However, the transformed coarse flow may not be accurate enough. Specifically, this flow cannot represent the motion of other parts other than the face, such as the upper body and hair.
[0171]
[0188] To solve this problem, M coarse After obtaining M coarse The U-Net architecture (coarse deformation frames of the fine motion field representation (i.e., F)) is used together with the downsampled keyframes. cdf ) which helps to generate the coarsely deformed frame F cdf , original keyframe F key , and the gross motor field M coarse By concatenating the U-Net predictor, we also obtain a pixel-wise dense motion map (i.e., M dense ) and the occlusion map (i.e., M occlusion ) Such an operation can fully exploit the implicit motion field characterization from the compact feature representation and can be beneficial for estimating the final video. M dense =P1(f U-Net (concat(F cdf ,M coarse ,F key )))、 M occlusion =P2(f U-Net (concat(F cdf ,M coarse ,F key ))) Here, P1(·) and P2(·) indicate two different prediction outputs.
[0172]
[0189] As mentioned above, unlike existing algorithms in which a drawn face from a 3D MMM template is directly input to a deep generative network, the face generation compression algorithm provided by the present disclosure provides a mechanism to leverage the motion between two 3D drawn face meshes (i.e., a key-frame mesh and an inter-frame mesh) for better 2D face generation.
[0173]
[0190] 14 again, in step 1440, the decoder generates a video according to the keyframes, the dense motion map, and the occlusion map. The dense motion map and the occlusion map can provide the optical flow of the picture, so that a video including a human face can be constructed by considering the optical flow.
[0174]
[0191] 16, step 1340 may be further implemented by the following steps: in step 1610, the decoder warps the keyframes according to the dense motion map, and in step 1620, the decoder generates the image by calculating the Hadamard product of the occlusion map and the warped result.
[0175]
[0192] As mentioned above, deep neural networks (especially deep generative networks) have strong estimation capabilities to reconstruct realistic images. A decoder may apply deep generative networks. To achieve promising generative results, a feature warping strategy is used to dense is used to warp the reconstructed keyframes according to the dense motion field M dense Compared with the real results
number
number
number
number
[0176]
[0193] Referring back to FIG. 13 , method 1300 may further include step 1350 of identifying a facial identifier in the keyframe. The identifier is a necessary parameter that distinguishes various people. In some examples, the WM3DR model and some other parts of the encoding scheme depend on the identifier of the object. Thus, the decoder may identify the facial identifier and determine which model parameters to use. In addition, the identifier parameters may also be used to determine the identifier preservation loss, which is described in more detail below. Although shown as the final step in FIG. 13 , those skilled in the art will understand that the identification may occur after generating the keyframe, and this disclosure is not limited thereto.
[0177]
[0194] Still referring to FIG. 13, the method 1300 may further include receiving 1360 a bitstream including a compressed frame and an inter prediction residual for each inter frame of the one or more inter frames.
[0178]
[0195] In some examples, the generated video may include a corresponding reconstructed frame for each interframe of one or more interframes. According to some embodiments described above, the method 1100 receives a bitstream including three-dimensional (3D) parameters, reconstructs a 3D mesh according to the 3D parameters, learns an optical flow for synthesis according to the 3D mesh, and reconstructs a video (e.g., a video with a speaking human face) according to the optical flow. The video is reconstructed using motion estimation information and source appearance. The 3D mesh includes a source mesh (keyframe mesh) and a driving mesh (interframe mesh).
[0179]
[0196] In some embodiments, model supervision and loss functions are introduced. For example, a perceptual loss, an adversarial loss, an identity-preserving loss, and a reconstruction-texture loss may be employed to supervise the end-to-end training process. It should be noted that these related loss functions do not need to be used together and may therefore be combined according to actual task requirements. The mathematical processes for determining the perceptual loss, the adversarial loss, the identity-preserving loss, and the reconstruction-texture loss have been suggested above, and therefore will not be repeated here for brevity in this disclosure.
[0180]
[0197] The deep neural network and generation module 1027 may be updated according to at least any of the following: a perceptual loss, an adversarial loss, an identity-preserving loss, and a reconstruction texture loss between each of one or more inter frames and its corresponding reconstructed frame.
[0181]
[0198] In some embodiments, a non-transitory computer-readable storage medium is also provided that stores the bitstream, which may be encoded and decoded according to the above-described encoder-decoder generated compression framework for talking face video (e.g., FIG. 10).
[0182]
[0199] In some embodiments, a non-transitory computer-readable storage medium containing instructions is also provided, and the instructions may be executed by a device (such as the disclosed encoders and decoders) to perform the above-described methods. Common types of non-transitory media include, for example, a floppy disk, a flexible disk, a hard disk, a solid-state drive, a magnetic tape, or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROM, and EPROM, FLASH-EPROM or any other flash memory, NVRAM, cache, registers, any other memory chip or cartridge, and networked versions thereof. A device may include one or more processors (CPUs), input / output interfaces, network interfaces, or memory.
[0183]
[0200] An embodiment of the present disclosure provides a computer program product including computer program instructions: the computer program instructions enable a computer to perform a method for decoding video data according to the above method embodiment for decoding video data.
[0184]
[0201] An embodiment of the present disclosure provides a computer program product including computer program instructions: the computer program instructions enable a computer to perform a method for encoding video data according to the above method embodiment for encoding video data.
[0185]
[0202] An embodiment of the present disclosure provides a computer program enabling a computer to perform the method for decoding video data according to the above method embodiments for decoding video data.
[0186]
[0203] An embodiment of the present disclosure provides a computer program enabling a computer to perform a method for encoding video data according to the above method embodiments for encoding video data.
[0187]
[0204] Some embodiments may be further described using the following clauses: 1. A method of decoding video data, comprising: decompressing the compressed frames to generate keyframes representing a face; generating a first set of parameters associated with a three-dimensional (3D) facial representation of the face for the keyframe; reconstructing a second set of parameters associated with a 3D facial representation of the face for each of the one or more inter frames according to the compressed inter prediction residual of the second set of parameters; and A method comprising generating a video including a face based on the keyframes, the first set of parameters, and the second set of parameters.
[0188] 2. The method of clause 1, wherein the compressed frames are coded according to one of the following standards: Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC).
[0189] 3. The method according to clause 1 or 2, wherein the first and second sets of parameters include the following items of the face: expression, translation and angle.
[0190] 4. The method of clause 3, further comprising identifying a face identifier in the keyframe, wherein the first set of parameters further comprises the face identifier, and the second set of parameters comprises the face identifier by inheriting from the first set of parameters.
[0191] 5. Generating a video including a face based on keyframes, a first set of parameters, and a second set of parameters includes: generating a keyframe mesh according to a first set of parameters by a 3D morphable model (3DMM); generating an interframe mesh for each interframe of the one or more interframes according to the second set of parameters with the 3DMM; determining a dense motion map and an occlusion map for each key frame and one or more inter frames according to the key frame mesh and the inter frame mesh; and 5. The method of any one of clauses 1 to 4, comprising generating video according to keyframes, a dense motion map, and an occlusion map.
[0192] 6. Determining a dense motion map and an occlusion map according to the keyframe mesh and the interframe mesh for each of the keyframe and one or more interframes, Obtaining a coarse motion field of the face according to the key-frame mesh and the inter-frame mesh; generating coarsely deformed frames by inputting the keyframes and the coarse motion field into a deep neural network; and 6. The method of clause 5, comprising estimating the dense motion map and the occlusion map by concatenating the gross motion frames, the key frames, and the gross motion field into a deep neural network.
[0193] 7. The method of clause 6, wherein generating an image including a face based on the keyframes, the first set of parameters, and the second set of parameters includes: warping the keyframes according to the dense motion map; and generating the image by calculating the Hadamard product of the occlusion map and the warping result.
[0194] 8. The method of clause 6, wherein the video is generated by a generation module, the generated video including a corresponding reconstructed frame for each of the one or more interframes, and the deep neural network or the generation module is updated by at least one of the following: a perceptual loss, an adversarial loss, an identity-preserving loss, and a reconstructed texture loss between each of the one or more interframes and its corresponding reconstructed frame.
[0195] 9. The method of any one of clauses 5 to 8, wherein the 3DMM is a weakly-supervised multi-face 3D reconstruction (WM3DR) model.
[0196] 10. The method of any one of clauses 1 to 9, further comprising receiving a bitstream comprising the compressed frame and an inter prediction residual for each inter frame of the one or more inter frames.
[0197] 11. A method for encoding video data, comprising: compressing keyframes representing a face; extracting a set of parameters associated with a three-dimensional (3D) facial expression from one or more interframes; and compressing an inter-prediction residual of the set of parameters.
[0198] 12. The method according to clause 11, wherein the key frames are compressed according to one of the following standards: Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC).
[0199] 13. The method according to clause 11 or 12, wherein the set of parameters includes the following items of the face: expression, translation or angle.
[0200] 14. A method according to any one of clauses 11 to 13, wherein compressing the inter-prediction residual of a set of parameters comprises: determining a difference between the sets of parameters of two adjacent frames; and encoding the difference as a compressed inter-prediction residual of the set of parameters.
[0201] 15. A method according to any one of clauses 11 to 14, wherein the set of parameters is extracted by a 3D morphable model (3DMM).
[0202] 16. The method of clause 15, wherein the 3DMM is a weakly-supervised multi-face 3D reconstruction (WM3DR) model.
[0203] 17. The method of any one of clauses 11 to 16, further comprising generating a bitstream comprising compressed keyframes and a set of parameter compressed inter prediction residuals.
[0204] 18. The method of clause 17, further comprising transmitting the bitstream over a public or private network.
[0205] 19. A method according to any one of clauses 11 to 18, wherein the key frame and one or more inter frames are generated in sequence by an image capture device or an array of image capture devices.
[0206] 20. A non-transitory computer-readable storage medium storing a bitstream of video, the bitstream including a compressed frame and an inter-prediction residual for each inter-frame of one or more inter-frames, the compressed inter-prediction frame and the compressed residual, when decoded by the decoder, providing to the decoder: decompressing the compressed frames to generate keyframes representing a face; generating a first set of parameters associated with a three-dimensional (3D) facial representation of the face for the keyframe; reconstructing a second set of parameters associated with a 3D facial representation of the face for each of the one or more inter frames according to the inter prediction residual; and A medium that performs a method that includes generating a facial image based on keyframes, a first set of parameters, and a second set of parameters.
[0207] 21. The medium of clause 20, wherein the compressed frames are coded according to one of the following standards: Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC).
[0208] 22. The medium of clause 20 or 21, wherein the first and second sets of parameters include the following items of the face: expression, translation and angle.
[0209] 23. The medium of clause 22, wherein the method further includes identifying a facial identifier in the keyframe, wherein the first set of parameters further includes the facial identifier, and wherein the second set of parameters includes the facial identifier by inheriting from the first set of parameters.
[0210] 24. Generating a video including a face based on keyframes, a first set of parameters, and a second set of parameters includes: generating a keyframe mesh according to a first set of parameters by a 3D morphable model (3DMM); generating an interframe mesh for each interframe of the one or more interframes according to the second set of parameters with the 3DMM; determining a dense motion map and an occlusion map for each key frame and one or more inter frames according to the key frame mesh and the inter frame mesh; and 24. The medium of any one of clauses 20 to 23, comprising generating video according to keyframes, dense motion maps, and occlusion maps.
[0211] 25. Determining a dense motion map and an occlusion map according to the keyframe mesh and the interframe mesh for each frame of the keyframe and one or more interframes includes: Obtaining a coarse motion field of the face according to the key-frame mesh and the inter-frame mesh; generating coarsely deformed frames by inputting the keyframes and the coarse motion field into a deep neural network; and 25. The medium of clause 24, comprising estimating the dense motion map and occlusion map by concatenating the gross motion frames, key frames, and gross motion fields into a deep neural network.
[0212] 26. The medium described in clause 25, wherein generating an image including a face based on keyframes, a first set of parameters, and a second set of parameters includes: warping the keyframes according to a dense motion map; and generating the image by calculating a Hadamard product of the occlusion map and the warping result.
[0213] 27. An image is generated by a generation module, the generated image including a corresponding reconstructed frame for each inter-frame of the one or more inter-frames; The medium of clause 25, wherein the deep neural network or generative module is updated by at least one of the following: a perceptual loss, an adversarial loss, an identity-preserving loss, and a reconstruction texture loss between each of the one or more inter frames and its corresponding reconstructed frame.
[0214] 28. The medium of any one of clauses 24 to 27, wherein the 3DMM is a weakly-supervised multi-face 3D reconstruction (WM3DR) model.
[0215] 29. The medium of any one of clauses 20-29, wherein the method further comprises: receiving a bitstream comprising the compressed frame and an inter prediction residual for each inter frame of the one or more inter frames.
[0216] 30. An apparatus for decoding video data, comprising: a decoder configured to decompress the compressed frames to generate keyframes representing a face; an extractor configured to generate a first set of parameters associated with a three-dimensional (3D) facial representation of a face for a key frame; and to reconstruct a second set of parameters associated with the 3D facial representation of the face for each of one or more inter frames according to a compressed inter prediction residual of the second set of parameters; and An apparatus comprising: a generator configured to generate a video including a face based on the keyframes, a first set of parameters, and a second set of parameters.
[0217] 31. The apparatus of clause 30, wherein the compressed frames are coded according to one of the following standards: Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC).
[0218] 32. The apparatus of clause 30 or 31, wherein the first and second sets of parameters include the following items of the face: expression, translation and angle.
[0219] 33. The apparatus of clause 32, wherein the extractor is further configured to identify a facial identifier in the keyframe, the first set of parameters further including the facial identifier, and the second set of parameters including the facial identifier by inheriting from the first set of parameters.
[0220] 34. The generator includes a generation module, the generator further comprising: generating a keyframe mesh according to a first set of parameters by a 3D morphable model (3DMM); generating an interframe mesh for each interframe of the one or more interframes according to the second set of parameters with the 3DMM; determining a dense motion map and an occlusion map for each key frame and one or more inter frames according to the key frame mesh and the inter frame mesh; and generating an image by a generation module according to the keyframes, the dense motion map, and the occlusion map; 34. The apparatus of any one of clauses 30 to 33, configured to:
[0221] 35. The generator further: Obtaining a coarse motion field of the face according to the key-frame mesh and the inter-frame mesh; generating coarsely deformed frames by inputting the keyframes and the coarse motion field into a deep neural network; and Estimating dense motion and occlusion maps by concatenating gross motion frames, keyframes, and gross motion fields into a deep neural network 35. The apparatus of clause 34, configured to:
[0222] 36. The generator further warps the keyframes according to the dense motion map; and generates the image by calculating the Hadamard product of the occlusion map and the warped result. 36. The apparatus of clause 35, configured to:
[0223] 37. The apparatus of clause 35, wherein the generated video includes a corresponding reconstructed frame for each of the one or more interframes, and the deep neural network or generation module is updated by at least one of the following: a perceptual loss, an adversarial loss, an identity-preserving loss, and a reconstructed texture loss between each of the one or more interframes and its corresponding reconstructed frame.
[0224] 38. An apparatus described in any one of clauses 34 to 37, wherein the 3DMM is a weakly-supervised multi-face 3D reconstruction (WM3DR) model.
[0225] 39. The apparatus of any one of clauses 30 to 38, further comprising: a receiver configured to receive a bitstream comprising the compressed frame and an inter prediction residual for each inter frame of the one or more inter frames.
[0226] 40. An apparatus for encoding video data, comprising: an encoder configured to compress key frames representing a face; an extractor configured to extract a set of parameters related to a three-dimensional (3D) facial expression from one or more inter frames; and an encoding module configured to compress an inter prediction residual of the set of parameters.
[0227] 41. The apparatus of clause 40, wherein the key frames are compressed according to one of the following standards: Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC).
[0228] 42. The apparatus according to clause 40 or 41, wherein the set of parameters includes the following items of the face: expression, translation or angle.
[0229] 43. The encoding module further comprises: determining a difference between the set of parameters of two adjacent frames; and encoding the difference as a compressed inter-prediction residual of the set of parameters. 43. The apparatus of any one of clauses 40 to 42, configured to:
[0230] 44. An apparatus according to any one of clauses 40 to 43, wherein the set of parameters is extracted by a 3D morphable model (3DMM).
[0231] 45. The apparatus of clause 44, wherein the 3DMM is a weakly-supervised multi-face 3D reconstruction (WM3DR) model.
[0232] 46. The apparatus of any one of clauses 40 to 45, wherein the encoder is further configured to generate a bitstream including compressed keyframes and a compressed inter-prediction residual of the set of parameters.
[0233] 47. The apparatus of clause 46, wherein the encoder is further configured to transmit the bitstream over a public or private network.
[0234] 48. An apparatus according to any one of clauses 40 to 47, wherein the key frame and one or more inter frames are generated in sequence by an image capture device or an array of image capture devices.
[0235] 49. An apparatus for decoding video data comprising: a memory configured to store instructions; and one or more processors, the one or more processors: decompressing the compressed frames to generate keyframes representing a face; generating a first set of parameters associated with a three-dimensional (3D) facial representation of the face for the keyframe; reconstructing a second set of parameters associated with a 3D facial representation of the face for each of the one or more inter frames according to the compressed inter prediction residual of the second set of parameters; and Generating a video including a face based on the keyframes, the first set of parameters, and the second set of parameters 10. An apparatus configured to execute instructions to cause the apparatus to:
[0236] 50. The apparatus of clause 49, wherein the compressed frames are coded according to one of the following standards: Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC).
[0237] 51. The apparatus of clause 49 or 50, wherein the first and second sets of parameters include the following items of the face: expression, translation and angle.
[0238] 52. The apparatus of clause 51, wherein the apparatus further identifies a face identifier in the key frame, the first set of parameters further including the face identifier, and the second set of parameters includes the face identifier by inheriting from the first set of parameters.
[0239] 53. Generating a video including a face based on keyframes, a first set of parameters, and a second set of parameters includes: generating a keyframe mesh according to a first set of parameters by a 3D morphable model (3DMM); generating an interframe mesh for each interframe of the one or more interframes according to the second set of parameters with the 3DMM; determining a dense motion map and an occlusion map for each key frame and one or more inter frames according to the key frame mesh and the inter frame mesh; and 53. The apparatus of any one of clauses 49 to 52, comprising generating video according to keyframes, a dense motion map, and an occlusion map.
[0240] 54. Determining a dense motion map and an occlusion map according to the keyframe mesh and the interframe mesh for each frame of the keyframe and one or more interframes includes: Obtaining a coarse motion field of the face according to the key-frame mesh and the inter-frame mesh; generating coarsely deformed frames by inputting the keyframes and the coarse motion field into a deep neural network; and 54. The apparatus of clause 53, comprising estimating the dense motion map and the occlusion map by concatenating the gross motion frames, the key frames, and the gross motion field into a deep neural network.
[0241] 55. The apparatus described in clause 54, wherein generating an image including a face based on the keyframes, the first set of parameters, and the second set of parameters includes: warping the keyframes according to the dense motion map; and generating the image by calculating the Hadamard product of the occlusion map and the warping result.
[0242] 56. The image is generated by a generation module; the generated video includes a corresponding reconstructed frame for each inter-frame of the one or more inter-frames; The apparatus of clause 54, wherein the deep neural network or generative module is updated by at least one of the following: a perceptual loss, an adversarial loss, an identity-preserving loss, and a reconstruction texture loss between each of the one or more inter frames and its corresponding reconstructed frame.
[0243] 57. An apparatus described in any one of clauses 53 to 56, wherein the 3DMM is a weakly-supervised multi-face 3D reconstruction (WM3DR) model.
[0244] 58. The apparatus of any one of clauses 49 to 57, wherein the apparatus is further configured to: receive a bitstream comprising the compressed frame and an inter prediction residual for each inter frame of the one or more inter frames.
[0245] 59. An apparatus for encoding video data, comprising: a memory configured to store instructions; and one or more processors, the one or more processors: Extracting a set of parameters related to a three-dimensional (3D) facial expression from one or more interframes; and Compressing inter-prediction residuals of a set of parameters - Patent Application 20070122997 10. An apparatus configured to execute instructions to cause the apparatus to:
[0246] 60. The apparatus of clause 59, wherein the key frames are compressed in accordance with one of the following standards: Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC).
[0247] 61. The device according to clause 59 or 60, wherein the set of parameters includes the following items of the face: expression, translation or angle.
[0248] 62. An apparatus described in any one of clauses 59 to 61, wherein compressing an inter-prediction residual of a set of parameters comprises: determining a difference between a set of parameters of two adjacent frames; and encoding the difference as a compressed inter-prediction residual of the set of parameters.
[0249] 63. An apparatus according to any one of clauses 59 to 62, wherein the set of parameters is extracted by a 3D morphable model (3DMM).
[0250] 64. The apparatus of clause 63, wherein the 3DMM is a weakly-supervised multi-face 3D reconstruction (WM3DR) model.
[0251] 65. The apparatus of any one of clauses 59 to 64, wherein the apparatus is further configured to generate a bitstream comprising compressed keyframes and a set of parameter compressed inter prediction residuals.
[0252] 66. The apparatus of clause 65, further comprising: transmitting the bitstream over a public or private network.
[0253] 67. An apparatus according to any one of clauses 59 to 66, wherein the key frame and one or more inter frames are generated in sequence by an image capture device or an array of image capture devices.
[0254]
[0205] It should be noted that relational terms herein, such as "first" and "second," are used only to distinguish one entity or action from another, and thus do not require or imply any actual relationship or order between those entities or actions. Furthermore, the words "comprising," "having," "containing," and "including," and other similar forms, are intended to be equivalent in meaning, and the item or items following any of these words are intended to be open-ended in that they do not imply an exclusive listing of such items or items or that they are limited to only the listed item or items.
[0255] As used herein, unless otherwise stated, the term "or" includes all possible combinations unless impracticable. For example, if it is stated that a database may include A or B, then the database may include A, B, A and B, unless otherwise stated or impracticable. As a second example, if it is stated that a database may include A, B, or C, then the database may include A, B, C, A and B, A and C, B and C, A and B and C, unless otherwise stated or impracticable.
[0256]
[0207] It is understood that the above-described embodiments can be implemented by hardware, or software (program code), or a combination of hardware and software. If implemented by software, the software can be stored in the above-described computer-readable medium. The software, when executed by a processor, can perform the disclosed methods. The computing units and other functional units described in this disclosure can be implemented by hardware, or software, or a combination of hardware and software. Those skilled in the art will also understand that multiple of the above-described modules / units can be combined into one module / unit, and that each of the above-described modules / units can be further divided into multiple sub-modules / sub-units.
[0257]
[0208] In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Some adaptations and modifications of the above-described embodiments may be made. Other embodiments may become apparent to those skilled in the art from consideration of the description and practice of the present disclosure disclosed herein. It is intended that the specification and examples be considered merely exemplary, with the true scope and spirit of the present disclosure being indicated by the appended claims. It is also intended that the sequence of steps shown in the figures is for illustrative purposes only and is not intended to be limited to any particular sequence of steps. Thus, one skilled in the art will appreciate that these steps may be performed in different orders while implementing the same method.
[0258]
[0209] Illustrative embodiments have been disclosed in the accompanying drawings and herein. However, many variations and modifications may be made to these embodiments. Accordingly, although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation.
Claims
1. 1. A method of decoding video data, comprising: decompressing the compressed frames to generate keyframes representing a face; generating a first set of parameters for the keyframes associated with a three-dimensional (3D) facial representation of the face; reconstructing a second set of parameters associated with a 3D facial representation of the face for each of one or more inter frames according to a compressed inter prediction residual of the second set of parameters; and generating a video including the face based on the keyframes, the first set of parameters, and the second set of parameters.
2. The method of claim 1 , wherein the compressed frames are encoded according to one of the following standards: AVC, HEVC, or VVC.
3. The method of claim 1 or 2, wherein the first and second sets of parameters include at least one of an expression, a translation, and an angle.
4. 4. The method of claim 3, wherein the method further comprises identifying an identifier of the face in the keyframe, the first set of parameters further comprising the identifier of the face, and the second set of parameters comprising the identifier of the face by inheriting from the first set of parameters.
5. Generating the video including the face based on the keyframes, the first set of parameters, and the second set of parameters includes: generating a keyframe mesh according to the first set of parameters using a 3D morphable model (3DMM); generating, by the 3DMM, an inter-frame mesh for each inter-frame of the one or more inter-frames according to the second set of parameters; determining a dense motion map and an occlusion map for each of the key frames and the one or more inter frames according to the key frame mesh and the inter frame mesh; and The method of claim 1 , further comprising generating the video according to the keyframes, the dense motion map, and the occlusion map.
6. determining the dense motion map and the occlusion map for each of the key frames and the one or more inter frames according to the key frame mesh and the inter frame mesh; obtaining a gross motion field of the face according to the key-frame mesh and the inter-frame mesh; generating coarsely deformed frames by inputting the keyframes and the coarse motion field into a deep neural network; and The method of claim 5 , comprising estimating the dense motion map and the occlusion map by concatenating the gross motion frames, the key frames, and the gross motion field to the deep neural network.
7. Generating the video including the face based on the keyframes, the first set of parameters, and the second set of parameters includes: warping the keyframes according to the dense motion map; and The method of claim 6 , comprising generating the image by computing a Hadamard product of the occlusion map and the result of the warping.
8. the video is generated by a generation module; the generated video includes, for each inter-frame of one or more inter-frames, a corresponding reconstructed frame; 7. The method of claim 6, wherein the deep neural network or the generation module is updated with at least one of a perceptual loss, an adversarial loss, an identity-preserving loss, and a reconstruction texture loss between each of the one or more inter frames and its corresponding reconstructed frame.
9. The method of any one of claims 5 to 8, wherein the 3DMM is a weakly-supervised multi-face 3D reconstruction (WM3DR) model.
10. The method of claim 1 , further comprising receiving a bitstream comprising the compressed frames and the inter prediction residual for each inter frame of the one or more inter frames.
11. 1. A method of encoding video data, comprising: Compressing keyframes representing faces; Extracting a set of parameters related to a three-dimensional (3D) facial expression from one or more interframes; and compressing an inter-prediction residual of the set of parameters.
12. The method of claim 11 , wherein the key frames are encoded according to one of the following standards: AVC, HEVC, or VVC.
13. The method of claim 11 or 12, wherein the set of parameters includes at least one of an expression, a translation, or an angle.
14. Compressing the inter prediction residual of the set of parameters comprises: determining a difference between the set of parameters of two adjacent frames; and 14. The method of any one of claims 11 to 13, comprising encoding the difference as the compressed inter prediction residual of the set of parameters.
15. 15. The method of any one of claims 11 to 14, wherein the set of parameters is extracted by a 3D morphable model (3DMM).
16. The method of claim 15 , wherein the 3DMM is a weakly-supervised multi-face 3D reconstruction (WM3DR) model.
17. The method of claim 11 , further comprising generating a bitstream comprising the compressed keyframes and the compressed inter-prediction residual of the set of parameters.
18. 20. The method of claim 17, further comprising transmitting the bitstream over a public or private network.
19. 19. The method of any one of claims 11 to 18, wherein the key frame and the one or more Inter frames are generated in turn by an image capture device or an array of image capture devices.
20. 1. An apparatus for decoding video data, comprising: a decoder configured to decompress the compressed frames to generate key frames representing a face; an extractor configured to generate a first set of parameters associated with a three-dimensional (3D) facial representation of the face for the keyframe; and to reconstruct a second set of parameters associated with the 3D facial representation of the face for each of one or more interframes according to a compressed inter prediction residual of the second set of parameters; and a generator configured to generate a video including the face based on the keyframes, the first set of parameters, and the second set of parameters.
21. 21. The apparatus of claim 20, wherein the compressed frames are encoded according to one of the following standards: AVC, HEVC, and VVC.
22. 22. The apparatus of claim 20 or 21, wherein the first and second sets of parameters include at least one of an expression, a translation, and an angle.
23. 23. The apparatus of claim 22, further comprising: an identification module configured to identify an identifier of the face in the keyframe, wherein the first set of parameters further comprises the identifier of the face, and the second set of parameters comprises the identifier of the face by inheriting from the first set of parameters.
24. The generator: generating a keyframe mesh according to the first set of parameters using a 3D morphable model (3DMM); generating an interframe mesh for each interframe of one or more interframes according to the second set of parameters by the 3DMM; determining a dense motion map and an occlusion map according to the key-frame mesh and the inter-frame mesh for each of the key-frames and the one or more inter-frames; 24. The device of any one of claims 20 to 23, configured to generate the video according to the keyframes, the dense motion map and the occlusion map.
25. The generator: Obtain a gross motion field of the face according to the key-frame mesh and the inter-frame mesh; generating coarsely deformed frames by inputting the keyframes and the coarse motion field into a deep neural network; 25. The apparatus of claim 24, configured to estimate the dense motion map and the occlusion map by concatenating the coarse motion frames, the key frames, and the coarse motion field to the deep neural network.
26. The generator: warping the keyframes according to the dense motion map; and 26. The apparatus of claim 25, configured to generate the image by computing a Hadamard product of the occlusion map and the warp result.
27. the video is generated by a generation module; the generated video includes, for each inter-frame of one or more inter-frames, a corresponding reconstructed frame; 26. The apparatus of claim 25, wherein the deep neural network or the generation module is updated with at least one of a perceptual loss, an adversarial loss, an identity-preserving loss, and a reconstruction texture loss between each of the one or more inter frames and its corresponding reconstructed frame.
28. 28. The apparatus of claim 24, wherein the 3DMM is a weakly-supervised multi-face 3D reconstruction (WM3DR) model.
29. 29. The apparatus of claim 20, further comprising: a receiver configured to receive a bitstream comprising the compressed frames and the inter prediction residual for each inter frame of the one or more inter frames.
30. 1. An apparatus for encoding video data, comprising: an encoder configured to compress keyframes representing faces; an extractor configured to extract a set of parameters related to a three-dimensional (3D) facial expression from one or more interframes; and an encoding module configured to compress an inter-prediction residual of the set of parameters.
31. 31. The apparatus of claim 30, wherein the key frames are encoded according to one of the following standards: AVC, HEVC, and VVC.
32. 32. The apparatus of claim 30 or 31, wherein the set of parameters includes at least one of an expression, a translation, or an angle.
33. 33. The apparatus of claim 30, wherein the encoding module is configured to: determine a difference between the sets of parameters of two adjacent frames; and encode the difference as the compressed inter-prediction residual of the sets of parameters.
34. 34. The apparatus of any one of claims 30 to 33, wherein the set of parameters is extracted by a 3D morphable model (3DMM).
35. 35. The apparatus of claim 34, wherein the 3DMM is a weakly-supervised multi-face 3D reconstruction (WM3DR) model.
36. 36. The apparatus of any one of claims 30 to 35, further comprising a generator configured to generate a bitstream comprising the compressed keyframes and the compressed inter-prediction residual of the set of parameters.
37. 37. The apparatus of claim 36, further comprising a transmitter configured to transmit the bitstream over a public or private network.
38. 38. The apparatus of any one of claims 30 to 37, wherein the key frame and the one or more Inter frames are generated in turn by an image capture device or an array of image capture devices.
39. 1. An apparatus for decoding video data, comprising: a memory configured to store instructions; and 11. An apparatus comprising one or more processors configured to execute the instructions to cause the apparatus to perform the method of any one of claims 1 to 10.
40. 1. An apparatus for encoding video data, comprising: a memory configured to store instructions; and 20. An apparatus comprising one or more processors configured to execute said instructions to cause said apparatus to perform the method of any one of claims 11 to 19.
41. 1. A non-transitory computer-readable storage medium for storing a video bitstream, comprising: the bitstream includes a compressed frame and an inter prediction residual for each inter frame of one or more inter frames; A non-transitory computer-readable storage medium, wherein the compressed frame and compressed inter-prediction residual, when decoded by a decoder, cause the decoder to perform the method of any one of claims 1 to 10.
42. 1. A non-transitory computer-readable storage medium for storing a video bitstream, comprising: the bitstream includes a compressed frame and an inter prediction residual for each inter frame of one or more inter frames; 20. A non-transitory computer-readable storage medium, wherein the compressed frame and compressed inter-prediction residual, when decoded by a decoder, cause the decoder to perform the method of any one of claims 11 to 19.
43. A computer program product comprising computer program instructions enabling a computer to carry out the method of any one of claims 1 to 10.
44. A computer program product comprising computer program instructions enabling a computer to carry out the method of any one of claims 11 to 19.
45. A computer program enabling a computer to carry out the method according to any one of claims 1 to 10.
46. A computer program enabling a computer to carry out the method according to any one of claims 11 to 19.