METHOD AND APPARATUS FOR USING COMPRESSED SEI MESSAGES FOR FACIAL VIDEO GENERATION - Patent application
By employing facial video generation compression SEI messages, the method addresses the challenge of high compression efficiency in advanced video coding standards, particularly for facial video data, achieving improved quality at reduced bandwidth.
Patent Information
- Application Number
- JP2025534139
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2023-12-27
- Publication Date
- 2026-01-27
AI Technical Summary
Existing video coding standards face challenges in achieving high compression efficiency while maintaining subjective quality, particularly with the development of advanced standards like VVC/H.266, which require improved methods for encoding and decoding facial video data.
The method involves using facial video generation compression supplemental enhancement information (SEI) messages to identify and reconstruct face pictures, incorporating identification numbers to determine the use of a face video generation compression scheme, and signaling whether such a scheme is used during encoding and decoding processes.
This approach enhances compression efficiency by specifically handling facial video data, allowing for higher quality at reduced bandwidth, aligning with the goals of advanced video coding standards like VVC/H.266.
Smart Images

Figure 2026502830000001_ABST
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This disclosure claims the benefit of priority to U.S. Provisional Application No. 63 / 436,626, filed January 1, 2023, and claims the benefit of U.S. Patent Application No. 18 / 392,557, entitled "Method and Apparatus for Using Facial Image Generation Compressed SEI Messages," filed December 21, 2023. Both of the above applications are incorporated herein by reference in their entireties.
[0002] The present disclosure relates generally to video processing, and more specifically to methods and apparatus for using facial video generation compression supplemental enhancement information (SEI) messages. [Background technology]
[0003] Video is a set of static pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and decompressed before display. The compression process is typically called encoding, and the decompression process is typically called decoding. There are a variety of video coding formats that use standardized video coding techniques, most commonly based on prediction, transform, quantization, entropy coding, and in-loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, which specify specific video coding formats, are developed by standardization organizations. As video standards adopt increasingly advanced video coding techniques, the coding efficiency of new video coding standards becomes increasingly higher. Summary of the Invention
[0004] In a first aspect, an embodiment of the present disclosure provides a method for decoding a bitstream to output one or more pictures for a video stream, the method including: receiving a bitstream; and decoding one or more pictures using coding information of the bitstream. The decoding includes determining whether a face video generation compression scheme is used based on an identification number; decoding a supplemental enhancement information (SEI) message including face information in response to determining that the face video generation compression scheme is used; and reconstructing a face picture based on the face information and a base picture associated with the SEI message.
[0005] In a second aspect, an embodiment of the present disclosure provides a method for encoding a video sequence into a bitstream, the method including receiving a video sequence, encoding one or more pictures of the video sequence, and generating a bitstream, the encoding including signaling an identification number indicating whether a facial video generation compression scheme is used.
[0006] In a third aspect, an embodiment of the present disclosure provides a decoding device including: a receiving module configured to receive a bitstream; and a decoding module configured to decode one or more pictures using coding information of the bitstream, wherein the decoding module is configured to determine whether a face video generation compression scheme is used based on an identification number, and in response to determining that the face video generation compression scheme is used, decode an SEI message including face information, and reconstruct a face picture based on the face information and a base picture associated with the SEI message.
[0007] In a fourth aspect, an embodiment of the present disclosure provides an encoding device including: a receiving module configured to receive a video sequence; an encoding module configured to encode one or more pictures of the video sequence; and a generating module configured to generate a bitstream, wherein the encoding module is configured to signal an identification number indicating whether a facial video generating compression scheme is used.
[0008] In a fifth aspect, an embodiment of the present disclosure provides an electronic device including: a memory storing a set of instructions; and one or more processors configured to execute the set of instructions to cause the one or more processors to perform the method of decoding a bitstream described in the first aspect and outputting one or more pictures for a video stream.
[0009] In a sixth aspect, an embodiment of the present disclosure provides an electronic device including: a memory storing a set of instructions; and one or more processors configured to execute the set of instructions to cause the one or more processors to perform the method of encoding a video sequence into a bitstream described in the second aspect.
[0010] In a seventh aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing a video bitstream that, when decoded by a processor, causes the processor to perform the method of decoding the bitstream and outputting one or more pictures for the video stream described in the first aspect.
[0011] In an eighth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing a video sequence of video that, when encoded by a processor, causes the processor to perform the method of encoding a video sequence into a bitstream described in the second aspect.
[0012] In a ninth aspect, an embodiment of the present disclosure provides a computer program product including computer program instructions that enable a computer to perform the method of decoding a bitstream according to the first aspect to output one or more pictures for a video stream. In a tenth aspect, an embodiment of the present disclosure provides a computer program product comprising computer program instructions that enable a computer to perform the method of encoding a video sequence into a bitstream according to the second aspect. In an eleventh aspect, an embodiment of the present disclosure provides a computer program enabling a computer to perform the method of decoding a bitstream according to the first aspect to output one or more pictures for a video stream. In a twelfth aspect, an embodiment of the present disclosure provides a computer program enabling a computer to perform the method for encoding a video sequence into a bitstream according to the second aspect.
[0013] Embodiments and various aspects of the present disclosure are illustrated in the following detailed description and accompanying drawings, in which various features illustrated in the drawings are not drawn to scale. [Brief explanation of the drawings]
[0014] FIG. 1 is a schematic diagram illustrating an exemplary system for pre-processing and coding image data according to some embodiments of the present disclosure.
[0015] FIG. 2A is a schematic diagram illustrating an exemplary encoding process of a hybrid video coding system, according to an embodiment of the present disclosure.
[0016] FIG. 2B is a schematic diagram illustrating another exemplary encoding process of a hybrid video coding system, according to an embodiment of the present disclosure.
[0017] FIG. 3A is a schematic diagram illustrating an exemplary decoding process for a hybrid video coding system, according to an embodiment of the present disclosure.
[0018] FIG. 3B is a schematic diagram illustrating another exemplary decoding process for a hybrid video coding system, according to an embodiment of the present disclosure.
[0019] FIG. 4 is a block diagram of an exemplary apparatus for preprocessing or coding image data according to some embodiments of the present disclosure.
[0020] FIG. 5 is a schematic diagram illustrating an exemplary deep learning-based video generative compression framework according to some embodiments of the present disclosure.
[0021] FIG. 6 is a schematic diagram illustrating an exemplary encoder-decoder coding framework with a compact feature size of 1x4x4 for talking face video, according to some embodiments of the present disclosure.
[0022] FIG. 7 is a schematic diagram illustrating a general encoder-decoder generated compression framework for 3DMM-assisted talking face video, according to some embodiments of the present disclosure.
[0023] FIG. 8 is a flowchart of an exemplary method for processing video based on a facial video generation compression supplemental enhancement information (SEI) message according to some embodiments of the present disclosure.
[0024] FIG. 9 is a flowchart of an exemplary method for processing video based on a facial video generation compression supplemental enhancement information (SEI) message according to some embodiments of the present disclosure.
[0025] FIG. 10A illustrates an example syntax of the disclosed facial image generation compression SEI message according to some embodiments of the present disclosure.
[0026] FIG. 10B illustrates an example syntax of the disclosed facial image generation compression SEI message according to some embodiments of the present disclosure.
[0027] FIG. 11 is a flowchart of an exemplary method for processing video based on a facial video generation compression supplemental enhancement information (SEI) message according to some embodiments of the present disclosure.
[0028] FIG. 12A illustrates another example syntax of the disclosed facial image generation compression SEI message according to some embodiments of the present disclosure.
[0029] FIG. 12B illustrates another example syntax of the disclosed facial image generation compression SEI message according to some embodiments of the present disclosure.
[0030] FIG. 13 is a flowchart of an exemplary method for processing video based on a facial video generation compression supplemental enhancement information (SEI) message according to some embodiments of the present disclosure.
[0031] FIG. 14A illustrates another example syntax of the disclosed facial image generation compression SEI message according to some embodiments of the present disclosure.
[0032] FIG. 14B illustrates another example syntax of the disclosed facial image generation compression SEI message according to some embodiments of the present disclosure.
[0033] FIG. 15 is a flowchart of an exemplary method for processing video based on a facial video generation compression supplemental enhancement information (SEI) message according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0034] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description will refer to the accompanying drawings, in which like numbers in different drawings represent the same or similar elements unless otherwise stated. The implementations set forth in the following description of exemplary embodiments do not represent all implementations according to the present disclosure. Instead, they are merely examples of apparatus and methods according to aspects of the present disclosure as recited in the claims. Specific aspects of the present disclosure are described in more detail below. In the event of a conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein shall control.
[0035] The ITU-T Video Coding Expert Group (ITU-T VCEG) and the ISO / IEC Moving Picture Expert Group (ISO / IEC MPEG) Joint Video Experts Team (JVET) are currently developing the General Purpose Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 at half the bandwidth.
[0036] To achieve the same subjective quality as HEVC / H.265 at half the bandwidth, JVET has developed technology beyond HEVC using the Joint Exploration Model (JEM) reference software. As coding techniques are incorporated into JEM, JEM achieves substantially higher coding performance than HEVC.
[0037] The VVC standard has evolved recently and continues to include more coding techniques that provide better compression performance. VVC is based on the same hybrid video coding system used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.
[0038] Video is a set of static pictures (or "frames") arranged in a temporal sequence to store visual information. A video capture device (e.g., a camera) can be used to capture and store the pictures in a temporal sequence, and a video playback device (e.g., a television, a computer, a smartphone, a tablet computer, a video player, or any end-user terminal with display capabilities) can be used to display the pictures in a temporal sequence. In some applications, the video capture device can also transmit the captured video in real time to a video playback device (e.g., a computer with a monitor) for purposes such as surveillance, conferencing, or live broadcasting.
[0039] To reduce the storage space and transmission bandwidth required for such applications, video can be compressed before storage and transmission and decompressed before display. Compression and decompression can be achieved by software executed by a processor (e.g., a general-purpose computer processor) or dedicated hardware. A module for compression is typically called an "encoder," and a module for decompression is typically called a "decoder." Encoders and decoders may be collectively referred to as a "codec." Encoders and decoders can be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of encoders and decoders may include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of encoders and decoders may include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed on a computer-readable medium. Video compression and decompression can be achieved by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, and the H.26x series. In some applications, a codec can decompress video from a first coding standard and recompress the decompressed video in a second coding standard, in which case the codec can be called a "transcoder."
[0040] A video coding process can identify and retain useful information that can be used to reconstruct an image and ignore information that is not important for reconstruction. If the ignored, unimportant information cannot be perfectly reconstructed, such a coding process can be called "lossy." Otherwise, it can be called "lossless." Most coding processes are lossy, which is a tradeoff to reduce the required storage space and transmission bandwidth.
[0041] Useful information about the picture being coded (called the "current picture") includes changes relative to a reference picture (e.g., a previously coded and reconstructed picture). Such changes can include pixel position changes, luminance changes, or color changes, of which position changes are of primary concern. Position changes of pixels representing an object can reflect the motion of the object between the reference picture and the current picture.
[0042] A picture coded without reference to another picture (i.e., its own reference picture) is called an "I-picture." If some or all of the blocks in a picture (e.g., blocks that generally refer to portions of a video picture) are predicted with one reference picture (e.g., uniprediction) by intra-prediction or inter-prediction, the picture is called a "P-picture." If at least one block in a picture is predicted with two reference pictures (e.g., biprediction), the picture is called a "B-picture."
[0043] 1 is a schematic diagram illustrating an exemplary system 100 for preprocessing and coding image data according to some embodiments of the present disclosure. Image data may include an image (also called a "picture" or "frame"), multiple images, or a video. An image is a static picture. Multiple images may be spatially or temporally related or unrelated. A video is a set of images arranged in a temporal sequence.
[0044] 1 , system 100 includes a source device 120 that provides encoded video data that is subsequently decoded by a destination device 140. Consistent with disclosed embodiments, each of source device 120 and destination device 140 may include any of a wide range of devices, such as a desktop computer, a notebook (e.g., laptop) computer, a server, a tablet computer, a set-top box, a mobile phone, a vehicle, a camera, an image sensor, a robot, a television, a camera, a wearable device (e.g., a smart watch, a wearable camera), a display device, a digital media player, a video game console, a video streaming device, etc. Source device 120 and destination device 140 may be configured for wireless or wired communication.
[0045] Referring to FIG. 1, source device 120 may include image / video preprocessor 122, image / video encoder 124, and output interface 126. Destination device 140 may include input interface 142, image / video decoder 144, and one or more machine vision applications 146. Image / video preprocessor 122 preprocesses image data, i.e., images or videos, and generates an input bitstream for image / video encoder 124. Image / video encoder 124 encodes the input bitstream and outputs it via output interface 126 as encoded bitstream 162. Encoded bitstream 162 is transmitted via communication medium 160 and received at input interface 142. Image / video decoder 144 decodes encoded bitstream 162 to generate decoded data usable by machine vision application 146.
[0046] More specifically, source device 120 may further include various devices (not shown) for providing source image data to be preprocessed by image / video preprocessor 122. Devices for providing source image data may include an image / video capture device such as a camera, an image / video archive or storage device containing previously captured images / video, or an image / video feed interface that receives images / video from an image / video content provider.
[0047] Image / video encoder 124 and image / video decoder 144 may each be implemented as any of a variety of suitable encoder or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. If encoding or decoding is partially realized in software, image / video encoder 124 or image / video decoder 144 may perform techniques according to this disclosure by storing software instructions on a suitable non-transitory computer-readable medium and executing them in hardware by one or more processors. Each of image / video encoder 124 or image / video decoder 144 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (codec) in the respective device.
[0048] Image / video encoder 124 and image / video decoder 144 may operate according to any video coding standard, such as Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), AOMedia Video 1 (AV1), Joint Photographic Experts Group (JPEG), Moving Picture Experts Group (MPEG), etc. Alternatively, image / video encoder 124 and image / video decoder 144 may be customized devices that do not conform to existing standards. Although not shown in FIG. 1 , in some embodiments, image / video encoder 124 and image / video decoder 144 may be integrated with an audio encoder and decoder, respectively, and may include appropriate MUX-DEMUX units or other hardware and software to handle the encoding of both audio and video in a common data stream or separate data streams.
[0049] Output interface 126 may include any type of medium or device capable of transmitting encoded bitstream 162 from source device 120 to destination device 140. For example, output interface 126 may include a transmitter or transceiver configured to transmit encoded bitstream 162 in real time from source device 120 directly to destination device 140. Encoded bitstream 162 may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device 140.
[0050] Communication medium 160 may include a transient medium, such as a wireless broadcast or a wired network transmission. For example, communication medium 160 may include a radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). Communication medium 160 may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. In some embodiments, communication medium 160 may include routers, switches, base stations, or other equipment useful for facilitating communication from source device 120 to destination device 140. For example, a network server (not shown) may receive encoded bitstream 162 from source device 120, e.g., via a network transmission, and provide encoded bitstream 162 to destination device 140.
[0051] Communication medium 160 may be in the form of a storage medium (e.g., a non-transitory storage medium), such as a hard disk, a flash drive, a compact disc, a digital video disc, a Blu-ray disc, volatile or non-volatile memory, or other suitable digital storage medium for storing encoded image data. In some embodiments, a computing device at a media production facility, such as a disc stamping facility, may receive the encoded image data from source device 120 and create a disc containing the encoded video data.
[0052] Input interface 142 may include any type of medium or device capable of receiving information from communication medium 160. The received information includes encoded bitstream 162. For example, input interface 142 may include a receiver or transceiver configured to receive encoded bitstream 162 in real time.
[0053] The machine vision application 146 may include various hardware or software components for utilizing the decoded image data generated by the image / video decoder 144. For example, the machine vision application 146 may include a display device for displaying the decoded image data to a user, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices. As another example, the machine vision application 146 may include one or more processors for using the decoded image data to perform various machine vision applications, such as object recognition and tracking, facial recognition, image matching, image / video retrieval, augmented reality, robotic vision and navigation, autonomous driving, 3D structure construction, stereo correspondence, motion tracking, etc.
[0054] Image data encoding and decoding techniques (eg, techniques utilized by image / video encoder 124 and image / video decoder 144) will now be described with reference to FIGS. 2A-2B and 3A-3B.
[0055] FIG. 2A illustrates a schematic diagram of an exemplary encoding process 200A according to an embodiment of the present disclosure. For example, encoding process 200A may be performed by an encoder, such as image / video encoder 124 in FIG. 1. As shown in FIG. 2A, the encoder may encode a video sequence 202 into a video bitstream 228 via process 200A. The video sequence 202 may include a set of pictures (referred to as "original pictures") arranged in a temporal order. Each original picture of the video sequence 202 may be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder may perform process 200A at the level of basic processing units for each original picture of the video sequence 202. For example, the encoder may perform process 200A iteratively, where the encoder may encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder may perform process 200A in parallel for each region of each original picture of the video sequence 202.
[0056] 2A , an encoder may generate prediction data 206 and prediction BPU 208 by providing a fundamental processing unit (referred to as an “original BPU”) of an original picture of a video sequence 202 to a prediction stage 204. The encoder may generate residual BPU 210 by subtracting prediction BPU 208 from the original BPU. The encoder may generate quantized transform coefficients 216 by providing residual BPU 210 to a transform stage 212 and a quantization stage 214. The encoder may generate video bitstream 228 by providing prediction data 206 and quantized transform coefficients 216 to a binary coding stage 226. The components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as a “forward pass.” In process 200A, after quantization stage 214, the encoder may generate a reconstructed residual BPU 222 by feeding quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220. The encoder may generate a predicted reference 224 used for the next iteration of process 200A in prediction stage 204 by adding the reconstructed residual BPU 222 to prediction BPU 208. The components 218, 220, 222, and 224 of process 200A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that both the encoder and decoder use the same reference data for prediction.
[0057] The encoder can iteratively perform process 200A to encode each original BPU of the original picture (in the forward pass) and generate a prediction reference 224 for encoding the next original BPU of the original picture (in the reconstruction pass). After encoding all original BPUs of the original picture, the encoder can proceed to encode the next picture in the video sequence 202.
[0058] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, "receive" may refer to receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or operating in any manner to input data.
[0059] In the prediction stage 204, in the current iteration, the encoder may receive the original BPU and a predicted reference 224 and perform a prediction operation to generate predicted data 206 and a predicted BPU 208. The predicted reference 224 may be generated from a reconstruction pass of a previous iteration of the process 200A. The purpose of the prediction stage 204 is to reduce information redundancy by extracting predicted data 206 from the predicted data 206 and the predicted reference 224 that can be used to reconstruct the original BPU as a predicted BPU 208.
[0060] Ideally, predicted BPU 208 would be identical to the original BPU. However, because prediction and reconstruction operations are not ideal, predicted BPU 208 typically differs slightly from the original BPU. To record such differences, an encoder can generate predicted BPU 208 and then subtract it from the original BPU to generate residual BPU 210. For example, the encoder can subtract pixel values (e.g., grayscale or RGB values) of predicted BPU 208 from corresponding pixel values of the original BPU. Each pixel of residual BPU 210 may have a residual value as a result of the subtraction between corresponding pixels of the original BPU and predicted BPU 208. Compared to the original BPU, predicted data 206 and residual BPU 210 may have fewer bits, which can be used to reconstruct the original BPU without significant quality degradation. This compresses the original BPU.
[0061] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce spatial redundancy of the residual BPU 210 by decomposing it into a set of two-dimensional "base patterns," each associated with a "transform coefficient." The base patterns may have the same size (e.g., the size of the residual BPU 210). Each base pattern may represent a variation frequency (e.g., frequency of luminance variation) component of the residual BPU 210. None of the base patterns can be reproduced from any combination (e.g., a linear combination) of any other base patterns. In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. Such a decomposition is analogous to a discrete Fourier transform of a function, in which the base patterns are analogous to the basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are analogous to the coefficients associated with the basis functions.
[0062] Different transform algorithms can use different base patterns. Transform stage 212 can use various transform algorithms, such as a discrete cosine transform or a discrete sine transform. The transform of transform stage 212 is reversible. That is, the encoder can reconstruct residual BPU 210 by inversely performing the transform (called an "inverse transform"). For example, to reconstruct a pixel of residual BPU 210, the inverse transform can multiply each coefficient by the value of the corresponding pixel in the base pattern and add the products to generate a weighted sum. For a video coding standard, both the encoder and decoder can use the same transform algorithm (and thus the same base pattern). Therefore, the encoder can record only the transform coefficients, from which the decoder can reconstruct residual BPU 210 without receiving the base pattern from the encoder. Compared to residual BPU 210, the transform coefficients may have fewer bits, which can be used to reconstruct residual BPU 210 without significant quality degradation. This further compresses residual BPU 210.
[0063] The encoder can further compress the transform coefficients in the quantization stage 214. In the transform process, different base patterns can represent different fluctuation frequencies (e.g., luminance fluctuation frequencies). Because the human eye is generally better at recognizing low-frequency fluctuations, the encoder can ignore high-frequency fluctuation information without significantly degrading the decoding quality. For example, in the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (referred to as a "quantization parameter") and rounding the quotient to the nearest integer. After such an operation, some transform coefficients of the high-frequency base pattern can be converted to zero, and the transform coefficients of the low-frequency base pattern can be converted to smaller integers. The encoder can ignore the zero-valued quantized transform coefficients 216, thereby further compressing the transform coefficients. The quantization process is also lossless, in that the quantized transform coefficients 216 can be reconstructed into transform coefficients in the inverse operation of quantization (referred to as "dequantization").
[0064] The quantization stage 214 can be lossy because the encoder ignores any remainder of such divisions in its rounding operation. Typically, the quantization stage 214 can result in the most information loss in the process 200A. The more information loss, the fewer bits the quantized transform coefficients 216 require. To achieve different levels of information loss, the encoder can use different values of the quantization parameter or any other parameter of the quantization process.
[0065] In binary coding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary coding technique, such as entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or other lossless or lossy compression algorithm. In some embodiments, the encoder may encode other information in binary coding stage 226 in addition to the prediction data 206 and the quantized transform coefficients 216, such as the prediction mode used in prediction stage 204, parameters of the prediction operation, the transform type in transform stage 212, parameters of the quantization process (e.g., quantization parameters), and encoder control parameters (e.g., bitrate control parameters). The encoder may generate a video bitstream 228 using the output data of binary coding stage 226. In some embodiments, the video bitstream 228 may be further packetized for network transmission.
[0066] Referring to the reconstruction path of process 200A, in an inverse quantization stage 218, the encoder may generate reconstructed transform coefficients by performing inverse quantization on the quantized transform coefficients 216. In an inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may generate a predicted reference 224 to be used in the next iteration of process 200A by adding the reconstructed residual BPU 222 to the predicted BPU 208.
[0067] It should be noted that other variations of process 200A can be used to encode video sequence 202. In some embodiments, the stages of process 200A may be performed in a different order by the encoder. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be split into multiple stages. For example, transform stage 212 and quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may omit one or more stages in FIG. 2A.
[0068] 2B shows a schematic diagram of another exemplary encoding process 200B according to an embodiment of the present disclosure. Process 200B can be modified from process 200A. For example, process 200B can be used by an encoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B further includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B further includes a loop filter stage 232 and a buffer 234.
[0069] Generally, prediction techniques can be categorized into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-prediction") can predict a current BPU using pixels from one or more already-coded neighboring BPUs in the same picture. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter-prediction") can predict a current BPU using regions from one or more already-coded pictures. That is, the prediction reference 224 in temporal prediction can include coded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.
[0070] Referring to process 200B, in the forward pass, the encoder performs prediction operations in a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra prediction. For an original BPU of a picture being encoded, the prediction reference 224 may include one or more neighboring BPUs of the same picture that were coded (in the forward pass) and reconstructed (in the reconstruction pass). The encoder may generate the predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, the encoder may perform extrapolation at the pixel level, such as by extrapolating the value of a corresponding pixel for each pixel of the predicted BPU 208. The neighboring BPUs used for extrapolation can be positioned relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., bottom left, bottom right, top left, or top right of the original BPU), or from any direction defined by the video coding standard being used. For intra prediction, the prediction data 206 may include, for example, the positions (e.g., coordinates) of the neighboring BPUs used, the sizes of the neighboring BPUs used, parameters of the extrapolation, the orientation of the neighboring BPUs used relative to the original BPU, etc.
[0071] Also for example, in the temporal prediction stage 2044, the encoder may perform inter-prediction. For an original BPU of a current picture, the prediction reference 224 may include one or more pictures (called "reference pictures") that have been coded (in the forward pass) and reconstructed (in the reconstruction pass). In some embodiments, the reference pictures may be coded and reconstructed for each BPU. For example, the encoder may generate a reconstructed BPU by adding the reconstructed residual BPU 222 to the predicted BPU 208. Once all the reconstructed BPUs of the same picture have been generated, the encoder may generate the reconstructed image as the reference picture. The encoder may perform a "motion estimation" operation to search for a matching region within a certain range (called a "search window") of the reference picture. The position of the search window in the reference picture may be determined based on the position of the original BPU in the current picture. For example, the search window may be centered at a location in the reference picture that has the same coordinates as the original BPU in the current picture and may extend to a predetermined distance. If the encoder identifies a region in the search window that is similar to the original BPU (e.g., by a pel recursion algorithm, a block matching algorithm, etc.), the encoder can determine such a region as a matching region. The matching region may have different dimensions than the original BPU (e.g., smaller than the original BPU, equal to the original BPU, larger than the original BPU, or a different shape). Because the reference picture and the current picture are temporally separated on the timeline, the matching region can be considered to "move" to the position of the original BPU over time. The encoder can record the direction and distance of the motion as a "motion vector." If multiple reference pictures are used, the encoder can find a matching region and determine its associated motion vector for each reference picture. In some embodiments, the encoder can assign weights to pixel values of the matching region in each matching reference picture.
[0072] Motion estimation can be used to identify various types of motion, such as, for example, translation, rotation, zoom, etc. For inter prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, a motion vector associated with the matching region, the number of reference pictures, weights associated with the reference pictures, etc.
[0073] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., motion vectors) and the prediction reference 224. For example, the encoder may shift the matching region of the reference picture by the motion vector, within which the encoder may predict the original BPU of the current picture. If multiple reference pictures are used, the encoder may shift the matching region of the reference picture by the average pixel value of the matching region and each motion vector. In some embodiments, if the encoder has assigned weights to the pixel values of the matching region of each matching reference picture, it may add a weighted sum of the pixel values of the shifted matching region.
[0074] In some embodiments, inter prediction may be unidirectional or bidirectional. Unidirectional inter prediction can use one or more reference pictures in the same temporal direction for the current picture. Unidirectional inter prediction uses a reference picture that precedes the current picture. Bidirectional inter prediction can use one or more reference pictures in both temporal directions for the current picture.
[0075] Still referring to the forward pass of process 200B, after spatial prediction 2042 and temporal prediction stage 2044, in mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra-prediction or inter-prediction) for the current iteration of process 200B. For example, the encoder can perform a rate-distortion optimization technique, in which the encoder can select a prediction mode to minimize the value of a cost function that depends on the distortion of the reconstructed reference picture in the candidate prediction mode and the bitrate of the candidate prediction mode. Depending on the selected prediction mode, the encoder can generate a corresponding predicted BPU 208 and predicted data 206.
[0076] In the reconstruction path of process 200B, if an intra prediction mode is selected in the forward path, after generating the prediction reference 224 (e.g., the coded and reconstructed current BPU in the current picture), the encoder can directly provide the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If an intra prediction mode is selected in the forward path, after generating the prediction reference 224 (e.g., the coded and reconstructed current picture for all BPUs), the encoder can provide the prediction reference 224 to the loop filter stage 232, where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate distortions (e.g., blocking artifacts) introduced by inter prediction. The encoder can apply various loop filter techniques to the loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filtering, etc. The loop-filtered reference picture may be stored in a buffer 234 (or "decoded picture buffer") for later use (e.g., for use as an inter-prediction reference picture for a future picture in the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) along with the quantized transform coefficients 216, the prediction data 206, and other information in the binary coding stage 226.
[0077] FIG. 3A shows a schematic diagram of an exemplary decoding process 300A according to an embodiment of the present disclosure. Process 300A may be a decompression process corresponding to compression process 200A in FIG. 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder (e.g., image / video decoder 144 in FIG. 1) can decode video bitstream 228 into video stream 304 based on process 300A. Video stream 304 may be very similar to video sequence 202. However, due to information loss in the compression and decompression processes (e.g., quantization stage 214 in FIGS. 2A-2B), video stream 304 is generally not identical to video sequence 202. Similar to processes 200A and 200B in FIGS. 2A-2B, the decoder can perform process 300A at the basic processing unit (BPU) level for each coded picture in video bitstream 228. For example, an encoder may perform process 300A iteratively, in which the encoder decodes a basic processing unit in one iteration of process 300A. In some embodiments, a decoder may perform process 300A in parallel for a region of each coded picture in video bitstream 228.
[0078] In FIG. 3A , a decoder may provide a portion of a video bitstream 228 associated with a basic processing unit (referred to as a “coding BPU”) of a coded picture to a binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may generate a reconstructed residual BPU 222 by providing the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220. The decoder may generate a prediction BPU 208 by providing the prediction data 206 to a prediction stage 204. The decoder may generate a prediction reference 224 by adding the reconstructed residual BPU 222 to the prediction BPU 208. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded picture buffer (DPB) in computer memory). The decoder may provide the prediction reference 224 to the prediction stage 204 for performing a prediction operation in a next iteration of the process 300A.
[0079] The decoder may iteratively perform process 300A to decode each of the coded BPUs of the coded picture and generate predictive references 224 for encoding the next coded BPU of the coded picture. After decoding all of the coded BPUs of the coded picture, the decoder may output the picture to the video stream 304 for display and proceed to decode the next coded picture in the video bitstream 228.
[0080] In binary decoding stage 302, the decoder may perform the inverse of the binary coding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or other lossless compression algorithm). In some embodiments, the decoder may decode other information in binary decoding stage 302, in addition to prediction data 206 and quantized transform coefficients 216, such as, for example, a prediction mode, parameters of the prediction operation, a transform type, parameters of the quantization process (e.g., quantization parameters), and encoder control parameters (e.g., bitrate control parameters). In some embodiments, if video bitstream 228 is transmitted in packets over a network, the decoder may depacketize video bitstream 228 before providing video bitstream 228 to binary decoding stage 302.
[0081] 3B shows a schematic diagram of another exemplary decoding process 300B according to an embodiment of the present disclosure. Process 300B can be modified from process 300A. For example, process 300B can be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 300A, process 300B further divides prediction stage 204 into spatial prediction stage 2042 and temporal prediction stage 2044, and further includes loop filter stage 232 and buffer 234.
[0082] In process 300B, for a coding basic processing unit (referred to as a "current BPU") of a coding picture being decoded (referred to as a "current picture"), prediction data 206 decoded by the decoder from binary decoding stage 302 may include various data depending on the prediction mode used by the encoder to encode the current BPU. For example, if intra prediction is used by the encoder to encode the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, etc. The parameters of the intra prediction operation may include, for example, the positions (e.g., coordinates) of one or more reference neighboring BPUs, the size of the neighboring BPUs, extrapolation parameters, and the direction of the neighboring BPUs relative to the original BPU. Also, for example, if inter prediction is used by the encoder to encode the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, etc. Parameters for the inter-prediction operation may include, for example, the number of reference pictures currently associated with the BPU, weights associated with each of the reference pictures, the locations (e.g., coordinates) of one or more matching regions in each reference picture, one or more motion vectors associated with each of the matching regions, etc.
[0083] Based on the prediction mode indicator, the decoder may determine whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. Details of performing such spatial or temporal prediction have been described in FIG. 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder may generate a prediction BPU 208. The decoder may generate a prediction reference 224 by adding the prediction BPU 208 and a reconstructed residual BPU 222, as shown in FIG. 3A.
[0084] In process 300B, the decoder may supply the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing a prediction operation in the next iteration of process 300B. For example, when the current BPU is decoded by intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder may supply the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). When the current BPU is decoded by inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture decoded by all BPUs), the encoder may reduce or remove distortion (e.g., blocking artifacts) by supplying the prediction reference 224 to the loop filter stage 232. The decoder may apply a loop filter to the prediction reference 224 as described in FIG. 2B. The loop-filtered reference picture may be stored in a buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., for use as an inter-predicted reference picture for future coded pictures of the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, if the prediction mode indicator of the prediction data 206 indicates that inter-prediction was used to encode the current BPU, the prediction data may further include parameters of the loop filter (e.g., loop filter strength).
[0085] Referring back to FIG. 1 , each of the image / video preprocessor 122, the image / video encoder 124, and the image / video decoder 144 may be implemented as any suitable hardware, software, or combination thereof. FIG. 4 is a block diagram of an exemplary apparatus 400 for processing image data according to an embodiment of the present disclosure. For example, the apparatus 400 may be a preprocessor, an encoder, or a decoder. As shown in FIG. 4 , the apparatus 400 may include a processor 402. When the processor 402 executes the instructions described herein, the apparatus 400 can become a dedicated machine for preprocessing, encoding, or decoding image data. The processor 402 may be any type of circuitry capable of manipulating or processing information. For example, processor 402 may include any combination of any number of central processing units (“CPUs”), graphics processing units (“GPUs”), neural processing units (“NPUs”), microcontroller units (“MCUs”), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, IP cores, programmable logic arrays (PLAs), programmable array logic (PALs), generic array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems on chips (SoCs), or application specific integrated circuits (ASICs), etc. In some embodiments, processor 402 may be a set of processors grouped as a single logical component. For example, as shown in FIG. 4, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0086] The apparatus 400 may include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, etc.). For example, as shown in FIG. 4, the stored data may include program instructions (e.g., program instructions for implementing stages of processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 may access the program instructions and data for processing (e.g., via bus 410) and perform operations or manipulations on the data for processing by executing the program instructions. The memory 404 may include a high-speed random access storage device or a non-volatile storage device. In some embodiments, the memory 404 may include any combination of any number of random access memory (RAM), read-only memory (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, compact flash cards, etc. The memory 404 may be a group of memories grouped as a single logical component (not shown in FIG. 4).
[0087] Bus 410 may be a communication device that transfers data between components within apparatus 400, such as an internal bus (e.g., a CPU memory bus) or an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Express port).
[0088] For ease of explanation and to avoid ambiguity, the processor 402 and other data processing circuitry will be collectively referred to in this disclosure as "data processing circuitry." The data processing circuitry may be implemented entirely as hardware or as a combination of software, hardware, or firmware. Additionally, the data processing circuitry may be a single, independent module or may be combined in whole or in part with other components of the device 400.
[0089] Device 400 may further include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.) In some embodiments, network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communications ("NFC") adapters, or cellular network chips, etc.
[0090] In some embodiments, apparatus 400 may further comprise a peripheral interface 408 to provide connection with one or more peripheral devices. As shown in Figure 4, the peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, touchpad, or touchscreen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera, or an input interface coupled to a video archive), etc.
[0091] It should be noted that a video codec (e.g., a codec performing process 200A, 200B, 300A, or 300B) may be implemented as any combination of software or hardware modules in device 400. For example, some or all of the stages of process 200A, 200B, 300A, or 300B may be implemented as one or more software modules in device 400, such as program instructions loadable into memory 404. Also, for example, some or all of the stages of process 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules in device 400, such as dedicated data processing circuitry (e.g., FPGA, ASIC, NPU, etc.).
[0092] Supplemental Enhancement Information (SEI) messages are intended to be conveyed within coded video bitstreams by methods specified in video coding specifications, or by other means determined by the specifications of systems utilizing such coded video bitstreams. SEI messages may indicate the timing of video pictures or contain various types of data that describe various characteristics of the coded video or how it can be used or enhanced. SEI messages that can contain arbitrary user-defined data are also defined. SEI messages do not affect the core decoding process, but may provide recommendations on how to post-process or display the video.
[0093] With the emergence of deep generative models, including variational autoencoding (VAE) and generative adversarial networks (GAN), facial video compression has achieved promising performance improvements. For example, X2Face can be used to control face generation via image, audio, and pose codes. Furthermore, realistic neural talking head models can be used via multi-shot adversarial learning. For video-to-video synthesis tasks, Face-vid2vid can be used. Schemes that leverage compact 3D keypoint representations to drive generative models for rendering target frames can also be used. Furthermore, a mobile-enabled video chat system based on FOMM can be used. VSBNet, which uses adversarial learning to reconstruct original frames from landmarks, can also be used. Furthermore, an end-to-end talking head video compression framework based on compact feature learning (CFTE), designed for highly efficient talking face video compression for ultra-low bandwidth scenes, can be used. The CFTE scheme leverages compact feature representations to compensate for temporal evolution and reconstruct target face video frames in an end-to-end manner. Furthermore, the CFTE scheme can be incorporated into video coding frameworks for rate-distortion supervision. Although these algorithms achieve frame reconstruction with a small number of facial parameters due to the powerful rendering capabilities of deep generative models, some head pose movements and facial expression movements are still not rendered accurately compared to the original video.
[0094] FIG. 5 is a schematic diagram illustrating an exemplary deep learning-based video generation and compression framework 500 according to some embodiments of the present disclosure. The framework 500 is suitable for compressing and generating talking face videos. For example, the framework 500 can be based on a first-order motion model (FOMM). The FOMM deforms a reference source frame to follow the motion of the driving video. This method works for various types of video (e.g., motion pictures, cartoons), but can also be used for facial animation applications. The FOMM follows an encoder-decoder architecture with a motion transfer component that includes the following steps:
[0095] First, the keypoint extractor (also called the motion module) is trained using equivariant loss without explicit labels. This keypoint extractor calculates two sets of 10 trained keypoints for the source frame and the driving frame. The trained keypoints are transformed from a channel size × 64 × 64 feature map through a Gaussian map function, so that each corresponding keypoint can represent different channel feature information. Note that every keypoint is an (x,y) point that can represent the most important information in the feature map.
[0096] A dense motion network then uses the landmarks and source frames to generate a dense motion field and occlusion map.
[0097] Then, the encoder 510 encodes the source frames using a conventional image / video compression scheme such as HEVC / VVC or JPEG / BPG, where VVC is used to compress the source frames.
[0098] In a later stage, the resulting feature map is warped by a dense motion field (by a differentiable grid sample operation) and multiplied with the occlusion map.
[0099] Finally, the decoder 520 generates the image from the warped map.
[0100] Figure 6 is a schematic diagram illustrating an example encoder-decoder coding framework 600 with a compact feature size of 1x4x4 for talking face video, according to some embodiments of the present disclosure. Figure 6 provides another basic framework for a deep-based video generative compression scheme based on compact feature representation, i.e., CFTE, which follows an encoder-decoder architecture that applies a context-based coding scheme.
[0101] On the encoder 610 side, the compression framework includes three modules: an encoder (also called a VVC encoding module) for compressing key frames, a feature extractor for extracting compact human features for other inter frames, and a feature coding module for compressing inter-predicted residuals of the compact human features. First, a key frame representing human texture is compressed by the VVC encoder. Through the compact feature extractor, each subsequent inter frame is represented by a compact feature matrix with a size of 1x4x4. Note that the size of the compact feature matrix is not fixed, and the number of feature parameters can be increased or decreased depending on the specific requirements for bit consumption. These extracted features are then inter-predicted and quantized, and the residuals are finally entropy coded as the final bitstream.
[0102] On the decoder 620 side, this compression framework also includes three main modules: decoding to reconstruct keyframes, reconstructing compact features through entropy decoding and compensation, and generating final videos using the reconstructed features and decoded keyframes. More specifically, during final video generation, compact feature extraction can represent decoded keyframes from a VVC bitstream as features. Next, considering features from keyframes and interframes, associated sparse motion fields are calculated to facilitate the generation of pixel-wise dense motion maps and occlusion maps. Finally, based on a deep generative model, the decoded keyframes, pixel-wise dense motion fields, and occlusion maps with implicit motion field characterization are used to generate final videos with accurate appearance, pose, and expression.
[0103] To further improve coding performance, numerous studies have focused on 3D faces. They employ a 3D head model to encode only the pose parameters for the task of face-specific video compression. Subsequently, both eigenspace and principal component analysis (PCA) models have been used for this task. However, based on these traditional 3D techniques, the visual quality of the reconstructed images is unacceptable. With the development of deep generative models, this 3D MMM-assisted face video generation task can yield promising results.
[0104] 7 is a schematic diagram illustrating a general encoder-decoder generation compression framework 700 for 3DMM-assisted talking face video, according to some embodiments of the present disclosure. In general, 3DMM-assisted face video generation can provide accurate 3D face reconstruction based on a combination of a shape S and a texture T, which are given as follows:
number
number
number
number
[0105] The existing SEI messages used in the current VVC standard are not designed to handle the task of facial video compression. However, facial videos can be described by the variation of feature structures with strong prior probabilities, such as landmarks and keypoints, and can be further parameterized into a set of facial semantic information representing head pose and facial expression states. Such compact facial semantics can provide a large space for implementing the syntax design and semantic description of SEI messages for facial video compression.
[0106] Furthermore, facial video communication is seeking more general use cases beyond facial video reconstruction, such as facial video retargeting and animation. For example, with the proliferation of metaverse activities, real-world facial movements may need to be transferred to the virtual metaverse world and represented by others. Furthermore, facial video reconstruction is expected to more closely match real-world situations and provide corresponding retargeting. Therefore, it is necessary to define an SEI message that can include various types of data indicating the timing of video pictures and / or describing various characteristics of the coded video or how it can be used or enhanced. In this way, post-processing or display of the reconstructed facial video can meet the user's actual needs in a user-friendly manner.
[0107] To address the above challenges, this disclosure proposes a new SEI message called the Face Video Generation and Compression SEI message. The proposed SEI has at least two functions: (1) reconstruct high-quality talking face videos at ultra-low bitrates; and (2) manipulate talking face videos for personalized characterization. Thus, the proposed SEI is applicable to video conferencing, live entertainment, and metaverse-related activities.
[0108] 8 is a flowchart of an example method 800 for processing video based on a facial video generation compression supplemental enhancement information (SEI) message according to some embodiments of the present disclosure. The method 800 describes a general syntax structure and syntax element order of a facial video generation compression SEI message. With reference to FIG. 8, the method 800 may include the following steps 802-808.
[0109] In step 802, an identification number is signaled to indicate whether a facial image generation compression scheme is to be used.
[0110] In step 804, if the identification number indicates that a facial image generation compression scheme is to be used, then multiple (e.g., five) facial present flags are further signaled to indicate the presence and length of specific syntax elements associated with the facial image generation compression scheme. In some embodiments, the facial present flags may include flags indicating head position, head rotation, head translation, eye blink, mouth motion, etc.
[0111] In step 806, one or more corresponding face information parameters are signaled. In some embodiments, the corresponding face information parameters may include head information parameters (e.g., head position parameters, head rotation parameters, head translation parameters, etc.), eye information parameters (e.g., blink parameters, etc.), and mouth information parameters (e.g., mouth motion parameters, etc.). In some embodiments, the head position corresponds to the face information included in the face image. The head rotation may further include three degrees of freedom, such as roll, pitch, or yaw. The head translation may further include, for example, translation in three directions, such as translation in a 3D coordinate system. The mouth motion parameters may further include multiple parameters indicating mouth motion information.
[0112] In step 808, the relevant facial parts are reconstructed or retargeted based on the base picture according to the corresponding signaled facial information parameters.
[0113] Some exemplary embodiments of the proposed facial image generation compressed SEI message are described in detail below.
[0114] In some embodiments, multiple facial present flags indicating the presence and length of specific syntax elements associated with the facial video generation compression scheme and parameters indicating quantization factors for processing facial semantic parameters are signaled for every face frame.
[0115] FIG. 9 is a flowchart of an example method 900 for processing video based on a facial video generation compression supplemental enhancement information (SEI) message according to some embodiments of the present disclosure. FIG. 10 illustrates an example syntax of the disclosed facial video generation compression SEI message according to some embodiments of the present disclosure. Method 900 describes the general syntax structure and syntax element order of the facial video generation compression SEI message. Method 900 may be performed by an encoder (e.g., by process 200A of FIG. 2A or 200B of FIG. 2B) or by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) may perform method 900. In some embodiments, method 900 may be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions, such as program code, executed by a computer (e.g., apparatus 400 of FIG. 4). Referring to FIGS. 9 and 10, method 900 may include the following steps 902-910.
[0116] In step 902, an identification number (eg, fi_id) is signaled to indicate whether a facial image generation compression scheme is used, see 1001 shown in FIG.
[0117] In step 904, a parameter indicating the number of face frames, for example, fi_num_set_of_parameter, is signaled, see 1002 in FIG.
[0118] In step 906, parameters indicating quantization factors for processing facial semantic parameters are signaled for all face frames, with reference to 1003 shown in Figure 10. The facial semantic parameters may include 14 facial semantic parameters (i.e., fi_location[i], fi_rotation_roll[i], fi_rotation_pitch[i], fi_rotation_yaw[i], fi_translation_x[i], fi_translation_y[i], fi_translation_z[i], fi_eye[i], fi_mouth_para1[i], fi_mouth_para2[i], fi_mouth_para3[i], fi_mouth_para4[i], fi_mouth_para5[i], and fi_mouth_para6[i]).
[0119] In step 908, and referring to 1004 shown in FIG. 10, for every face frame, a number of facial present flags are further signaled to indicate the presence and length of specific syntax elements associated with the facial video generation compression scheme.
[0120] In step 910, for each face frame, the corresponding face information parameters are signaled.
[0121] The process according to Figures 9 and 10 is explained as follows:
[0122] First, the facial image generation compression includes an identification number that can be used to identify the facial image generation compression filter.
[0123] Second, every SEI message always has a base picture (i.e., the first face frame in the face sequence) included in the PU, which can provide a texture reference so that the face frame can be reconstructed using the face information parameters carried in the SEI message.
[0124] Third, for these face information parameters in the SEI message, five corresponding face information parameter present flags (i.e., fi_head_location_present_flag, fi_head_rotation_present_flag, fi_head_translation_present_flag, fi_eye_blinking_present_flag, and fi_mouth_motion_present_flag in sequence 400) are signaled to determine whether the related face information parameters are transmitted. If these parameters are not transmitted, the corresponding face information parameters from the base picture are copied to generate the subsequent face frame.
[0125] Fourth, if fi_head_location_present_flag is present, the head location parameter fi_location[i] shall be carried in the SEI message. If fi_head_rotation_present_flag is present, these head rotation parameters (fi_rotation_roll[i], fi_rotation_pitch[i], and fi_rotation_yaw[i]) shall be carried in the SEI message. If fi_head_translation_present_flag is present, these head translation parameters (fi_translation_x[i], fi_translation_y[i], and fi_translation_z[i]) shall be carried in the SEI message. If fi_eye_blinking_present_flag is present, the blink parameter fi_eye[i] shall be carried in the SEI message. If fi_mouth_motion_present_flag is present, these mouth motion parameters (fi_mouth_para1[i], fi_mouth_para2[i], fi_mouth_para3[i], fi_mouth_para4[i], fi_mouth_para5[i], and fi_mouth_para6[i]) shall be carried in the SEI message. If the relevant flag is not present, the corresponding facial information parameters from the base picture are copied to generate subsequent face pictures, i.e., the 14 facial parameters of these generated face pictures are kept the same as those of the base picture.
[0126] Fifth, if these corresponding facial information parameters are carried in this SEI message, the face video can be reconstructed towards personalized characterization or a user-friendly way through the powerful generative capabilities of generative adversarial networks.
[0127] The semantics associated with the syntax in Figure 10 are explained as follows:
[0128] fi_id contains an identification number that can be used to identify the face image generation compression filter. The value of fi_id ranges from 0 to 2. 32 It can be in the range of ∼2 (inclusive).
[0129] fi_num_set_of_parameter indicates the number of face frames that can be used to realize face image generation compression using the face image generation compression filter. The value of fi_num_set_of_parameter is between 0 and 2. 10 (inclusive), beyond which other Face frames may be packaged into the next SEI message.
[0130] fi_quantization_factor is the quantization factor for processing these 14 facial semantic parameters (i.e., fi_location[i], fi_rotation_roll[i], fi_rotation_pitch[i], fi_rotation_yaw[i], fi_translation_x[i], fi_translation_y[i], fi_translation_z[i], fi_eye[i], fi_mouth_para1[i], fi_mouth_para2[i], fi_mouth_para3[i], fi_mouth_para4[i], fi_mouth_para5[i], and fi_mouth_para6[i]) in float16 type. These float16 parameters can be further expanded via fi_quantization_factor, where the value of fi_quantization_factor ranges from 0 to 10. 16 For example, the value of the original facial parameter is 0.1234567891234567, and the value of fi_quantization_factor is 10 6 , so the corresponding quantized facial parameter is 123456.
[0131] If present, fi_head_location_present_flag is equal to 1; if not present, fi_head_location_present_flag is equal to 0.
[0132] fi_location[i] specifies the quantized residual parameter corresponding to the head location between the ith face frame and the (i-1)th face frame via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_location[0] specifies the quantized head location parameter from the 0th face frame (base picture).
[0133] If present, fi_head_rotation_present_flag is equal to 1; if not present, fi_head_rotation_present_flag is equal to 0.
[0134] fi_rotation_roll[i] specifies the quantized residual parameter corresponding to the head rotation (called roll) around the front-to-back axis between the ith and (i-1)th face frames via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_rotation_roll[0] specifies the quantized front-to-back axis head rotation parameter from the 0th face frame (base picture).
[0135] fi_rotation_pitch[i] specifies the quantized residual parameter corresponding to the head rotation (called pitch) around the side-to-side axis between the ith and (i-1)th face frames via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_rotation_pitch[0] specifies the quantized side-to-side axis head rotation parameter from the 0th face frame (base picture).
[0136] fi_rotation_yaw[i] specifies the quantized residual parameter corresponding to the head rotation around the vertical axis (called yaw) between the ith and (i-1)th face frames via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_rotation_yaw[0] specifies the quantized vertical axis head rotation parameter from the 0th face frame (base picture).
[0137] If present, fi_head_translation_present_flag is equal to 1; if not present, fi_head_translation_present_flag is equal to 0.
[0138] fi_translation_x[i] specifies the quantized residual parameter corresponding to the head translation about the x-axis between the ith and (i-1)th face frames via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_translation_x[0] specifies the quantized x-axis head translation parameter from the 0th face frame (base picture).
[0139] fi_translation_y[i] specifies the quantized residual parameter corresponding to the head translation about the y-axis between the ith and (i-1)th face frames via fi_quantization_factor, if i is not equal to 0. If i is equal to 0, fi_translation_y[0] specifies the quantized y-axis head translation parameter from the 0th face frame (base picture).
[0140] fi_translation_z[i] specifies the quantized residual parameter corresponding to the head translation about the z-axis between the ith and (i-1)th face frames via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_translation_z[0] specifies the quantized z-axis head translation parameter from the 0th face frame (base picture).
[0141] If present, fi_eye_blinking_present_flag is equal to 1; if not present, fi_eye_blinking_present_flag is equal to 0.
[0142] fi_eye[i] specifies the quantized residual parameter corresponding to the degree of blinking between the ith and (i-1)th face frames via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_eye[0] specifies the quantized blinking parameter from the 0th face frame (base picture).
[0143] If present, fi_mouth_motion_present_flag is equal to 1; if not present, fi_mouth_motion_present_flag is equal to 0.
[0144] fi_mouth_para1[i] specifies the quantized residual parameters corresponding to the mouth motion between the ith face frame and the (i-1)th face frame via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_mouth_para1[0] specifies the quantized mouth motion parameters from the 0th face frame (base picture).
[0145] fi_mouth_para2[i] specifies the quantized residual parameters corresponding to the mouth motion between the ith face frame and the (i-1)th face frame via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_mouth_para2[0] specifies the quantized mouth motion parameters from the 0th face frame (base picture).
[0146] fi_mouth_para3[i] specifies the quantized residual parameters corresponding to the mouth motion between the ith face frame and the (i-1)th face frame via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_mouth_para3[0] specifies the quantized mouth motion parameters from the 0th face frame (base picture).
[0147] fi_mouth_para4[i] specifies the quantized residual parameters corresponding to the mouth motion between the ith face frame and the (i-1)th face frame via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_mouth_para4[0] specifies the quantized mouth motion parameters from the 0th face frame (base picture).
[0148] fi_mouth_para5[i] specifies the quantized residual parameters corresponding to the mouth motion between the ith face frame and the (i-1)th face frame via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_mouth_para5[0] specifies the quantized mouth motion parameters from the 0th face frame (base picture).
[0149] fi_mouth_para6[i] specifies the quantized residual parameters corresponding to the mouth motion between the ith face frame and the (i-1)th face frame via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_mouth_para6[0] specifies the quantized mouth motion parameters from the 0th face frame (base picture).
[0150] In some embodiments, multiple facial present flags are signaled for each face frame, each flag indicating the presence and length of a particular syntax element associated with a facial video generation compression scheme.
[0151] Furthermore, although the syntax definitions above and below specify what happens when the value is either 0 or 1, these values are configurable and can be changed, for example, so that fi_head_location_present_flag is equal to 1 when not present and fi_head_location_present_flag is equal to 0 when present.
[0152] FIG. 11 is a flowchart of an example method 1100 for processing video based on a facial video generation compression supplemental enhancement information (SEI) message according to some embodiments of the present disclosure. FIG. 12 illustrates another example syntax of the disclosed facial video generation compression SEI message according to some embodiments of the present disclosure. Method 1100 describes the general syntax structure and syntax element order of the facial video generation compression SEI message. Method 1100 may be performed by an encoder (e.g., by process 200A of FIG. 2A or 200B of FIG. 2B) or by one or more software or hardware components of an apparatus (e.g., apparatus 400 in FIG. 4). For example, a processor (e.g., processor 402 in FIG. 4) may perform method 1100. In some embodiments, method 1100 may be implemented by a computer program product embodied in a computer-readable medium including computer-executable instructions, such as program code, executed by a computer (e.g., apparatus 400 in FIG. 4). 11 and 12, the method 1100 may include the following steps 1102 to 1110.
[0153] In step 1102, an identification number (eg, fi_id) is signaled to indicate whether a facial image generation compression scheme is used, see 1201 shown in FIG.
[0154] In step 1104, a parameter indicating the number of face frames, for example, fi_num_set_of_parameter, is signaled, as shown in FIG. 12 at 1202.
[0155] In step 1106, parameters indicating quantization factors for processing facial semantic parameters are signaled for every face frame, with reference to 1203 shown in Figure 12. The facial information parameters may include 14 facial information parameters (i.e., fi_location[i], fi_rotation_roll[i], fi_rotation_pitch[i], fi_rotation_yaw[i], fi_translation_x[i], fi_translation_y[i], fi_translation_z[i], fi_eye[i], fi_mouth_para1[i], fi_mouth_para2[i], fi_mouth_para3[i], fi_mouth_para4[i], fi_mouth_para5[i], and fi_mouth_para6[i]).
[0156] In step 1108, referring to 1204 shown in FIG. 12, for each face frame, a number of facial present flags are further signaled to indicate the presence and length of specific syntax elements associated with the facial video generation compression scheme.
[0157] In step 1110, for each face frame, the corresponding face information parameters are signaled.
[0158] The process according to Figures 11 and 12 is explained as follows.
[0159] First, the facial image generation compression includes an identification number that can be used to identify the facial image generation compression filter.
[0160] Second, every SEI message always has a base picture (i.e., the first face frame in the face sequence) included in the PU, which can provide a texture reference so that the face frame can be reconstructed using the face information parameters carried in the SEI message.
[0161] Third, for these facial information parameters in the SEI message, it is proposed to set five corresponding facial information present flags (i.e., fi_head_location_present_flag[i], fi_head_rotation_present_flag[i], fi_head_translation_present_flag[i], fi_eye_blinking_present_flag[i], fi_mouth_motion_present_flag[i]) to determine whether the related facial information parameters of each face picture[i] are transmitted or not.
[0162] Fourth, if fi_head_location_flag[i] is present, the head location parameter fi_location[i] may be carried in the SEI message. If fi_head_rotation_present_flag[i] is present, these head rotation parameters (fi_rotation_roll[i], fi_rotation_pitch[i], and fi_rotation_yaw[i]) shall be carried in the SEI message. If fi_head_translation_present_flag[i] is present, these head translation parameters (fi_translation_x[i], fi_translation_y[i], and fi_translation_z[i]) shall be carried in the SEI message. If fi_eye_blinking_flag[i] is present, the blink parameter fi_eye[i] shall be carried in the SEI message. If fi_mouth_motion_present_flag[i] is present, these mouth motion parameters (fi_mouth_para1[i], fi_mouth_para2[i], fi_mouth_para3[i], fi_mouth_para4[i], fi_mouth_para5[i], and fi_mouth_para6[i]) shall be carried in the SEI message. If any of these five facial information present flags is not present, there are two situations according to this disclosure.
[0163] In some embodiments, the facial information parameters of the current face frame (i.e., face frame[i]) are copied from the corresponding information from the previous face frame (i.e., face frame[i-1]). For the first face frame ([i=0]), the facial information parameters are copied from the base picture.
[0164] In some embodiments, the face information parameters of the current face frame (ie, face frame[i]) are copied from the base picture.
[0165] Fifth, if these corresponding facial information parameters are carried in this SEI message, the face video can be reconstructed towards personalized characterization or a user-friendly way through the powerful generative capabilities of generative adversarial networks.
[0166] The semantics associated with the syntax in Figure 12 are as follows:
[0167] fi_id contains an identification number that can be used to identify the face image generation compression filter. The value of fi_id ranges from 0 to 2. 32 It is assumed to be in the range of ~2 (inclusive).
[0168] fi_num_set_of_parameter indicates the number of face frames that can be used to realize face image generation compression using the face image generation compression filter. The value of fi_num_set_of_parameter is between 0 and 2. 10 (inclusive), beyond which other Face frames must be packaged into the next SEI message.
[0169] fi_quantization_factor is the quantization factor for processing these 14 facial semantic parameters (i.e., fi_location[i], fi_rotation_roll[i], fi_rotation_pitch[i], fi_rotation_yaw[i], fi_translation_x[i], fi_translation_y[i], fi_translation_z[i], fi_eye[i], fi_mouth_para1[i], fi_mouth_para2[i], fi_mouth_para3[i], fi_mouth_para4[i], fi_mouth_para5[i], and fi_mouth_para6[i]) in float16 type. These float16 parameters shall be further expanded via fi_quantization_factor, where the value of fi_quantization_factor shall be between 0 and 10. 16 For example, the original facial parameter value is 0.1234567891234567, and the fi_quantization_factor value is 10 6 , so the corresponding quantized facial parameter is 123456.
[0170] If present, fi_head_location_present_flag[i] is equal to 1; if not present, fi_head_location_present_flag is equal to 0.
[0171] fi_location[i] specifies the quantized residual parameter corresponding to the head location between the ith face frame and the (i-1)th face frame via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_location[0] specifies the quantized head location parameter from the 0th face frame (base picture).
[0172] If present, fi_head_rotation_present_flag[i] is equal to 1; if not present, fi_head_rotation_present_flag is equal to 0.
[0173] fi_rotation_roll[i] specifies the quantized residual parameter corresponding to the head rotation (called roll) around the front-to-back axis between the ith and (i-1)th face frames via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_rotation_roll[0] specifies the quantized front-to-back axis head rotation parameter from the 0th face frame (base picture).
[0174] fi_rotation_pitch[i] specifies the quantized residual parameter corresponding to the head rotation (called pitch) around the side-to-side axis between the ith and (i-1)th face frames via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_rotation_pitch[0] specifies the quantized side-to-side axis head rotation parameter from the 0th face frame (base picture).
[0175] fi_rotation_yaw[i] specifies the quantized residual parameter corresponding to the head rotation around the vertical axis (called yaw) between the ith and (i-1)th face frames via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_rotation_yaw[0] specifies the quantized vertical axis head rotation parameter from the 0th face frame (base picture).
[0176] If present, fi_head_translation_present_flag[i] is equal to 1; if not present, fi_head_translation_present_flag is equal to 0.
[0177] fi_translation_x[i] specifies the quantized residual parameter corresponding to the head translation about the x-axis between the ith and (i-1)th face frames via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_translation_x[0] specifies the quantized x-axis head translation parameter from the 0th face frame (base picture).
[0178] fi_translation_y[i] specifies the quantized residual parameter corresponding to the head translation about the y-axis between the ith and (i-1)th face frames via fi_quantization_factor, if i is not equal to 0. If i is equal to 0, fi_translation_y[0] specifies the quantized y-axis head translation parameter from the 0th face frame (base picture).
[0179] fi_translation_z[i] specifies the quantized residual parameter corresponding to the head translation about the z-axis between the ith and (i-1)th face frames via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_translation_z[0] specifies the quantized z-axis head translation parameter from the 0th face frame (base picture).
[0180] If present, fi_eye_blinking_present_flag[i] is equal to 1; if not present, fi_eye_blinking_present_flag is equal to 0.
[0181] fi_eye[i] specifies the quantized residual parameter corresponding to the degree of blinking between the ith and (i-1)th face frames via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_eye[0] specifies the quantized blinking parameter from the 0th face frame (base picture).
[0182] If present, fi_mouth_motion_present_flag[i] is equal to 1; if not present, fi_mouth_motion_present_flag is equal to 0.
[0183] fi_mouth_para1[i] specifies the quantized residual parameters corresponding to the mouth motion between the ith face frame and the (i-1)th face frame via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_mouth_para1[0] specifies the quantized mouth motion parameters from the 0th face frame (base picture).
[0184] fi_mouth_para2[i] specifies the quantized residual parameters corresponding to the mouth motion between the ith face frame and the (i-1)th face frame via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_mouth_para2[0] specifies the quantized mouth motion parameters from the 0th face frame (base picture).
[0185] fi_mouth_para3[i] specifies the quantized residual parameters corresponding to the mouth motion between the ith face frame and the (i-1)th face frame via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_mouth_para3[0] specifies the quantized mouth motion parameters from the 0th face frame (base picture).
[0186] fi_mouth_para4[i] specifies the quantized residual parameters corresponding to the mouth motion between the ith face frame and the (i-1)th face frame via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_mouth_para4[0] specifies the quantized mouth motion parameters from the 0th face frame (base picture).
[0187] fi_mouth_para5[i] specifies the quantized residual parameters corresponding to the mouth motion between the ith face frame and the (i-1)th face frame via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_mouth_para5[0] specifies the quantized mouth motion parameters from the 0th face frame (base picture).
[0188] fi_mouth_para6[i] specifies the quantized residual parameters corresponding to the mouth motion between the ith face frame and the (i-1)th face frame via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_mouth_para6[0] specifies the quantized mouth motion parameters from the 0th face frame (base picture).
[0189] In some embodiments, multiple facial present flags indicating the presence and length of specific syntax elements associated with the facial video generation compression scheme and parameters indicating quantization factors for processing facial semantic parameters are signaled for each face frame, respectively.
[0190] FIG. 13 is a flowchart of an example method 1300 for processing video based on a facial video generation compression supplemental enhancement information (SEI) message according to some embodiments of the present disclosure. FIG. 14 illustrates another example syntax of the disclosed facial video generation compression SEI message according to some embodiments of the present disclosure. Method 1300 describes the general syntax structure and syntax element order of the facial video generation compression SEI message. Method 1300 may be performed by an encoder (e.g., by process 200A of FIG. 2A or 200B of FIG. 2B) or by one or more software or hardware components of an apparatus (e.g., apparatus 400 in FIG. 4). For example, a processor (e.g., processor 402 in FIG. 4) may perform method 1300. In some embodiments, method 1300 may be implemented by a computer program product embodied in a computer-readable medium including computer-executable instructions, such as program code, executed by a computer (e.g., apparatus 400 in FIG. 4). 13 and 14, the method 1300 may include the following steps 1302 to 1310.
[0191] In step 1302, an identification number (eg, fi_id) is signaled to indicate whether a facial image generation compression scheme is used, see 1401 shown in FIG.
[0192] In step 1304, a parameter indicating the number of face frames, for example, fi_num_set_of_parameter, is signaled, see 1402 shown in FIG.
[0193] In step 1306, parameters indicating quantization factors for processing facial semantic parameters are signaled for each face frame, respectively, with reference to 1403 shown in Figure 14. The facial semantic parameters may include 14 facial semantic parameters (i.e., fi_location[i], fi_rotation_roll[i], fi_rotation_pitch[i], fi_rotation_yaw[i], fi_translation_x[i], fi_translation_y[i], fi_translation_z[i], fi_eye[i], fi_mouth_para1[i], fi_mouth_para2[i], fi_mouth_para3[i], fi_mouth_para4[i], fi_mouth_para5[i], and fi_mouth_para6[i]).
[0194] In step 1308, and with reference to 1404 shown in FIG. 14, for each face frame, a number of facial present flags are further signaled to indicate the presence and length of specific syntax elements associated with the facial video generation compression scheme.
[0195] In step 1310, for each face frame, corresponding face information parameters are signaled.
[0196] The difference between Table 2 in Figure 12 and Table 3 in Figure 14 is that in Figure 14, different fi_quantization_factor for face information parameters are signaled for different face pictures. The process according to Figures 13 and 14 is described as follows:
[0197] First, the facial image generation compression includes an identification number that can be used to identify the facial image generation compression filter.
[0198] Second, every SEI message always has a base picture (i.e., the first face frame in a face sequence) included in the PU, which can provide rich texture references so that the face frame can be reconstructed using the face information parameters carried in the SEI message.
[0199] Third, for these facial information parameters in the SEI message, it is proposed to set five corresponding facial information present flags (i.e., fi_head_location_present_flag[i], fi_head_rotation_present_flag[i], fi_head_translation_present_flag[i], fi_eye_blinking_present_flag[i], fi_mouth_motion_present_flag[i]) to determine whether the related facial information parameters of each face picture[i] are transmitted or not.
[0200] Fourth, if fi_head_location_present_flag[i] is present, the head location parameter fi_location[i] shall be carried in the SEI message. If fi_head_rotation_present_flag[i] is present, these head rotation parameters (fi_rotation_roll[i], fi_rotation_pitch[i], and fi_rotation_yaw[i]) shall be carried in the SEI message. If fi_head_translation_present_flag[i] is present, these head translation parameters (fi_translation_x[i], fi_translation_y[i], and fi_translation_z[i]) shall be carried in the SEI message. If fi_eye_blinking_flag[i] is present, the blink parameter fi_eye[i] shall be carried in the SEI message. If fi_mouth_motion_present_flag[i] is present, these mouth motion parameters (fi_mouth_para1[i], fi_mouth_para2[i], fi_mouth_para3[i], fi_mouth_para4[i], fi_mouth_para5[i], and fi_mouth_para6[i]) shall be carried in the SEI message. If any of these five facial information present flags is not present, there are two situations according to this disclosure.
[0201] In some embodiments, the facial information parameters of the current face frame (i.e., face frame[i]) are copied from the corresponding information from the previous face frame (i.e., face frame[i-1]). For the first face frame ([i=0]), the facial information parameters are copied from the base picture.
[0202] In some embodiments, the face information parameters of the current face frame (ie, face frame[i]) are copied from the base picture.
[0203] Fifth, if these corresponding facial information parameters are carried in this SEI message, the face video can be reconstructed towards personalized characterization or a user-friendly way through the powerful generative capabilities of generative adversarial networks.
[0204] The semantics associated with the syntax in Figure 14 are as follows:
[0205] fi_id contains an identification number that can be used to identify the face image generation compression filter. The value of fi_id ranges from 0 to 2. 32 It is assumed to be in the range of ~2 (inclusive).
[0206] fi_num_set_of_parameter indicates the number of face frames that can be used to realize face image generation compression using the face image generation compression filter. The value of fi_num_set_of_parameter is between 0 and 2. 10 (inclusive), beyond which other Face frames must be packaged into the next SEI message.
[0207] fi_quantization_factor[i] is the quantization factor for processing these 14 facial semantic parameters (i.e., fi_location[i], fi_rotation_roll[i], fi_rotation_pitch[i], fi_rotation_yaw[i], fi_translation_x[i], fi_translation_y[i], fi_translation_z[i], fi_eye[i], fi_mouth_para1[i], fi_mouth_para2[i], fi_mouth_para3[i], fi_mouth_para4[i], fi_mouth_para5[i], and fi_mouth_para6[i]) in float16 type. These float16 parameters shall be further expanded via fi_quantization_factor, where the value of fi_quantization_factor[i] shall be between 0 and 10.16 For example, the original facial parameter value is 0.1234567891234567, and the fi_quantization_factor value is 10 6 , so the corresponding quantized face information parameter is 123456.
[0208] If present, fi_head_location_present_flag[i] is equal to 1; if not present, fi_head_location_present_flag is equal to 0.
[0209] fi_location[i] specifies the quantized residual parameter corresponding to the head location between the ith face frame and the (i-1)th face frame via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_location[0] specifies the quantized head location parameter from the 0th face frame (base picture).
[0210] If present, fi_head_rotation_present_flag[i] is equal to 1; if not present, fi_head_rotation_present_flag is equal to 0.
[0211] fi_rotation_roll[i] specifies the quantized residual parameter corresponding to the head rotation (called roll) around the front-to-back axis between the ith and (i-1)th face frames via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_rotation_roll[0] specifies the quantized front-to-back axis head rotation parameter from the 0th face frame (base picture).
[0212] fi_rotation_pitch[i] specifies the quantized residual parameter corresponding to the head rotation (called pitch) around the side-to-side axis between the ith and (i-1)th face frames via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_rotation_pitch[0] specifies the quantized side-to-side axis head rotation parameter from the 0th face frame (base picture).
[0213] fi_rotation_yaw[i] specifies the quantized residual parameter corresponding to the head rotation around the vertical axis (called yaw) between the ith and (i-1)th face frames via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_rotation_yaw[0] specifies the quantized vertical axis head rotation parameter from the 0th face frame (base picture).
[0214] If present, fi_head_translation_present_flag[i] is equal to 1; if not present, fi_head_translation_present_flag is equal to 0.
[0215] fi_translation_x[i] specifies the quantized residual parameter corresponding to the head translation about the x-axis between the ith and (i-1)th face frames via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_translation_x[0] specifies the quantized x-axis head translation parameter from the 0th face frame (base picture).
[0216] fi_translation_y[i] specifies the quantized residual parameter corresponding to the head translation about the y-axis between the ith and (i-1)th face frames via fi_quantization_factor, if i is not equal to 0. If i is equal to 0, fi_translation_y[0] specifies the quantized y-axis head translation parameter from the 0th face frame (base picture).
[0217] fi_translation_z[i] specifies the quantized residual parameter corresponding to the head translation about the z-axis between the ith and (i-1)th face frames via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_translation_z[0] specifies the quantized z-axis head translation parameter from the 0th face frame (base picture).
[0218] If present, fi_eye_blinking_present_flag[i] is equal to 1; if not present, fi_eye_blinking_present_flag is equal to 0.
[0219] fi_eye[i] specifies the quantized residual parameter corresponding to the degree of blinking between the ith and (i-1)th face frames via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_eye[0] specifies the quantized blinking parameter from the 0th face frame (base picture).
[0220] If present, fi_mouth_motion_present_flag[i] is equal to 1; if not present, fi_mouth_motion_present_flag is equal to 0.
[0221] fi_mouth_para1[i] specifies the quantized residual parameters corresponding to the mouth motion between the ith face frame and the (i-1)th face frame via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_mouth_para1[0] specifies the quantized mouth motion parameters from the 0th face frame (base picture).
[0222] fi_mouth_para2[i] specifies the quantized residual parameters corresponding to the mouth motion between the ith face frame and the (i-1)th face frame via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_mouth_para2[0] specifies the quantized mouth motion parameters from the 0th face frame (base picture).
[0223] fi_mouth_para3[i] specifies the quantized residual parameters corresponding to the mouth motion between the ith face frame and the (i-1)th face frame via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_mouth_para3[0] specifies the quantized mouth motion parameters from the 0th face frame (base picture).
[0224] fi_mouth_para4[i] specifies the quantized residual parameters corresponding to the mouth motion between the ith face frame and the (i-1)th face frame via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_mouth_para4[0] specifies the quantized mouth motion parameters from the 0th face frame (base picture).
[0225] fi_mouth_para5[i] specifies the quantized residual parameters corresponding to the mouth motion between the ith face frame and the (i-1)th face frame via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_mouth_para5[0] specifies the quantized mouth motion parameters from the 0th face frame (base picture).
[0226] fi_mouth_para6[i] specifies the quantized residual parameters corresponding to the mouth motion between the ith face frame and the (i-1)th face frame via fi_quantization_factor if i is not equal to 0. If i is equal to 0, fi_mouth_para6[0] specifies the quantized mouth motion parameters from the 0th face frame (base picture).
[0227] FIG. 15 is a flowchart of an exemplary method 1500 for processing video based on a facial video generation compression supplemental enhancement information (SEI) message according to some embodiments of the present disclosure. Method 1500 describes the general syntax structure and syntax element order of a facial video generation compression SEI message. Method 1500 may be performed by a decoder (e.g., by process 300A of FIG. 3A or 300B of FIG. 3B) or by one or more software or hardware components of an apparatus (e.g., apparatus 400 in FIG. 4). For example, a processor (e.g., processor 402 in FIG. 4) may perform method 1500. In some embodiments, method 1500 may be implemented by a computer program product embodied in a computer-readable medium, including computer-executable instructions, such as program code, executed by a computer (e.g., apparatus 400 in FIG. 4). Referring to FIG. 15, method 1500 may include the following steps 1502-1506:
[0228] In step 1502, it is determined whether a facial image generation compression scheme is used based on the identification number.
[0229] In step 1504, in response to determining that a facial image generation compression scheme is to be used, a Supplementary Enhancement Information (SEI) message is decoded. The SEI message includes facial information, for example, as shown in FIG. 10, FIG. 12, or FIG. 14.
[0230] In some embodiments, decoding the SEI message further includes decoding a flag indicating whether a syntax element associated with the face information is present in the SEI message, and, in response to the presence of the syntax element associated with the face information, decoding the face information based on the syntax element associated with the face information.
[0231] In some embodiments, decoding the SEI message further includes decoding a syntax element indicating the number of face frames that use the facial video generation compression scheme, and decoding a flag for each face frame indicating whether a syntax element associated with face information is present in the SEI message.
[0232] In some embodiments, decoding the SEI message further includes decoding a syntax element indicating the number of face frames that use the facial image generation compression scheme, and decoding a plurality of facial present flags, each for each face frame, that indicate the presence and length of a particular syntax element associated with the facial image generation compression scheme.
[0233] In some embodiments, decoding the SEI message further includes decoding, for each face frame, a factor indicating a quantization factor for processing the face information.
[0234] In step 1506, a face picture is reconstructed based on the face information and the base picture associated with the SEI message.
[0235] In some embodiments, the facial information for the current face frame (i.e., face frame[i]) is copied from the corresponding information from the previous face frame (i.e., face frame[i-1]). For the first face frame ([i=0]), the facial information is copied from the base picture.
[0236] In some embodiments, the face information of the current face frame (ie, face frame[i]) is copied from the base picture.
[0237] In some embodiments, a receiving module configured to receive the bitstream; a decoding module configured to decode one or more pictures using coding information of the bitstream; The decoding module determining whether a facial image generation compression scheme is to be used based on the identification number; In response to determining that the facial image generating compression scheme is to be used, decoding a supplemental enhancement information (SEI) message including facial information; configured to reconstruct a face picture based on the face information and a base picture associated with the SEI message. A decoding device is provided.
[0238] In one implementation, the SEI message further includes a flag indicating whether a syntax element associated with face information is present in the SEI message, and the decoding module: In response to the presence of a syntax element associated with the face information, the facial information is configured to decode the face information based on the syntax element associated with the face information.
[0239] In one implementation, the SEI message further includes a syntax element indicating a number of face frames using the facial image generation compression scheme, and the decoding module: The flag indicating whether a syntax element associated with face information is present in the SEI message is decoded for each face frame.
[0240] In one implementation, the SEI message further includes a syntax element indicating a number of face frames using the facial image generation compression scheme, and the decoding module: The facial information is decoded for each face frame based on syntax elements associated with the facial information.
[0241] In one implementation, the decoding device comprises: The device further includes a copy module configured to copy corresponding face information from a previous face frame in response to a flag indicating that no syntax element associated with face information exists for the current face frame.
[0242] In one implementation, the SEI message further includes a base picture as a reference, and the copy module: For a first face frame, the method is configured to copy corresponding face information from the base picture in response to a flag indicating that a syntax element associated with face information is not present.
[0243] In one implementation, the SEI message further includes a base picture as a reference, and the decoding device: The apparatus further includes a copy module configured to copy corresponding face information from the base picture in response to a flag indicating that a syntax element associated with face information does not exist for a current face frame.
[0244] In one implementation, the SEI message further includes a factor indicating a quantization factor for processing the face information, and the decoding module: configured to decrypt the factor; The decoding device The device further includes a processing module configured to process the facial information based on the factors.
[0245] In one implementation, the SEI message further includes a syntax element indicating a number of face frames using the facial image generation compression scheme, and the decoding module: configured to decode the factors for each face frame, The processing module includes: The facial information is processed based on the factors for each face frame.
[0246] In some embodiments, a receiving module configured to receive a video sequence; an encoding module configured to encode one or more pictures of the video sequence; a generating module configured to generate a bitstream; The encoding module: configured to signal an identification number indicating whether a facial image generation compression scheme is used; An encoding device is provided.
[0247] In one implementation, when the facial image generating compression scheme is used, the encoding module: signaling a flag indicating whether a syntax element associated with face information is present; If the flag indicates that a syntax element associated with the face information exists, the syntax element associated with the corresponding face information is configured to signal.
[0248] In one implementation, the encoding module: signaling a syntax element indicating a number of face frames using the facial image generation compression scheme; A flag indicating whether a syntax element associated with face information is present or not is configured to be signaled for each face frame.
[0249] In one implementation, the encoding module: signaling a syntax element indicating a number of face frames using the facial image generation compression scheme; The device is configured to signal, for each face frame, a factor indicative of a quantization factor for processing the face information.
[0250] In some embodiments, an electronic device is provided that includes a memory that stores a set of instructions and one or more processors that are configured to execute the set of instructions to cause the one or more processors to perform a method for decoding a bitstream and outputting one or more pictures for a video stream as described in the embodiment of the method for decoding a bitstream and outputting one or more pictures for a video stream described above.
[0251] In some embodiments, an electronic device is provided that includes a memory that stores a set of instructions and one or more processors that are configured to execute the set of instructions to cause the one or more processors to perform a method for encoding a video sequence into a bitstream as described in the above-mentioned embodiment of the method for encoding a video sequence into a bitstream.
[0252] In some embodiments, a non-transitory computer-readable storage medium is provided that stores a video bitstream that, when decoded by a processor, causes the processor to perform the method for decoding a bitstream to output one or more pictures for a video stream described in the embodiment of the method for decoding a bitstream to output one or more pictures for a video stream described above.
[0253] In some embodiments, a non-transitory computer-readable storage medium is provided that stores a video sequence of video that, when encoded by a processor, causes the processor to perform a method for encoding a video sequence into a bitstream as described in an embodiment of the method for encoding a video sequence into a bitstream.
[0254] In some embodiments, a computer program product is provided that includes computer program instructions that enable a computer to perform the method for decoding a bitstream and outputting one or more pictures for a video stream described in the embodiment of the method for decoding a bitstream and outputting one or more pictures for a video stream described above.
[0255] In some embodiments, a computer program product is provided that includes computer program instructions that enable a computer to perform the method for encoding a video sequence into a bitstream described in the embodiments of the method for encoding a video sequence into a bitstream described above.
[0256] In some embodiments, a computer program is provided that enables a computer to perform the method for decoding a bitstream and outputting one or more pictures for a video stream described in the above-described embodiment of the method for decoding a bitstream and outputting one or more pictures for a video stream.
[0257] In some embodiments, a computer program is provided that enables a computer to perform the method for encoding a video sequence into a bitstream described in the embodiments of the method for encoding a video sequence into a bitstream described above.
[0258] The following clauses may be used to further describe the embodiments. 1. A method for decoding a bitstream to output one or more pictures for a video stream, comprising: receiving a bitstream; decoding one or more pictures using coding information of the bitstream; The step of decoding comprises: determining whether a facial image generation compression scheme is to be used based on the identification number; In response to determining that the facial image generating compression scheme is to be used, decoding a supplemental enhancement information (SEI) message containing facial information; and reconstructing a face picture based on the face information and a base picture associated with the SEI message. method. 2. The SEI message further includes a flag indicating whether a syntax element associated with face information is present in the SEI message, and the method further comprises: Item 1, the method of Item 1 further comprising, in response to the presence of a syntax element associated with the face information, decoding the face information based on the syntax element associated with the face information. 3. The SEI message further includes a syntax element indicating a number of face frames using the facial image generation compression scheme, and the method further comprises: Item 3. The method of item 2, further comprising the step of decoding the flag indicating whether a syntax element associated with face information is present in the SEI message for each face frame. 4. The SEI message further includes a syntax element indicating a number of face frames using the facial image generation compression scheme, and the method further comprises: Item 3. The method of item 2, further comprising decoding the face information for each face frame based on syntax elements associated with the face information. 5. Further including the step of copying corresponding face information from a previous face frame according to a flag indicating that no syntax element associated with face information exists for the current face frame; The method described in item 4. 6. The SEI message further includes a base picture as a reference, and the method further comprises: The method of clause 5, further comprising the step of, for a first face frame, copying corresponding face information from the base picture in response to a flag indicating that a syntax element associated with face information is not present. 7. The SEI message further includes a base picture as a reference, and the method further comprises: Item 5. The method of item 4, further comprising the step of copying corresponding face information from the base picture in response to a flag indicating that a syntax element associated with face information does not exist for the current face frame. 8. The SEI message further includes a factor indicating a quantization factor for processing the face information, and the method further comprises: Decrypting the factor; Item 8. The method according to any one of items 1 to 7, further comprising: processing the face information based on the factor. 9. The SEI message further includes a syntax element indicating a number of face frames using the facial image generation compression scheme, and the method further comprises: decoding the factors for each face frame; Item 9. The method of item 8, further comprising the step of processing the face information based on the factors for each face frame. 10. A method of encoding a video sequence into a bitstream, comprising: receiving a video sequence; encoding one or more pictures of the video sequence; generating a bitstream; The step of encoding comprises: signaling an identification number indicating whether a facial image generation compression scheme is used; method. 11. When the facial image generation compression scheme is used, the method further comprises: signaling a flag indicating whether a syntax element associated with face information is present; 11. The method of claim 10, further comprising the step of: if the flag indicates that a syntax element associated with the face information exists, signaling a syntax element associated with the corresponding face information. 12. Signaling a syntax element indicating the number of face frames that use the facial image generation compression scheme; signaling a flag for each face frame indicating whether a syntax element associated with face information is present or not; Item 12. The method according to item 11, further comprising: 13. Signaling a syntax element indicating the number of face frames that use the facial image generation compression scheme; signaling, for each face frame, a factor indicative of a quantization factor for processing the face information; Item 12. The method according to item 11, further comprising: 14. A non-transitory computer-readable storage medium storing a video bitstream, the bitstream comprising: a supplemental enhancement information (SEI) message containing face information; A non-transitory computer-readable storage medium, wherein the face information is used to reconstruct a face picture based on a base picture associated with the face picture. 15. The non-transitory computer-readable storage medium of clause 14, wherein the SEI message further includes a flag indicating whether a syntax element associated with face information is present in the SEI message. 16. The non-transitory computer-readable storage medium of clause 15, wherein the SEI message further includes a syntax element indicating the number of face frames using a facial image generation compression scheme. 17. The non-transitory computer-readable storage medium of clause 16, wherein the SEI message further includes one or more syntax elements associated with face information for each face frame. 18. A non-transitory computer-readable storage medium as described in clause 14, wherein the SEI message further includes a syntax element indicating the number of face frames that use a facial image generation compression scheme, and a flag indicating for each face frame whether a syntax element associated with facial information is present in the SEI message. 19. The non-transitory computer-readable storage medium of clause 14, wherein the SEI message further includes a factor indicating a quantization factor for processing the face information. 20. A non-transitory computer-readable storage medium as described in clause 14, wherein the SEI message further includes a syntax element indicating the number of face frames using a facial image generation compression scheme and a factor indicating, for each face frame, a quantization factor for processing the facial information.
[0259] In some embodiments, a non-transitory computer-readable storage medium containing instructions is also provided, which may be executed by a device (e.g., the disclosed encoder and decoder) to perform the above-described method. In some embodiments, a non-transitory computer-readable storage medium storing a bitstream or an SEI message is also provided. The bitstream can be encoded and decoded using the facial video generation compression supplemental enhancement information (SEI) message described above (e.g., FIGS. 10, 12, and 14). Common forms of non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape, or other magnetic data storage media, CD-ROMs, any other optical data storage media, physical media with patterns of holes, RAM, PROMs, EPROMs, flash EPROMs or other flash memory, NVRAM, cache, registers, other memory chips or cartridges, and networked versions thereof. A device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.
[0260] It should be noted that, in this specification, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another and do not require or imply an actual relationship or order between those entities or operations. Furthermore, the words "comprising," "having," "containing," "containing," and other similar forms are intended to be equivalent in meaning and open-ended in that they do not imply that the item or items following any of these words are an exhaustive list of the item or items, or that the items are limited to only the listed item or items.
[0261] As used herein, unless otherwise stated, the term "or" includes all possible combinations unless impossible. For example, if it is stated that a database may include A or B, then the database may also include A, or B, or A and B, unless otherwise stated or impossible. As a second example, if it is stated that a database may include A, B, or C, then the database may also include A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C, unless otherwise stated or impossible.
[0262] It is understood that the above embodiments can be realized by hardware, or software (program code), or a combination of hardware and software. If realized by software, it may be stored in the computer-readable medium. When executed by a processor, the software can perform the disclosed methods. The computing units and other functional units described in this disclosure can be implemented by hardware, or software, or a combination of hardware and software. Those skilled in the art can understand that multiple of the above modules / units may be combined into one module / unit, or that each of the above modules / units may be further divided into multiple sub-modules / sub-units.
[0263] In the foregoing specification, embodiments have been described with reference to numerous specific details that vary from embodiment to embodiment. Certain adaptations and variations can be made to the above-described embodiments. Other embodiments will be apparent to those skilled in the art from consideration of the detailed description and practice of the present disclosure disclosed herein. It is intended that the specification and examples be considered exemplary, with a true scope and spirit of the present disclosure being indicated by the following claims. It is also intended that the order of steps depicted in the figures is for illustrative purposes only and is not intended to be limited to the particular order of steps. Thus, one skilled in the art will appreciate that steps can be performed in different orders when performing the same method.
[0264] In the drawings and specification, illustrative embodiments are disclosed. However, these embodiments are susceptible to many variations and modifications. Thus, although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation.
Claims
1. 1. A method for decoding a bitstream to output one or more pictures for a video stream, comprising: receiving a bitstream; decoding one or more pictures using coding information of the bitstream; The step of decoding comprises: determining whether a facial image generation compression scheme is to be used based on the identification number; In response to determining that the facial image generating compression scheme is to be used, decoding a supplemental enhancement information (SEI) message that includes facial information; reconstructing a face picture based on the face information and a base picture associated with the SEI message; method.
2. The SEI message further includes a flag indicating whether a syntax element associated with face information is present in the SEI message, and the method further includes: The method of claim 1 , further comprising the step of: responsive to the presence of a syntax element associated with the face information, decoding the face information based on the syntax element associated with the face information.
3. The SEI message further includes a syntax element indicating a number of face frames using the facial image generation compression scheme, and the method further includes: The method of claim 2 , further comprising the step of respectively decoding the flag for each face frame, the flag indicating whether a syntax element associated with face information is present in the SEI message.
4. The SEI message further includes a syntax element indicating a number of face frames using the facial image generation compression scheme, and the method further includes: The method of claim 2 , further comprising decoding the face information for each face frame based on syntax elements associated with the face information.
5. and further comprising the step of copying corresponding face information from a previous face frame in response to a flag indicating that no syntax element associated with face information exists for the current face frame. The method of claim 4.
6. The SEI message further includes a base picture as a reference, and the method further comprises: The method of claim 5 , further comprising the step of, for a first face frame, copying corresponding face information from the base picture in response to a flag indicating that a syntax element associated with face information is absent.
7. The SEI message further includes a base picture as a reference, and the method further comprises: The method of claim 4 , further comprising the step of: copying corresponding face information from the base picture in response to a flag indicating that a syntax element associated with face information is absent for a current face frame.
8. The SEI message further includes a factor indicating a quantization factor for processing the face information, and the method further comprises: Decrypting the factor; The method of claim 1 or 2, further comprising the step of: processing the face information based on the factors.
9. The SEI message further includes a syntax element indicating a number of face frames using the facial image generation compression scheme, and the method further includes: decoding the factors for each face frame; The method of claim 8 , further comprising: processing the face information based on the factors for each face frame.
10. 1. A method for encoding a video sequence into a bitstream, comprising: receiving a video sequence; encoding one or more pictures of the video sequence; generating a bitstream; The step of encoding comprises: signaling an identification number indicating whether a facial image generation compression scheme is used; method.
11. When the facial image generation compression scheme is used, the method comprises: signaling a flag indicating whether a syntax element associated with face information is present; The method of claim 10 , further comprising: if the flag indicates that a syntax element associated with the face information exists, signaling a syntax element associated with the corresponding face information.
12. signaling a syntax element indicating a number of face frames using the facial image generation compression scheme; signaling a flag for each face frame indicating whether a syntax element associated with face information is present or not; The method of claim 11 further comprising:
13. signaling a syntax element indicating a number of face frames using the facial image generation compression scheme; signaling, for each face frame, a factor indicative of a quantization factor for processing the face information; The method of claim 11 further comprising:
14. a receiving module configured to receive the bitstream; a decoding module configured to decode one or more pictures using coding information of the bitstream; The decoding module determining whether a facial image generation compression scheme is to be used based on the identification number; In response to determining that the facial image generating compression scheme is to be used, decoding a supplemental enhancement information (SEI) message including facial information; configured to reconstruct a face picture based on the face information and a base picture associated with the SEI message. Decoding device.
15. The SEI message further includes a flag indicating whether a syntax element associated with face information is present in the SEI message, and the decoding module: The decoding device of claim 14 , configured to, in response to a presence of a syntax element associated with the face information, decode the face information based on the syntax element associated with the face information.
16. The SEI message further includes a syntax element indicating a number of face frames using the facial image generation compression scheme, and the decoding module: The decoding device according to claim 15 , configured to decode the flag indicating whether a syntax element associated with face information is present in the SEI message, respectively, for each face frame.
17. The SEI message further includes a syntax element indicating a number of face frames using the facial image generation compression scheme, and the decoding module: The decoding device according to claim 15 , configured to decode the face information for each face frame based on syntax elements associated with the face information.
18. and a copy module configured to copy corresponding face information from a previous face frame in response to a flag indicating that a syntax element associated with face information does not exist for the current face frame.
18. A decoding device according to claim 17.
19. The SEI message further includes a base picture as a reference, and the copy module:
20. The decoding device of claim 18, configured to, for a first face frame, copy corresponding face information from the base picture in response to a flag indicating the absence of a syntax element associated with face information.
20. The SEI message further includes a base picture as a reference, and the decoding device:
18. The decoding device of claim 17, further comprising: a copy module configured to copy corresponding face information from the base picture in response to a flag indicating that a syntax element associated with face information is not present for a current face frame.
21. The SEI message further includes a factor indicating a quantization factor for processing the face information, and the decoding module: configured to decrypt the factor; The decoding device 16. The decoding device according to claim 14 or 15, further comprising a processing module configured to process the face information based on the factors.
22. The SEI message further includes a syntax element indicating a number of face frames using the facial image generation compression scheme, and the decoding module: configured to decode the factors for each face frame, The processing module includes: The decoding device of claim 21 , configured to process the face information based on the factors for each face frame.
23. a receiving module configured to receive a video sequence; an encoding module configured to encode one or more pictures of the video sequence; a generating module configured to generate a bitstream; The encoding module: configured to signal an identification number indicating whether a facial image generation compression scheme is used; Encoding device.
24. When the facial image generating compression scheme is used, the encoding module: signaling a flag indicating whether a syntax element associated with face information is present; 24. The encoding device of claim 23, configured to signal a syntax element associated with corresponding face information if the flag indicates that the syntax element associated with the face information is present.
25. The encoding module: signaling a syntax element indicating a number of face frames using the facial image generation compression scheme; 25. The encoding device of claim 24, configured to signal a flag for each face frame indicating whether a syntax element associated with face information is present or not.
26. The encoding module: signaling a syntax element indicating a number of face frames using the facial image generation compression scheme; The encoding device of claim 24, configured to signal, for each face frame, a factor indicative of a quantization factor for processing the face information.
27. 10. An electronic device comprising: a memory storing a set of instructions; and one or more processors configured to execute the set of instructions to cause the one or more processors to perform the method of decoding a bitstream and outputting one or more pictures for a video stream according to any one of claims 1 to 9.
28. 14. An electronic device comprising: a memory storing a set of instructions; and one or more processors configured to execute the set of instructions to cause the one or more processors to perform the method for encoding a video sequence into a bitstream according to any one of claims 10 to 13.
29. A non-transitory computer-readable storage medium storing a video bitstream that, when decoded by a processor, causes the processor to perform the method of decoding a bitstream and outputting one or more pictures for a video stream according to any one of claims 1 to 9.
30. A non-transitory computer readable storage medium storing a video sequence of videos that, when encoded by a processor, causes the processor to perform the method of encoding a video sequence into a bitstream according to any one of claims 10 to 13.
31. 10. A computer program product comprising computer program instructions that enable a computer to perform the method of decoding a bitstream to output one or more pictures for a video stream according to any one of claims 1 to 9.
32. A computer program product comprising computer program instructions that enable a computer to carry out the method for encoding a video sequence into a bitstream according to any one of claims 10 to 13.
33. A computer program product enabling a computer to carry out the method for decoding a bitstream according to any one of claims 1 to 9 to output one or more pictures for a video stream.
34. A computer program product enabling a computer to carry out the method for encoding a video sequence into a bitstream according to any one of claims 10 to 13.