Method and apparatus for generating compressed SEI message using face video
By using facial video to generate compressed supplementary enhanced information (SEI) messages, the problem of low coding efficiency in the prior art is solved, and more efficient facial video compression and decoding is achieved, and video quality and storage/transmission efficiency are improved.
Patent Information
- Application Number
- CN202380089823.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2023-12-27
- Publication Date
- 2025-08-19
AI Technical Summary
The existing video encoding technology is less efficient in facial video processing, and it is difficult to effectively use facial feature information for compression and decoding.
Using the method of generating compressed supplementary enhancement information (SEI) messages on the face video, the SEI message is decoded to reconstruct the facial image by receiving a bitstream and determining whether to use the face video to generate a compression scheme based on the identification number.
It improves the encoding efficiency of facial videos, can more effectively use facial feature information for compression and decoding, and improves video quality and storage/transmission efficiency.
Smart Images

Figure CN120513631A_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This disclosure claims priority to U.S. Provisional Application No. 63 / 436,626, filed on January 1, 2023, and U.S. Patent Application No. 18 / 392,557, filed on December 21, 2023, entitled “METHOD AND APPARATUSES FOR USING FACE VIDEO GENERATIVE COMPRESSION SEI MESSAGE.” The entire contents of all of the foregoing applications are incorporated herein by reference. Technical Field
[0002] The present disclosure relates generally to video processing and, more particularly, to methods and apparatus for generating compressed Supplemental Enhancement Information (SEI) messages using facial videos. Background Art
[0003] A video is a set of static images (or "frames") that capture visual information. In order to reduce storage memory and transmission bandwidth, videos can be compressed before storage or transmission, and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, usually based on prediction, transform, quantization, entropy coding and loop filtering. Standardization organizations have developed video coding standards that specify specific video coding formats, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard and the AVS standard. As more and more advanced video coding technologies are adopted by video standards, the coding efficiency of new video coding standards is getting higher and higher. Summary of the Invention
[0004] In a first aspect, an embodiment of the present disclosure provides a method for decoding a bitstream to output one or more images of a video stream, the method comprising: receiving a bitstream; and decoding one or more images using encoding information of the bitstream; the decoding comprising: determining whether a facial video generation compression scheme is used based on an identification number; in response to a determination that the facial video generation compression scheme is used, decoding a supplemental enhancement information SEI message, the SEI message including facial information; and reconstructing a facial image based on the facial information and a base image associated with the SEI message.
[0005] In a second aspect, an embodiment of the present disclosure provides a method for encoding a video sequence into a bit stream, the method comprising: receiving a video sequence; encoding one or more images of the video sequence; and generating a bit stream; the encoding comprises: sending an identification number, the identification number indicating whether a facial video generation compression scheme is used.
[0006] In a third aspect, embodiments of the present disclosure provide a decoding device, comprising: a receiving module configured to receive a bitstream; and a decoding module configured to decode one or more images using encoding information in the bitstream. The decoding module is configured to determine, based on an identification number, whether a facial video generation compression scheme is used; in response to a determination that the facial video generation compression scheme is used, decode an SEI message, the SEI message including facial information; and reconstruct a facial image based on the facial information and a base image associated with the SEI message.
[0007] In a fourth aspect, an embodiment of the present invention provides an encoding apparatus, comprising: a receiving module configured to receive a video sequence; an encoding module configured to encode one or more images of the video sequence; and a generation module configured to generate a bitstream. The encoding module is configured to signal an identification number indicating whether a facial video generation compression scheme is used.
[0008] In a fifth aspect, an embodiment of the present invention provides an electronic device comprising: a memory storing an instruction set; and one or more processors configured to execute the instruction set, so that the one or more processors perform the method of decoding a bit stream to output one or more images of a video stream according to the first aspect.
[0009] In a sixth aspect, an embodiment of the present invention provides an electronic device comprising: a memory storing an instruction set; and one or more processors configured to execute the instruction set so that the one or more processors perform the method of encoding a video sequence into a bit stream according to the second aspect.
[0010] In the seventh aspect, an embodiment of the present invention provides a non-temporary computer-readable storage medium that stores a video bit stream. When the bit stream is decoded by a processor, the processor executes the method of decoding the bit stream according to the first aspect to output one or more images of the video stream.
[0011] In an eighth aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium that stores a video sequence of a video. When the video sequence is encoded by a processor, the processor executes the method of encoding the video sequence into a bit stream according to the second aspect.
[0012] In a ninth aspect, an embodiment of the present invention provides a computer program product comprising: a plurality of computer program instructions, and wherein the plurality of computer program instructions enable a computer to execute the method of decoding a bit stream to output one or more images of a video stream according to the first aspect.
[0013] In a tenth aspect, an embodiment of the present invention provides a computer program product, comprising: a plurality of computer program instructions, and wherein the plurality of computer program instructions enable a computer to execute the method of encoding a video sequence into a bit stream according to the second aspect.
[0014] In an eleventh aspect, an embodiment of the present invention provides a computer program, and the computer program enables a computer to execute the method of decoding a bit stream to output one or more images of a video stream according to the first aspect.
[0015] In a twelfth aspect, an embodiment of the present invention provides a computer program, and the computer program enables a computer to execute the method of encoding a video sequence into a bit stream according to the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Embodiments and aspects of the present disclosure are illustrated in the following detailed description and accompanying drawings.The various features shown in the accompanying drawings are not drawn to scale.
[0017] Figure 1 A schematic diagram of an exemplary system for preprocessing and encoding image data according to some embodiments of the present disclosure is shown.
[0018] Figure 2A A schematic diagram illustrating an exemplary encoding process of a hybrid video coding system consistent with an embodiment of the present disclosure is shown.
[0019] Figure 2B A schematic diagram illustrating another exemplary encoding process of a hybrid video coding system consistent with an embodiment of the present disclosure is shown.
[0020] Figure 3A A schematic diagram illustrating an exemplary decoding process of a hybrid video coding system consistent with an embodiment of the present disclosure is shown.
[0021] Figure 3B A schematic diagram illustrating another exemplary decoding process of a hybrid video coding system consistent with an embodiment of the present disclosure is shown.
[0022] Figure 4 A block diagram of an exemplary apparatus for pre-processing or encoding image data according to some embodiments of the present disclosure is shown.
[0023] Figure 5A schematic diagram of an exemplary deep learning-based video generation and compression framework according to some embodiments of the present disclosure is shown.
[0024] Figure 6 A schematic diagram illustrating an exemplary encoder-decoder encoding framework with a compact feature size of 1×4×4 for speaking face videos according to some embodiments of the present disclosure is shown.
[0025] Figure 7 A schematic diagram illustrating a general encoder-decoder generation compression framework for 3DMM-assisted speaking face videos according to some embodiments of the present disclosure is shown.
[0026] Figure 8 is a flowchart of an exemplary method for generating compressed supplemental enhancement information (SEI) messages based on facial video to process video according to some embodiments of the present disclosure
[0027] Figure 9 is a flow chart of an exemplary method for processing video by generating compressed Supplemental Enhancement Information (SEI) messages based on facial video, according to some embodiments of the present disclosure.
[0028] Figure 10 An exemplary syntax of the disclosed facial video generation compression SEI message according to some embodiments of the present disclosure is shown.
[0029] Figure 11 is a flow chart of an exemplary method for processing video by generating compressed Supplemental Enhancement Information (SEI) messages based on facial video, according to some embodiments of the present disclosure.
[0030] Figure 12 Another exemplary syntax of the disclosed facial video generation compression SEI message according to some embodiments of the present disclosure is shown.
[0031] Figure 13 is a flow chart of an exemplary method for processing video by generating compressed Supplemental Enhancement Information (SEI) messages based on facial video, according to some embodiments of the present disclosure.
[0032] Figure 14 Another exemplary syntax of the disclosed facial video generation compression SEI message according to some embodiments of the present disclosure is shown.
[0033] Figure 15 is a flow chart of an exemplary method for processing video by generating compressed Supplemental Enhancement Information (SEI) messages based on facial video, according to some embodiments of the present disclosure. DETAILED DESCRIPTION
[0034] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which, unless otherwise specified, the same reference numerals in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with the aspects related to the present disclosure recited in the appended claims. Specific aspects of the present disclosure are described in more detail below. In the event of a conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein shall prevail.
[0035] The Joint Video Experts Group (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) is currently developing a Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, VVC aims to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.
[0036] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET has been developing technologies beyond HEVC using the Joint Exploration Model (JEM) reference software. As coding technologies are incorporated into JEM, JEM achieves significantly higher coding performance than HEVC.
[0037] The VVC standard has recently been finalized and is continually incorporating new coding techniques to provide even better compression performance. VVC uses the same hybrid video coding system used by modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.
[0038] A video is a set of static images (or "frames") arranged in a temporal sequence to store visual information. A video capture device (e.g., a camera) can be used to capture and store those images in temporal order, and a video playback device (e.g., a television, computer, smartphone, tablet, video player, or any end-user terminal with a display) can be used to display such images in temporal order. In addition, in some applications, the video capture device can transmit the captured video to the video playback device (e.g., a computer with a display) in real time, for example, for video observation, conferencing, or live broadcasting.
[0039] To reduce the storage space and transmission bandwidth required for these applications, the video can be compressed before storage and transmission, and decompressed before display. The compression and decompression can be implemented by software or dedicated hardware executed by a processor (e.g., a processor of a general-purpose computer). The module used for compression is generally referred to as an "encoder," while the module used for decompression is generally referred to as a "decoder." The encoder and decoder can be collectively referred to as "codecs." The encoder and decoder can be implemented as any of various suitable hardware, software, or combinations thereof. For example, the hardware implementation of the encoder and decoder can include circuits—such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. The software implementation of the encoder and decoder can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression can be implemented by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, and the like. In some applications, the codec may decompress the video from a first coding standard and recompress the decompressed video using a second coding standard, in which case the codec may be referred to as a "transcoder."
[0040] The video encoding process identifies and retains useful information that can be used to reconstruct an image, while ignoring less important information for that reconstruction. If the omitted, less important information cannot be fully reconstructed, the encoding process is called "lossy." Otherwise, it is called "lossless." Most encoding processes are lossy as a trade-off to reduce required storage space and transmission bandwidth.
[0041] Useful information about an image being encoded (called the "current image") includes changes relative to a reference image (e.g., a previously encoded and reconstructed image). Such changes can include changes in pixel position, brightness, or color, with position changes being of most interest. Changes in the position of a group of pixels representing an object can reflect the object's motion between the reference image and the current image.
[0042] A picture that is encoded without reference to another picture (i.e., it is its own reference picture) is called an "I-picture." If some or all of the blocks in the picture (e.g., these blocks typically represent portions of a video picture) are predicted using intra-frame prediction or inter-frame prediction with the help of a reference picture (e.g., unidirectional prediction), then the picture is called a "P-picture" (or "P-frame"). If at least one block in the picture is predicted using two reference pictures (e.g., bidirectional prediction), then the picture is called a "B-picture."
[0043] Figure 1 A block diagram of a system 100 for preprocessing and encoding image data according to some embodiments of the present disclosure is shown. The image data may include an image (also referred to as a "picture" or "frame"), multiple images, or a video. An image is a static image. Multiple images may be spatially or temporally related or unrelated. A video is a set of images arranged in a time sequence.
[0044] like Figure 1 As shown, system 100 includes a source device 120 that provides encoded video data that is subsequently decoded by a destination device 140. Consistent with the disclosed embodiments, source device 120 and destination device 140 can each include any of a variety of devices, including: a desktop computer, a notebook (e.g., laptop) computer, a server, a tablet computer, a set-top box, a mobile phone, a vehicle, a camera, an image sensor, a robot, a television, a wearable device (e.g., a smartwatch or wearable camera), a display device, a digital media player, a video game console, a video streaming device, etc. Source device 120 and destination device 140 can be equipped for wireless or wired communication.
[0045] refer to Figure 1 , the source device 120 may include an image / video preprocessor 122, an image / video encoder 124, and an output interface 126. The target device 140 may include an input interface 142, an image / video decoder 144, and one or more machine vision applications 146. The image / video preprocessor 122 preprocesses image data, i.e., one or more images or one or more videos, and generates an input bitstream for the image / video encoder 124. The image / video encoder 124 encodes the input bitstream and outputs an encoded bitstream 162 via the output interface 126. The encoded bitstream 162 is transmitted over the communication medium 160 and received by the input interface 142. The image / video decoder 144 then decodes the encoded bitstream 162 to generate decoded data, which can be utilized by the machine vision application 146.
[0046] More specifically, the source device 120 may further include various devices (not shown) for providing source image data to be pre-processed by the image / video pre-processor 122. The devices for providing source image data may include image / video acquisition devices, such as cameras, image / video archives or storage devices containing previously acquired images / videos, or image / video feed interfaces for receiving images / videos from image / video content providers.
[0047] The image / video encoder 124 and the image / video decoder 144 can each be implemented in any of a variety of suitable encoder or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When encoding or decoding is partially implemented in software, the image / video encoder 124 or the image / video decoder 144 can store multiple instructions for the software in a suitable, non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform techniques consistent with the present disclosure. The image / video encoder 124 or the image / video decoder 144 can each be included in one or more encoders or decoders, which can be integrated as part of a composite encoder / decoder (CODEC) in the corresponding device.
[0048] The image / video encoder 124 and the image / video decoder 144 may operate according to any video coding standard, such as Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), AOMedia Video 1 (AV1), Joint Photographic Experts Group (JPEG), Moving Picture Experts Group (MPEG), etc. Alternatively, the image / video encoder 124 and the image / video decoder 144 may be custom devices that do not conform to existing standards. Figure 1 Not shown, but in some embodiments, the image / video encoder 124 and the image / video decoder 144 may each be integrated with an audio encoder and decoder, and may include appropriate multiplexer-demultiplexer units (MUX-DEMUX), or other hardware and software, to handle the encoding of audio and video, including encoding both in a common data stream or in separate data streams.
[0049] Output interface 126 may include any type of medium or device capable of transmitting encoded bitstream 162 from source device 120 to destination device 140. For example, output interface 126 may include a transmitter or transceiver configured to transmit encoded bitstream 162 directly from source device 120 to destination device 140 in real-time. Encoded bitstream 162 may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device 140.
[0050] The communication medium 160 may include a transient medium, such as a wireless broadcast or a wired network transmission. For example, the communication medium 160 may include a radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). The communication medium 160 may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. In some embodiments, the communication medium 160 may include a router, a switch, a base station, or any other device that can be used to facilitate communication from the source device 120 to the target device 140. For example, a network server (not shown) may receive the encoded bit stream 162 from the source device 120 and provide the encoded bit stream 162 to the target device 140, for example, via network transmission.
[0051] Communication medium 160 may also take the form of a storage medium (e.g., a non-transitory storage medium) such as a hard drive, a flash drive, an optical disc, a digital video disc, a Blu-ray disc, a volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded image data. In some embodiments, a computer device of a media generation device (e.g., an optical disc pressing device) may receive the encoded image data from source device 120 and create an optical disc containing the encoded video data.
[0052] Input interface 142 may include any type of medium or device capable of receiving information from communication medium 160. The received information includes encoded bitstream 162. For example, input interface 142 may include a receiver or transceiver configured to receive encoded bitstream 162 in real time.
[0053] The machine vision application 146 includes various hardware and / or software for utilizing the decoded image data generated by the image / video decoder 144. For example, the machine vision application 146 may include a display device that displays the decoded image data to a user, and may include any of a variety of display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices. As another example, the machine vision application 146 may include one or more processors configured to use the decoded image data to perform various machine vision applications, such as object recognition and tracking, facial recognition, image matching, image / video search, augmented reality, robotic vision and navigation, autonomous driving, three-dimensional structure construction, stereo correspondence, motion tracking, etc.
[0054] Next, combine Figures 2A-2B and Figures 3A-3B Detailed description of exemplary image data encoding and decoding techniques, such as those utilized by image / video encoder 124 and image / video decoder 144, is provided below.
[0055] Figure 2A Schematic diagram of an exemplary encoding process 200A consistent with an embodiment of the present disclosure is shown. For example, the encoding process 200A may be performed by a processor such as Figure 1 The image / video encoder 124 in FIG. Figure 2A As shown, the encoder can encode a video sequence 202 into a video bitstream 228 according to process 200A. The video sequence 202 may include a set of images (referred to as "original images") arranged in chronological order. Each original image of the video sequence 202 may be divided by the encoder into a plurality of basic processing units, a plurality of basic processing sub-units, or a plurality of regions for processing. In some embodiments, the encoder may perform process 200A at the basic processing unit level for each original image of the video sequence 202. For example, the encoder may perform process 200A in an iterative manner, wherein the encoder may encode one basic processing unit in one iteration of process 200A. In some embodiments, the encoder may perform process 200A in parallel for multiple regions of each original image of the video sequence 202.
[0056] exist Figure 2A 200A, the encoder may feed the basic processing units of the original images of the video sequence 202 (referred to as "original BPUs") to a prediction stage 204 to produce prediction data 206 and prediction BPUs 208. The encoder may subtract the prediction BPUs 208 from the original BPUs to generate residual BPUs 210. The encoder may feed the residual BPUs 210 to a transform stage 212 and a quantization stage 214 to produce quantized transform coefficients 216. The encoder may feed the prediction data 206 and the quantized transform coefficients 216 to a binary encoding stage 226 to produce a video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as a "forward path." During process 200A, after the quantization stage 214, the encoder may feed the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder may add the reconstruction residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224, which is used in the prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that both the encoder and the decoder use the same reference data for prediction.
[0057] The encoder may iteratively perform process 200A to encode each original BPU of the original image (in the forward path) and generate a prediction reference 224 for encoding the next original BPU of the original image (in the reconstruction path). After encoding all the original BPUs of the original image, the encoder may proceed to encode the next image in the video sequence 202.
[0058] Referring to process 200A, the encoder may receive a video sequence 202 generated by a video capture device (eg, a camera). As used herein, the term "receive" may refer to any action of receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or in any way inputting data.
[0059] In the prediction phase 204, in the current iteration, the encoder may receive the original BPU and the prediction reference 224 and perform a prediction operation to generate prediction data 206 and a predicted BPU 208. The prediction reference 224 may be generated from the reconstruction path in the previous iteration of process 200A. The purpose of the prediction phase 204 is to reduce information redundancy by extracting the prediction data 206 from the prediction data 206 and the prediction reference 224, which can be used to reconstruct the original BPU into the predicted BPU 208.
[0060] Ideally, the predicted BPU 208 would be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 typically differs slightly from the original BPU. To account for such differences, after generating the predicted BPU 208, the encoder may subtract the predicted BPU 208 from the original BPU to generate a residual BPU 210. For example, the encoder may subtract the pixel values (e.g., grayscale values or RGB values) of the predicted BPU 208 from the corresponding pixel values of the original BPU. Each pixel value of the residual BPU 210 may have a residual value that is the result of this subtraction between the corresponding pixel values of the original BPU and the predicted BPU 208. Compared to the original BPU, the predicted data 206 and the residual BPU 210 may have fewer bits, but they can be used to reconstruct the original BPU without noticeable quality degradation. Thus, the original BPU is compressed.
[0061] To further compress the residual BPU 210, during the transform stage 212, the encoder can reduce spatial redundancy in the residual BPU 210 by decomposing it into a set of two-dimensional "basis patterns," each of which is associated with a "transform coefficient." The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a frequency-varying component of the residual BPU 210 (e.g., the frequency of luminance variation). No basis pattern can be reproduced from any combination (e.g., linear combination) of any other basis patterns. In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. This decomposition is analogous to a discrete Fourier transform of a function, where the basis patterns are analogous to the basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are analogous to the coefficients associated with the basis functions.
[0062] Different transform algorithms can use different basis patterns. Various transform algorithms can be used in the transform stage 212, such as discrete cosine transform, discrete sine transform, and so on. The transform at the transform stage 212 is reversible. That is, the encoder can recover the residual BPU 210 by performing the inverse operation of the transform (referred to as an "inverse transform"). For example, to recover a pixel of the residual BPU 210, the inverse transform may be performed by multiplying the corresponding pixel values of the multiple basis patterns by their respective associated coefficients and adding the products to produce a weighted sum. For video coding standards, both the encoder and decoder can use the same transform algorithm (and therefore the same basis patterns). Therefore, the encoder can only record the transform coefficients, and the decoder can reconstruct the residual BPU 210 based on these transform coefficients without receiving the basis patterns from the encoder. Compared to the residual BPU 210, the transform coefficients may have fewer bits, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. As a result, the residual BPU 210 is further compressed.
[0063] The encoder can further compress the transform coefficients during the quantization stage 214. During the transform process, different basis patterns can represent different frequencies of variation (e.g., the frequency of brightness variations). Because the human eye is generally better at detecting low-frequency variations, the encoder can ignore information about high-frequency variations without significantly degrading decoding quality. For example, during the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (called a "quantization parameter") and rounding the quotient to the nearest integer. This operation can convert some transform coefficients of high-frequency basis patterns to zero, while the transform coefficients of low-frequency basis patterns can be converted to smaller integers. The encoder can further compress the transform coefficients by ignoring quantized transform coefficients 216 with zero values. The quantization process described above is also reversible, where the quantized transform coefficients 216 can be reconstructed as the transform coefficients in the inverse operation of quantization (called "inverse quantization").
[0064] Because the encoder ignores the remainder of this division during rounding operations, the quantization stage 214 can be lossy. Typically, the quantization stage 214 can cause the greatest information loss in process 200A. The greater the information loss, the fewer bits may be required to quantize the transform coefficients 216. To achieve different levels of information loss, the encoder can use different values for the quantization parameter or any other parameter in the quantization process.
[0065] In the binary encoding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary encoding technique, such as, for example, entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may also encode other information at the binary encoding stage 226, such as, for example, the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the transform type in the transform stage 212, the parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bit rate control parameters), etc. The encoder may use the output data of the binary encoding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 may be further packaged for network transmission.
[0066] Referring to the reconstruction path of process 200A, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients in an inverse quantization stage 218. In an inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224 to be used in the next iteration of process 200A.
[0067] It should be noted that other variations of process 200A may be used to encode video sequence 202. In some embodiments, the stages of process 200A may be performed by the encoder in a different order. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be divided into multiple stages. For example, transform stage 212 and quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may be omitted. Figure 2A one or more stages in a process.
[0068] Figure 2B A schematic diagram of another exemplary encoding process 200B consistent with embodiments of the present disclosure is shown. Process 200B can be modified from process 200A. For example, process 200B can be used by an encoder that complies with a hybrid video coding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B additionally includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.
[0069] Generally speaking, prediction techniques can be divided into two categories: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-image prediction or "intra-frame prediction") can use pixels from one or more coded adjacent BPUs in the same image to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include the adjacent BPUs. The spatial prediction can reduce the spatial redundancy inherent in the image. Temporal prediction (e.g., inter-image prediction or "inter-frame prediction") can use regions from one or more coded images to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include the coded images. The temporal prediction can reduce the temporal redundancy inherent in the image.
[0070] Referring to process 200B, in the forward path, the encoder performs the prediction operations in the spatial prediction stage 2042 and the temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra-frame prediction. For a certain original BPU of a certain picture being encoded, the prediction reference 224 may include one or more neighboring BPUs that have been encoded (in the forward path) and reconstructed (in the reconstruction path) in the same picture. The encoder may generate the predicted BPU 208 by extrapolating the neighboring BPUs. The extrapolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, the encoder may perform the extrapolation at the pixel level, for example, by extrapolating the corresponding pixel value of each pixel of the predicted BPU 208. The neighboring BPU used for extrapolation can be positioned in various directions relative to the original BPU, for example, vertically (e.g., on top of the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., to the lower left, lower right, upper left, or upper right of the original BPU), or any direction defined in the video coding standard used. For intra-frame prediction, the prediction data 206 may include, for example, the position (e.g., coordinates) of the neighboring BPU used, the size of the neighboring BPU used, the parameters of the extrapolation, the direction of the neighboring BPU used relative to the original BPU, etc.
[0071] As another example, during the temporal prediction stage 2044, the encoder may perform inter-frame prediction. For a particular original BPU of the current image, the prediction reference 224 may include one or more images (referred to as "reference images") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference images may be encoded and reconstructed on a BPU-by-BPU basis. For example, the encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. After all reconstructed BPUs for the same image are generated, the encoder may generate a reconstructed image as a reference image. The encoder may perform a "motion estimation" operation to search for a matching region within the reference image (referred to as a "search window"). The position of the search window in the reference image may be determined based on the position of the original BPU in the current image. For example, the search window may be centered at a location in the reference image with the same coordinates as the original BPU in the current image and may extend outward by a predetermined distance. When the encoder identifies an area similar to the original BPU in the search window (for example, by using a pixel recursion algorithm, a block matching algorithm, etc.), the encoder can determine such an area as a matching area. The matching area can have different specifications from the original BPU (for example, smaller than, equal to, larger than, or having a different shape). Since the reference image and the current image are temporally separated on the time axis, it can be considered that the matching area "moves" to the position of the original BPU over time. The encoder can record the direction and distance of this movement as a "motion vector". When multiple reference images are used, the encoder can search for matching areas for each reference image and determine its associated motion vector. In some embodiments, the encoder can assign weights to the pixel values of the matching areas of each matching reference image.
[0072] The motion estimation may be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter-frame prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, a motion vector associated with the matching region, the number of reference images, a plurality of weights associated with a plurality of the reference images, etc.
[0073] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., the motion vector) and the prediction reference 224. For example, the encoder may shift the matching region of the reference image according to the motion vector, thereby enabling the encoder to predict the original BPU of the current image. When multiple reference images are used, the encoder may shift the matching regions of the multiple reference images according to their respective motion vectors and average pixel values. In some embodiments, if the encoder has assigned weights to the pixel values of the matching regions of each matching reference image, the encoder may perform a weighted sum of the pixel values of the shifted matching regions.
[0074] In some embodiments, the inter-frame prediction can be unidirectional or bidirectional. Unidirectional inter-frame prediction can use one or more reference pictures in the same temporal direction relative to the current picture. Unidirectional inter-frame prediction uses a reference picture before the current picture. Bidirectional inter-frame prediction can use one or more reference pictures in both temporal directions relative to the current picture.
[0075] Still referring to the forward path of process 200B, after the spatial prediction 2042 and temporal prediction stages 2044, the encoder can select a prediction mode (e.g., intra prediction or inter prediction) for the current iteration of process 200B at the mode decision stage 230. For example, the encoder can perform a rate-distortion optimization technique, wherein the encoder can select a prediction mode based on the bit rate of a candidate prediction mode and the distortion of the reconstructed reference image under the candidate prediction mode to minimize the value of a cost function. Based on the selected prediction mode, the encoder can generate a corresponding prediction BPU 208 and prediction data 206.
[0076] In the reconstruction path of process 200B, if intra-prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU that has been encoded and reconstructed in the current picture), the encoder can directly feed the prediction reference 224 to the spatial prediction stage 2042 for subsequent use (e.g., for extrapolation of the next BPU of the current picture). If inter-prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current picture where all BPUs have been encoded and reconstructed), the encoder can feed the prediction reference 224 to the loop filtering stage 232, where the encoder can apply loop filtering to the prediction reference 224 to reduce or eliminate distortion (e.g., blocking artifacts) introduced by the inter-prediction. The encoder can apply various loop filtering techniques in the loop filtering stage 232, such as deblocking, sample adaptive offset, adaptive loop filtering, etc. The loop-filtered reference pictures may be stored in a buffer 234 (or "decoded picture buffer") for subsequent use (e.g., as inter-frame prediction reference pictures for future pictures in the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use in a temporal prediction stage 2044. In some embodiments, the encoder may encode loop filtering parameters (e.g., loop filter strength) as well as quantized transform coefficients 216, prediction data 206, and other information in a binary encoding stage 226.
[0077] Figure 3A FIG. 3 is a schematic diagram illustrating an exemplary decoding process 300A consistent with an embodiment of the present disclosure. The process 300A may be Figure 2A In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder (e.g., Figure 1 The image / video decoder 144 in FIG. 300A may decode the video bitstream 228 into a video stream 304 according to process 300A. The video stream 304 may be very similar to the video sequence 202. However, due to information loss during compression and decompression (e.g., Figures 2A-2B Typically, the video stream 304 is not identical to the video sequence 202. Figures 2A-2B 200A and 200B in the decoder, the decoder may perform process 300A at the basic processing unit (BPU) level for each picture encoded in the video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, where the decoder may decode one BPU in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for multiple regions of each picture encoded in the video bitstream 228.
[0078] exist Figure 3A In the process 300A, the decoder may feed a portion of the video bitstream 228 associated with a basic processing unit (referred to as an "encoded BPU") of an encoded picture to the binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may feed the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may feed the prediction data 206 to the prediction stage 204 to generate the prediction BPU 208. The decoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded picture buffer (DPB) in computer memory). The decoder may feed the prediction reference 224 to the prediction stage 204 for use in performing a prediction operation in the next iteration of the process 300A.
[0079] The decoder may iteratively perform process 300A to decode each coded BPU of the coded picture and generate a prediction reference 224 for encoding the next coded BPU of the coded picture. After decoding all coded BPUs of the coded picture, the decoder may output the picture to a video stream 304 for display and continue decoding the next coded picture in the video bitstream 228.
[0080] In the binary decoding stage 302, the decoder may perform the inverse of the binary coding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder may also decode other information in the binary decoding stage 302, such as, for example, prediction mode, parameters of the prediction operation, transform type, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), etc. In some embodiments, if the video bitstream 228 is transmitted in packets over the network, the decoder may depacketize the video bitstream 228 before feeding it to the binary decoding stage 302.
[0081] Figure 3BA schematic diagram of another exemplary decoding process 300B consistent with embodiments of the present disclosure is shown. Process 300B can be modified from process 300A. For example, process 300B can be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 300A, process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and also includes a loop filter stage 232 and a buffer 234.
[0082] In process 300B, for an encoded basic processing unit (referred to as a "current BPU") of an encoded picture being decoded (referred to as the "current picture"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 may include various types of data, depending on the prediction mode used by the encoder to encode the current BPU. For example, if the encoder used intra-frame prediction to encode the current BPU, the prediction data 206 may include a prediction mode identifier (e.g., a flag value) indicating the use of intra-frame prediction, parameters for the intra-frame prediction operation, and so on. The parameters for the intra-frame prediction operation may include, for example, the locations (e.g., coordinates) of one or more neighboring BPUs used as reference, the sizes of the neighboring BPUs, extrapolation parameters, the orientation of the neighboring BPUs relative to the original BPU, and so on. For another example, if the encoder used inter-frame prediction to encode the current BPU, the prediction data 206 may include a prediction mode identifier (e.g., a flag value) for the inter-frame prediction, parameters for the inter-frame prediction operation, and so on. The parameters of the inter-frame prediction operation may include, for example, the number of reference images associated with the current BPU, the weights associated with the reference images respectively, the positions (e.g., coordinates) of one or more matching regions in the respective reference images, one or more motion vectors associated with the matching regions respectively, and the like.
[0083] Based on the prediction mode identifier, the decoder can decide whether to perform spatial prediction (e.g., intra prediction) in a spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in a temporal prediction stage 2044. Details of performing such spatial prediction or temporal prediction are described in Figure 2B After performing such spatial prediction or temporal prediction, the decoder may generate a prediction BPU 208. The decoder may add the prediction BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, as shown in FIG. Figure 3A As shown in .
[0084] In process 300B, the decoder may feed the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing a prediction operation in the next iteration of process 300B. For example, if the current BPU is decoded using the intra prediction in the spatial prediction stage 2042, then after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder may feed the prediction reference 224 directly to the spatial prediction stage 2042 for subsequent use (e.g., for extrapolation of the next BPU of the current picture). If the current BPU is decoded using the inter prediction in the temporal prediction stage 2044, then after generating the prediction reference 224 (e.g., a reference picture in which all BPUs have been decoded), the encoder may feed the prediction reference 224 to the loop filtering stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may perform the following operations: Figure 2B In-loop filtering is applied to the prediction reference 224 in the manner described in
[15] . The loop-filtered reference picture may be stored in a buffer 234 (e.g., a decoded picture buffer in a computer memory) for subsequent use (e.g., as an inter-frame prediction reference picture for a subsequently encoded picture in the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, when the prediction mode identifier of the prediction data 206 indicates that inter-frame prediction was used to encode the current BPU, the prediction data may further include parameters of the loop filtering (e.g., loop filtering strength).
[0085] Return Reference Figure 1 The image / video preprocessor 122, the image / video encoder 124 and the image / video decoder 144 may be implemented using any suitable hardware, software or a combination thereof. Figure 4 is a block diagram of an example apparatus 400 for processing image data consistent with embodiments of the present disclosure. For example, apparatus 400 may be a preprocessor, an encoder, or a decoder. Figure 4As shown, the device 400 may include a processor 402. When the processor 402 executes the instructions described herein, the device 400 may become a special-purpose machine for pre-processing, encoding, or decoding image data. The processor 402 may be any type of circuit capable of manipulating or processing information. For example, the processor 402 may include any number of CPU processing units (or "CPUs"), graphics processing units (or "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property cores (IP cores), programmable logic arrays (PLAs), programmable array logic (PALs), general array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems on chip (SoCs), application-specific integrated circuits (ASICs), or other similar combinations. In some embodiments, the processor 402 may also be a group of processors grouped into a single logical component. For example, as Figure 4 As shown, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0086] The apparatus 400 may also include a memory 404 configured to store data (eg, instruction sets, computer code, intermediate data, etc.). Figure 4 As shown, the stored data may include program instructions (e.g., program instructions for implementing the stages in process 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). Processor 402 may access the program instructions and data for processing (e.g., via bus 410) and execute the program instructions to perform operations or manipulations on the data for processing. Memory 404 may include a high-speed random access memory device or a non-transitory storage device. In some embodiments, memory 404 may include any number of random access memories (RAMs), read-only memories (ROMs), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (CF) cards, etc., or any combination of the like. Memory 404 may also be a group of memories (e.g., multiple memories) grouped into a single logical component. Figure 4 not shown).
[0087] The bus 410 may be a communication device that transmits data between components within the apparatus 400 , such as an internal bus (eg, a CPU-memory bus), an external bus (eg, a Universal Serial Bus port, a Peripheral Component Interconnect Express port), and the like.
[0088] For ease of explanation and to avoid ambiguity, the processor 402 and other data processing circuitry are collectively referred to as "data processing circuitry" in this disclosure. The data processing circuitry may be implemented entirely as hardware, or as a combination of software, hardware, or firmware. Furthermore, the data processing circuitry may be a single standalone module or may be fully or partially integrated into any other component of the apparatus 400.
[0089] The apparatus 400 may further include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.). In some embodiments, the network interface 406 may include any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, or any combination thereof.
[0090] In some embodiments, the apparatus 400 may further include a peripheral interface 408 to provide a connection to one or more peripheral devices. Figure 4 As shown, the peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, touchpad, or touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or input interface coupled to a video archive), and the like.
[0091] It should be noted that a video codec (e.g., a codec that performs processes 200A, 200B, 300A, or 300B) can be implemented as any combination of software or hardware modules in apparatus 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more software modules of apparatus 400, such as program instructions that can be loaded into memory 404. For another example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of apparatus 400, such as dedicated data processing circuits (e.g., FPGAs, ASICs, NPUs, etc.).
[0092] Supplemental Enhancement Information (SEI) messages are intended to be transmitted within a coded video bitstream in a manner specified in a video coding specification, or by other means determined by a system specification using such a coded video bitstream. SEI messages can contain various types of data that indicate the timing of the video pictures or describe various properties of the coded video or how it may be used or enhanced. SEI messages are also defined to contain arbitrary user-defined data. SEI messages do not affect the core decoding process, but may indicate suggestions for post-processing or display of the video.
[0093] With the emergence of deep generative models, including variational autoencoders (VAEs) and generative adversarial networks (GANs), facial video compression can achieve significant performance improvements. For example, X2Face can be used to control facial generation through image, audio, and pose encoding. Furthermore, realistic neural talking head models can be applied through few-shot adversarial learning. For video-to-video synthesis tasks, Face-vid2vid can be used. Furthermore, schemes that use compact 3D keypoint representations to drive generative models to render the target frame can be adopted. Furthermore, mobile-compatible video chat systems based on FOMM can be used. VSBNet, which uses adversarial learning to reconstruct the original frame from landmarks, can also be used. Furthermore, an end-to-end talking head video compression framework based on compact feature learning (CFTE) can be used, which is designed for efficient talking face video compression in ultra-low bandwidth scenarios. The CFTE scheme uses the compact feature representation to compensate for temporal evolution and reconstruct the target facial video frame in an end-to-end manner. Furthermore, the CFTE scheme can be integrated into the video coding framework under the supervision of a rate-distortion objective. Although these algorithms achieve frame reconstruction with a small number of facial parameters through the powerful rendering capabilities of multiple deep generative models, some head pose motions and facial expression motions cannot be accurately rendered compared to the original motion videos.
[0094] Figure 5 is a schematic diagram of a deep learning-based video generation and compression framework 500 according to some embodiments of the present disclosure. The framework 500 is suitable for compressing and generating multiple speaking face videos. For example, the framework 500 can be based on a first-order motion model (FOMM). The FOMM deforms the reference source frame so that it follows the motion trajectory in the driving video. Although this method is applicable to various types of videos (e.g., dynamic images, cartoons), this method can also be used for facial animation applications. The FOMM adopts an encoder-decoder architecture and integrates a motion migration module, and includes the following steps.
[0095] First, a key point extractor (also known as a motion module) is trained using an equivariant loss without explicit annotations. Through this key point extractor, two sets of ten learned key points are calculated for the source frame and the driving frame. The learned key points are transformed from the feature map of channel dimension ×64×64 by the Gaussian mapping function, so each corresponding key point can represent the feature information of different channels. It should be noted that each key point is a point represented by the coordinates (x, y), which can represent the most important information of the feature map.
[0096] Second, a dense motion network uses the landmarks and the source frames to generate a dense motion field and an occlusion map.
[0097] Then, the encoder 510 encodes the source frame via the conventional image / video compression method such as HEVC / VVC or JPEG / BPG. In this embodiment, the source frame is compressed using VVC.
[0098] In a subsequent stage, the generated feature map is warped using the dense motion field (implemented using a differentiable grid sampling operation) and then multiplied with the occlusion map.
[0099] Finally, the decoder 520 generates an image from the warped map.
[0100] Figure 6 is a diagram illustrating an exemplary encoder-decoder encoding framework 600 with a compact feature size of 1×4×4 for speaking face videos, according to some embodiments of the present disclosure. Figure 6 Another basic framework for a deep video generation compression scheme based on compact feature representation, namely CFTE, is proposed. It follows an encoder-decoder architecture that applies a context-based coding scheme.
[0101] On the encoder 610 side, the compression framework includes three modules: an encoder for compressing the key frame (also called a VVC encoding module), a feature extractor for extracting compact human features of other inter-frames, and a feature encoding module for inter-frame prediction residuals of compressed compact human features. First, the key frame representing the human body texture is compressed using the VVC encoder. Through the compact feature extractor, each frame in the subsequent inter-frame frames is represented by a compact feature matrix of size 1×4×4. It should be noted that the size of the compact feature matrix is not fixed, and the number of feature parameters can also be increased or decreased according to the specific requirements of bit consumption. Then, these extracted features are inter-frame predicted and quantized, and the residual is finally entropy encoded into the final bit stream.
[0102] On the decoder 620 side, the compression framework also includes three main modules, including decoding for reconstructing the key frames, reconstructing the compact features through entropy decoding and compensation, and generating the final video by utilizing the reconstructed features and the decoded key frames. More specifically, in the process of generating the final video, the key frames decoded from the VVC bitstream can be further represented in the form of features through compact feature extraction. Subsequently, given the features from the key frames and inter-frames, the associated sparse motion fields are calculated, thereby facilitating the generation of pixel-by-pixel dense motion maps and occlusion maps. Finally, based on the deep generative model, the decoded key frames, pixel-by-pixel dense motion maps, and occlusion maps with implicit motion field features are used to produce the final video with accurate appearance, posture, and expression.
[0103] To further pursue coding performance, numerous studies have been conducted on 3D faces. A 3D head model was employed, and only the pose parameters were encoded for face-specific video compression tasks. Subsequently, feature space and principal component analysis (PCA) models were used for this task. However, based on these traditional 3D techniques, the visual quality of the reconstructed images was unacceptable. With the development of deep generative models, our proposed 3DMM-assisted (3D morphable model-assisted) facial video generation task can provide promising results.
[0104] Figure 7 FIG is a diagram illustrating a general encoder-decoder generation compression framework 700 for 3DMM-assisted speaking face videos according to some embodiments of the present disclosure. In general, 3DMM-assisted face video generation can provide accurate shape-based and texture The combined three-dimensional (3D) facial reconstruction is given by: in, and Denotes the average identity and texture, and the basis vectors of the identity, expression and texture space are denoted by B id 、B exp 、B tRepresented. The facial identity, expression and texture are represented by α, β and δ, which are the corresponding feature vectors that control the reconstructed face. In addition, the pose and position of the 3D face are controlled by the angle θ and the translation l. Therefore, on the encoder side (e.g., transmitter 710), the 3DMM parameters used as the 3D facial feature descriptor are compressed. In addition, the decoder (e.g., receiver 720) receives the bitstream to reconstruct a 3DMM template (e.g., 3D facial mesh, 3D facial landmarks, etc.). The 3D information reconstructed from the source image and the driving image is used as a guide to learn the optical flow required in the face reenactment synthesis process.
[0105] The existing SEI messages used in the current VVC standard were not designed to handle facial video compression. However, facial videos can be described by variations in feature structures with strong priors, such as landmarks or key points, and can even be parameterized as a series of facial semantic information to characterize the state of head pose and facial expression. For facial video compression, this compact facial semantics provides a large space for the syntax design and semantic description of SEI messages.
[0106] In addition, facial video communication has invoked more common use cases besides facial video reconstruction, such as facial video redirection or animation. For example, with the prevalence of metaverse activities, real-world facial movements may need to be transferred to the virtual metaverse world and represented by another person. Moreover, the reconstruction of the facial video is expected to be more realistic and redirected accordingly. Therefore, it is necessary to define a facial video compression SEI message, which may contain various types of data that are used to indicate the timing of the video images and / or describe various properties of the encoded video or how to use or enhance the encoded video. In this way, the post-processing or display of the reconstructed facial video can meet the actual needs of the user in a user-friendly manner.
[0107] To address the aforementioned issues, this disclosure proposes a new SEI message, called the Face Video Generation Compression SEI message. The proposed SEI has at least two functions: (1) reconstructing high-quality speaking face videos at an ultra-low bitrate, and (2) manipulating the speaking face videos into personalized representations. Therefore, the proposed SEI is suitable for video conferencing, real-time entertainment, and metaverse-related activities.
[0108] Figure 8 8 is a flow chart of an exemplary method 800 for processing a video by generating a compressed supplemental enhancement information (SEI) message based on a facial video according to some embodiments of the present disclosure. The method 800 describes the general syntax structure and syntax element order of the facial video generated compressed SEI message. Figure 8As shown, the method 800 may include the following steps 802 to 808 .
[0109] In step 802, an identification number is signaled to indicate whether a facial video generation compression scheme is used.
[0110] At step 804, when the identification number indicates that the facial video generation compression scheme is used, a plurality of (e.g., five) facial presence flags are further transmitted to indicate the presence and length of certain syntax elements associated with the facial video generation compression scheme. In some embodiments, the facial presence flags may include flags indicating head position, head rotation, head translation, eye blinking, and mouth movement.
[0111] In step 806, one or more corresponding facial information parameters are sent. In some embodiments, the corresponding facial information parameters may include head information parameters (e.g., head position parameters, head rotation parameters, head translation parameters, etc.), eye information parameters (e.g., blink parameters, etc.), and mouth information parameters (e.g., mouth motion parameters, etc.). In some embodiments, the head position corresponds to the facial information contained in the facial video. The head rotation further includes three degrees of freedom, such as roll, pitch, or yaw. The head translation may further include, for example, translation in three directions in a 3D coordinate system. The mouth motion parameters may further include a plurality of parameters indicating motion information of the mouth.
[0112] At step 808, the relevant facial parts are reconstructed or reoriented based on the base image using the corresponding facial information parameters sent.
[0113] Some exemplary embodiments regarding the proposed generation of compressed SEI messages for facial videos are described in detail below.
[0114] In some embodiments, multiple face presence flags and parameters are sent for all facial frames, wherein the flags are used to indicate the presence and length of certain syntax elements related to the facial video generation compression scheme, and the parameters indicate quantization factors used to process facial semantic parameters.
[0115] Figure 9 is a flow chart of an exemplary method 900 for processing a video by generating a compressed Supplemental Enhancement Information (SEI) message based on a plurality of facial videos, according to some embodiments of the present disclosure. Figure 10 1 shows an exemplary syntax for the disclosed facial video generation compression SEI message according to some embodiments of the present disclosure. Method 900 describes the general syntax structure and syntax element order of the facial video generation compression SEI message. Method 900 can be performed by an encoder (e.g., by Figure 2AProcess 200A or Figure 2B 200B) is performed by or by a device (e.g., Figure 4 For example, a processor (e.g., Figure 4 The method 900 may be performed by a processor 402 of the computer system. In some embodiments, the method 900 may be implemented by a computer program product contained in a computer-readable medium, the computer program product including a program executed by a computer (e.g., Figure 4 Computer executable instructions, such as program codes, executed by the device 400. Figure 9 and Figure 10 , method 900 may include the following steps 902 to 910.
[0116] In step 902, an identification number (e.g., fi_id) is sent to indicate whether to use the facial video generation compression scheme. Figure 10 1001 shown in .
[0117] In step 904, a parameter indicating the number of face frames is sent, such as fi_num_set_of_parameter, see Figure 10 1002 shown.
[0118] In step 906, a parameter is sent for all facial frames, the parameter indicating the quantization factor for processing facial semantic parameters, reference Figure 10 1003 shown in . The plurality of facial semantic parameters may include 14 facial semantic parameters (i.e., fi_location[i], fi_rotation_roll[i], fi_rotation_pitch[i], fi_rotation_yaw[i], fi_translation_x[i], fi_translation_y[i], fi_translation_z[i], fi_eye[i], fi_mouth_para1[i], fi_mouth_para2[i], fi_mouth_para3[i], fi_mouth_para4[i], fi_mouth_para5[i], and fi_mouth_para6[i]).
[0119] In step 908, a plurality of face presence flags are further sent for all face frames to indicate the presence and length of certain syntax elements associated with the face video generation compression scheme, referring to Figure 10 1004 shown in .
[0120] At step 910 , a plurality of corresponding facial information parameters are sent for each facial frame.
[0121] The following pairs Figure 9 and 10 Describe the consistent process.
[0122] First, the facial video generation compression includes an identification number that can be used to identify the facial video generation compression filter.
[0123] Second, each SEI message always has a base image included in the PU (i.e., the first face frame in the face sequence). The base image can provide a texture reference so that the face information parameters carried in the SEI message can be used to reconstruct the face frame.
[0124] Third, for these facial information parameters in the SEI message, five corresponding facial information parameter presence flags (i.e., fi_head_location_present_flag, fi_head_rotation_present_flag, fi_head_translation_present_flag, fi_eye_blinking_present_flag, and fi_mouth_motion_flag of sequence 1004) are sent to determine whether the relevant facial information parameters are sent. If these parameters are not sent, the corresponding facial information parameters from the base image are copied to generate multiple subsequent facial frames.
[0125] Fourth, when fi_head_location_present_flag is present, the head position parameters fi_location[i] shall be carried in the SEI message. When fi_head_rotation_present_flag is present, the head rotation parameters (fi_rotation_roll[i], fi_rotation_pitch[i], and fi_rotation_yaw[i]) shall be carried in the SEI message. When fi_head_translation_present_flag is present, the head translation parameters (fi_translation_x[i], fi_translation_y[i], and fi_translation_z[i]) shall be carried in the SEI message. When fi_eye_blinking_present_flag is present, the eye blink parameters fi_eye[i] shall be carried in the SEI message. When fi_mouth_motion_present_flag is present, these mouth motion parameters (fi_mouth_para1[i], fi_mouth_para2[i], fi_mouth_para3[i], fi_mouth_para4[i], fi_mouth_para5[i], and fi_mouth_para6[i]) should be carried in the SEI message. If the relevant flag is not present, the corresponding facial information parameters from the base image will be copied to generate the subsequent facial images, that is, the 14 facial parameters of these generated facial images remain the same as those of the base image.
[0126] Fifth, when these corresponding facial information parameters are carried in the SEI message, the facial video can be reconstructed in a personalized or user-friendly direction through the strong generation capability of the generative adversarial network.
[0127] and Figure 10 The semantics associated with the syntax in are described below.
[0128] fi_id contains an identification number that can be used to identify the compression filter used to generate the facial video. The value of fi_id can be between 0 and 2. 32 -2 (inclusive).
[0129] fi_num_set_of_parameter indicates the number of face frames that can be compressed using the facial video compression filter. The value of fi_num_set_of_parameter should be between 0 and 2. 10If exceeded, multiple facial frames can be encapsulated into the next SEI message.
[0130] fi_quantization_factor is the quantization factor used to process these 14 float16 type facial semantic parameters (i.e. fi_location[i], fi_rotation_roll[i], fi_rotation_pitch[i], fi_rotation_yaw[i], fi_translation_x[i], fi_translation_y[i], fi_translation_z[i], fi_eye[i], fi_mouth_para1[i], fi_mouth_para2[i], fi_mouth_para3[i], fi_mouth_para4[i], fi_mouth_para5[i] and fi_mouth_para6[i]). These float16 parameters can be further amplified by fi_quantization_factor, where the value of fi_quantization_factor can be between 0 and 10. 16 For example, the original facial parameter value is 0.1234567891234567, and the fi_quantization_factor value is 10. 6 , then the corresponding quantized facial parameters will be 123456.
[0131] When the head location exists, fi_head_location_present_flag is equal to 1, and when the head location does not exist, fi_head_location_present_flag is equal to 0.
[0132] When i of fi_location[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameters corresponding to the head position between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_location[0] specifies the quantized head position parameters from the 0th face frame (base image).
[0133] When head rotation exists, fi_head_rotation_present_flag is equal to 1, and when head rotation does not exist, fi_head_rotation_present_flag is equal to 0.
[0134] When i in fi_rotation_roll[i] is not equal to 0, fi_rotation_roll[i] specifies the quantized residual parameters corresponding to the front-to-back rotation of the head (called roll) between the i-th facial frame and the (i-1)-th facial frame by fi_quantization_factor. When i is equal to 0, fi_rotation_roll[0] specifies the quantized front-to-back head rotation parameters from the 0th facial frame (base image).
[0135] When i in fi_rotation_pitch[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameter (called pitch) of the head rotation around the left-right axis between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_rotation_pitch[0] specifies the quantized left-right head rotation parameter from the 0th face frame (base image).
[0136] When i in fi_rotation_yaw[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameter corresponding to the head rotation around the vertical axis (called head rotation, yaw) between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_rotation_yaw[0] specifies the quantized vertical axis head rotation parameter from the 0th face frame (base image).
[0137] When head translation is present, fi_head_translation_present_flag is equal to 1, and when head translation is not present, fi_head_translation_present_flag is equal to 0.
[0138] When i of fi_translation_x[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameter corresponding to the head translation around the x-axis between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_translation_x[0] specifies the quantized x-axis head translation parameter from the 0th face frame (base image).
[0139] When i of fi_translation_y[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameter corresponding to the head translation around the y-axis between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_translation_y[0] specifies the quantized y-axis head translation parameter from the 0th face frame (base image).
[0140] When i of fi_translation_z[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameter corresponding to the head translation around the z-axis between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_translation_z[0] specifies the quantized z-axis head translation parameter from the 0th face frame (base image).
[0141] When blinking is present, fi_eye_blinking_present_flag is equal to 1, and when blinking is not present, fi_eye_blinking_present_flag is equal to 0.
[0142] When i of fi_eye[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameter corresponding to the blink degree between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_eye[0] specifies the quantized blink parameter from the 0th face frame (base image).
[0143] When mouth motion is present, fi_mouth_motion_present_flag is equal to 1, and when mouth motion is not present, fi_mouth_motion_present_flag is equal to 0.
[0144] When i of fi_mouth_paral[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameters corresponding to the mouth motion between the i-th facial frame and the (i-1)-th facial frame. When i is equal to 0, fi_mouth_paral[0] specifies the quantized mouth motion parameters from the 0th facial frame (base image).
[0145] When i of fi_mouth_para2[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameters corresponding to the mouth motion between the i-th facial frame and the (i-1)-th facial frame. When i is equal to 0, fi_mouth_para2[0] specifies the quantized mouth motion parameters from the 0th facial frame (base image).
[0146] When i of fi_mouth_para3[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameters corresponding to the mouth motion between the i-th facial frame and the (i-1)-th facial frame. When i is equal to 0, fi_mouth_para3[0] specifies the quantized mouth motion parameters from the 0th facial frame (base image).
[0147] When i of fi_mouth_para4[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameters corresponding to the mouth motion between the i-th facial frame and the (i-1)-th facial frame. When i is equal to 0, fi_mouth_para4[0] specifies the quantized mouth motion parameters from the 0th facial frame (base image).
[0148] When i of fi_mouth_para5[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameters corresponding to the mouth motion between the i-th facial frame and the (i-1)-th facial frame. When i is equal to 0, fi_mouth_para5[0] specifies the quantized mouth motion parameters from the 0th facial frame (base image).
[0149] When i of fi_mouth_para6[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameters corresponding to the mouth motion between the i-th facial frame and the (i-1)-th facial frame. When i is equal to 0, fi_mouth_para6[0] specifies the quantized mouth motion parameters from the 0th facial frame (base image).
[0150] In some embodiments, for each facial frame, a plurality of face presence flag bits are sent respectively indicating the presence and length of certain syntax elements associated with the facial video generation compression scheme.
[0151] In addition, although the syntax definitions above and below specify what happens when the value is 0 or 1, it should be understood that these values are configurable and can be changed. For example, it can be modified so that when there is no head position information, fi_head_location_present_flag is equal to 1, and when the head position information is present, fi_head_location_present_flag is equal to 0.
[0152] Figure 11 is a flow chart of an exemplary method 1100 for processing video by generating compressed Supplemental Enhancement Information (SEI) messages based on facial video, according to some embodiments of the present disclosure. Figure 12 Another exemplary syntax of the disclosed facial video generation compression SEI message according to some embodiments of the present disclosure is shown. Method 1100 describes the general syntax structure and syntax element order of the facial video generation compression SEI message. Method 1100 can be performed by an encoder (e.g., by Figure 2A Process 200A or Figure 2B 200B) is performed by or by a device (e.g., Figure 4 For example, a processor (e.g., Figure 4 The method 1100 may be performed by a processor 402 of the computer system. In some embodiments, the method 1100 may be implemented by a computer program product contained in a computer-readable medium, the computer program product comprising a computer (e.g., Figure 4 Computer executable instructions, such as program codes, executed by the device 400. Figure 11 and Figure 12 , method 1100 may include the following steps 1102 to 1110.
[0153] In step 1102, an identification number (e.g., fi_id) is sent to indicate whether to use the facial video generation compression scheme. Figure 12 1201 shown in .
[0154] In step 1104, a parameter indicating the number of facial frames is sent, such as fi_num_set_of_parameter, see Figure 12 1202 shown.
[0155] In step 1106, a parameter is sent for all facial frames, the parameter indicating the quantization factor for processing facial semantic parameters, reference Figure 121203 shown in . The facial information parameters may include 14 facial information parameters (i.e., fi_location[i], fi_rotation_roll[i], fi_rotation_pitch[i], fi_rotation_yaw[i], fi_translation_x[i], fi_translation_y[i], fi_translation_z[i], fi_eye[i], fi_mouth_para1[i], fi_mouth_para2[i], fi_mouth_para3[i], fi_mouth_para4[i], fi_mouth_para5[i], and fi_mouth_para6[i]).
[0156] At step 1108, a plurality of face presence flags are further sent for each face frame to indicate the presence and length of certain syntax elements associated with the face video generation compression scheme, referring to Figure 12 1204 shown in .
[0157] In step 1110 , corresponding facial information parameters are sent for each facial frame.
[0158] The following pairs Figure 11 and 12 Describe the consistent process.
[0159] First, the facial video generation compression includes an identification number that can be used to identify the facial video generation compression filter.
[0160] Second, each SEI message always has a base image included in the PU (i.e., the first face frame in the face sequence). The base image can provide a texture reference so that the face information parameters carried in the SEI message can be used to reconstruct the face frame.
[0161] Third, for these facial information parameters in the SEI message, it is recommended to set 5 corresponding facial information presence flags (i.e., fi_head_location_present_flag[i], fi_head_rotation_present_flag[i], fi_head_translation_present_flag[i], fi_eye_blinking_present_flag[i], and fi_mouth_motion_present_flag[i]) to determine whether to send the relevant facial information parameters of each facial image [i].
[0162] Fourth, when fi_head_location_flag[i] is present, the head position parameter fi_location[i] can be carried in the SEI message. When fi_head_rotation_present_flag[i] is present, these head rotation parameters (fi_rotation_roll[i], fi_rotation_pitch[i], and fi_rotation_yaw[i]) should be carried in the SEI message. When fi_head_translation_present_flag[i] is present, these head translation parameters (fi_translation_x[i], fi_translation_y[i], and fi_translation_z[i]) should be carried in the SEI message. When fi_eye_blinking_flag[i] is present, the blink parameter fi_eye[i] should be carried in the SEI message. When fi_mouth_motion_present_flag[i] is present, these mouth motion parameters (fi_mouth_para1[i], fi_mouth_para2[i], fi_mouth_para3[i], fi_mouth_para4[i], fi_mouth_para5[i], and fi_mouth_para6[i]) should be carried in the SEI message. When any of the five facial information present flags is not present, there are two cases according to the present disclosure.
[0163] In some embodiments, the facial information parameters of the current facial frame (i.e., facial frame [i]) are copied from corresponding information from the previous facial frame (i.e., facial frame [i-1]). For the first facial frame ([i=0]), the facial information parameters are copied from the base image.
[0164] In some embodiments, the facial information parameters of the current facial frame (ie, facial frame [i]) are copied from the base image.
[0165] Fifth, when these corresponding facial information parameters are carried in the SEI message, the facial video can be reconstructed in a personalized or user-friendly direction through the strong generation capability of the generative adversarial network.
[0166] and Figure 12 The various semantics associated with the syntax are described below.
[0167] fi_id contains an identification number that can be used to identify the compression filter used to generate the facial video. The value of fi_id should be between 0 and 2. 32 -2 (inclusive).
[0168] fi_num_set_of_parameter indicates the number of face frames that can be compressed using the facial video compression filter. The value of fi_num_set_of_parameter should be between 0 and 2. 10 If it exceeds the range (including the end value), other facial frames need to be encapsulated into the next SEI message.
[0169] fi_quantization_factor is the quantization factor used to process these 14 float16 type facial semantic parameters (i.e. fi_location[i], fi_rotation_roll[i], fi_rotation_pitch[i], fi_rotation_yaw[i], fi_translation_x[i], fi_translation_y[i], fi_translation_z[i], fi_eye[i], fi_mouth_para1[i], fi_mouth_para2[i], fi_mouth_para3[i], fi_mouth_para4[i], fi_mouth_para5[i] and fi_mouth_para6[i]). These float16 parameters can be further amplified by fi_quantization_factor, where the value of fi_quantization_factor should be between 0 and 10. 16 For example, the original facial parameter value is 0.1234567891234567, and the fi_quantization_factor value is 10. 6 , then the corresponding quantized facial parameters will be 123456.
[0170] When the head location exists, fi_head_location_present_flag[i] is equal to 1, and when the head location does not exist, fi_head_location_present_flag is equal to 0.
[0171] When i of fi_location[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameters corresponding to the head position between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_location[0] specifies the quantized head position parameters from the 0th face frame (base image).
[0172] When head rotation exists, fi_head_rotation_present_flag[i] is equal to 1, and when head rotation does not exist, fi_head_rotation_present_flag is equal to 0.
[0173] When i in fi_rotation_roll[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameter corresponding to the front-to-back rotation of the head between the i-th face frame and the (i-1)-th face frame (called head roll). When i is equal to 0, fi_rotation_roll[0] specifies the quantized front-to-back head rotation parameter from the 0th face frame (base image).
[0174] When i in fi_rotation_pitch[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameter (called pitch) of the head rotation around the left-right axis between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_rotation_pitch[0] specifies the quantized left-right head rotation parameter from the 0th face frame (base image).
[0175] When i in fi_rotation_yaw[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameter corresponding to the head rotation around the vertical axis (called head rotation, yaw) between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_rotation_yaw[0] specifies the quantized vertical axis head rotation parameter from the 0th face frame (base image).
[0176] When head translation exists, fi_head_translation_present_flag[i] is equal to 1, and when head translation does not exist, fi_head_translation_present_flag is equal to 0.
[0177] When i of fi_translation_x[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameter corresponding to the head translation around the x-axis between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_translation_x[0] specifies the quantized x-axis head translation parameter from the 0th face frame (base image).
[0178] When i of fi_translation_y[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameter corresponding to the head translation around the y-axis between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_translation_y[0] specifies the quantized y-axis head translation parameter from the 0th face frame (base image).
[0179] When i of fi_translation_z[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameter corresponding to the head translation around the z-axis between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_translation_z[0] specifies the quantized z-axis head translation parameter from the 0th face frame (base image).
[0180] When blinking is present, fi_eye_blinking_present_flag[i] is equal to 1, and when blinking is not present, fi_eye_blinking_present_flag is equal to 0.
[0181] When i of fi_eye[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameter corresponding to the blink degree between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_eye[0] specifies the quantized blink parameter from the 0th face frame (base image).
[0182] When mouth motion is present, fi_mouth_motion_present_flag[i] is equal to 1, and when mouth motion is not present, fi_mouth_motion_present_flag is equal to 0.
[0183] When i of fi_mouth_paral[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameters corresponding to the mouth motion between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_mouth_paral[0] specifies the quantized mouth motion parameters from the 0th face frame (base image).
[0184] When i of fi_mouth_para2[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameters corresponding to the mouth motion between the i-th facial frame and the (i-1)-th facial frame. When i is equal to 0, fi_mouth_para2[0] specifies the quantized mouth motion parameters from the 0th facial frame (base image).
[0185] When i of fi_mouth_para3[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameters corresponding to the mouth motion between the i-th facial frame and the (i-1)-th facial frame. When i is equal to 0, fi_mouth_para3[0] specifies the quantized mouth motion parameters from the 0th facial frame (base image).
[0186] When i of fi_mouth_para4[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameters corresponding to the mouth motion between the i-th facial frame and the (i-1)-th facial frame. When i is equal to 0, fi_mouth_para4[0] specifies the quantized mouth motion parameters from the 0th facial frame (base image).
[0187] When i of fi_mouth_para5[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameters corresponding to the mouth motion between the i-th facial frame and the (i-1)-th facial frame. When i is equal to 0, fi_mouth_para5[0] specifies the quantized mouth motion parameters from the 0th facial frame (base image).
[0188] When i of fi_mouth_para6[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameters corresponding to the mouth motion between the i-th facial frame and the (i-1)-th facial frame. When i is equal to 0, fi_mouth_para6[0] specifies the quantized mouth motion parameters from the 0th facial frame (base image).
[0189] In some embodiments, for each facial frame, multiple facial presence flags and parameters are sent respectively, wherein the facial presence flags represent the presence and length of certain syntax elements associated with the facial video generation compression scheme, and the parameters represent quantization factors for processing facial semantic parameters.
[0190] Figure 13is a flow chart of an exemplary method 1300 for processing video by generating compressed supplemental enhancement information (SEI) messages based on facial video, according to some embodiments of the present disclosure. Figure 14 Another exemplary syntax of the disclosed facial video generation compression SEI message according to some embodiments of the present disclosure is shown. Method 1300 describes the general syntax structure and syntax element order of the facial video generation compression SEI message. Method 1300 can be performed by an encoder (e.g., by Figure 2A Process 200A or Figure 2B 200B) is performed by or by a device (e.g., Figure 4 For example, a processor (e.g., Figure 4 The method 1300 may be performed by a processor 402 of the computer system. In some embodiments, the method 1300 may be implemented by a computer program product contained in a computer-readable medium, the computer program product comprising a computer (e.g., Figure 4 Computer executable instructions, such as program codes, executed by the device 400. Figure 13 and Figure 14 , method 1300 may include the following steps 1302 to 1310.
[0191] In step 1302, an identification number (e.g., fi_id) is sent to indicate whether to use the facial video generation compression scheme. Figure 14 1401 shown in .
[0192] In step 1304, a parameter indicating the number of facial frames is sent, such as fi_num_set_of_parameter, see Figure 14 1402 shown.
[0193] In step 1306, a parameter is sent for each facial frame, wherein the parameter indicates a quantization factor for processing facial semantic parameters. Figure 14The facial semantic parameters may include 14 facial semantic parameters (i.e., fi_location[i], fi_rotation_roll[i], fi_rotation_pitch[i], fi_rotation_yaw[i], fi_translation_x[i], fi_translation_y[i], fi_translation_z[i], fi_eye[i], fi_mouth_para1[i], fi_mouth_para2[i], fi_mouth_para3[i], fi_mouth_para4[i], fi_mouth_para5[i], and fi_mouth_para6[i]).
[0194] In step 1308, a plurality of face presence flags are further sent for each face frame to indicate the presence and length of certain syntax elements associated with the face video generation compression scheme, referring to Figure 14 1404 shown in .
[0195] At step 1310 , corresponding facial information parameters are sent for each facial frame.
[0196] Figure 12 Table 2 in Figure 14 The difference between Table 3 in Figure 14 In the example, different fi_quantization_factors for the facial information parameters are sent for different facial images. Figure 13 and 14 A description of the consistent process.
[0197] First, the facial video generation compression includes an identification number that can be used to identify the facial video generation compression filter.
[0198] Second, each SEI message always has a base image included in the PU (i.e., the first face frame in the face sequence). The base image can provide a rich texture reference so that the face information parameters carried in the SEI message can be used to reconstruct the face frame.
[0199] Third, for these facial information parameters in the SEI message, it is recommended to set 5 corresponding facial information presence flags (i.e., fi_head_location_present_flag[i], fi_head_rotation_present_flag[i], fi_head_translation_present_flag[i], fi_eye_blinking_present_flag[i], and fi_mouth_motion_present_flag[i]) to determine whether to send the relevant facial information parameters of each facial image [i].
[0200] Fourth, when fi_head_location_present_flag[i] is present, the head position parameter fi_location[i] shall be carried in the SEI message. When fi_head_rotation_present_flag[i] is present, the head rotation parameters (fi_rotation_roll[i], fi_rotation_pitch[i], and fi_rotation_yaw[i]) shall be carried in the SEI message. When fi_head_translation_present_flag[i] is present, the head translation parameters (fi_translation_x[i], fi_translation_y[i], and fi_translation_z[i]) shall be carried in the SEI message. When fi_eye_blinking_flag[i] is present, the eye blink parameter fi_eye[i] shall be carried in the SEI message. When fi_mouth_motion_present_flag[i] is present, these mouth motion parameters (fi_mouth_para1[i], fi_mouth_para2[i], fi_mouth_para3[i], fi_mouth_para4[i], fi_mouth_para5[i], and fi_mouth_para6[i]) should be carried in the SEI message. When any of the five facial information present flags is not present, there are two cases according to the present disclosure.
[0201] In some embodiments, the facial information parameters of the current facial frame (i.e., facial frame [i]) are copied from corresponding information from the previous facial frame (i.e., facial frame [i-1]). For the first facial frame ([i=0]), the facial information parameters are copied from the base image.
[0202] In some embodiments, the facial information parameters of the current facial frame (ie, facial frame [i]) are copied from the base image.
[0203] Fifth, when these corresponding facial information parameters are carried in the SEI message, the facial video can be reconstructed in a personalized or user-friendly direction through the strong generation capability of the generative adversarial network.
[0204] and Figure 14 The various semantics associated with the syntax are described below.
[0205] fi_id contains an identification number that can be used to identify the compression filter used to generate the facial video. The value of fi_id should be between 0 and 2. 32 -2 (inclusive).
[0206] fi_num_set_of_parameter indicates the number of face frames that can be compressed using the facial video compression filter. The value of fi_num_set_of_parameter should be between 0 and 2. 10 If it exceeds the range, other facial frames need to be encapsulated into the next SEI message.
[0207] fi_quantization_factor[i] is the quantization factor used to process these 14 float16 type facial semantic parameters (i.e. fi_location[i], fi_rotation_roll[i], fi_rotation_pitch[i], fi_rotation_yaw[i], fi_translation_x[i], fi_translation_y[i], fi_translation_z[i], fi_eye[i], fi_mouth_para1[i], fi_mouth_para2[i], fi_mouth_para3[i], fi_mouth_para4[i], fi_mouth_para5[i] and fi_mouth_para6[i]). These float16 parameters can be further amplified by fi_quantization_factor, where the value of fi_quantization_factor[i] should be between 0 and 10. 16 For example, the original facial parameter value is 0.1234567891234567, and the fi_quantization_factor value is 10. 6 , then the corresponding quantized facial information parameters will be 123456.
[0208] When the head location exists, fi_head_location_present_flag[i] is equal to 1, and when the head location does not exist, fi_head_location_present_flag is equal to 0.
[0209] When i of fi_location[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameters corresponding to the head position between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_location[0] specifies the quantized head position parameters from the 0th face frame (base image).
[0210] When head rotation exists, fi_head_rotation_present_flag[i] is equal to 1, and when head rotation does not exist, fi_head_rotation_present_flag is equal to 0.
[0211] When i in fi_rotation_roll[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameter corresponding to the front-to-back rotation of the head between the i-th face frame and the (i-1)-th face frame (called head roll). When i is equal to 0, fi_rotation_roll[0] specifies the quantized front-to-back head rotation parameter from the 0th face frame (base image).
[0212] When i in fi_rotation_pitch[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameter (called pitch) of the head rotation around the left-right axis between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_rotation_pitch[0] specifies the quantized left-right head rotation parameter from the 0th face frame (base image).
[0213] When i in fi_rotation_yaw[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameter corresponding to the head rotation around the vertical axis (called head rotation, yaw) between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_rotation_yaw[0] specifies the quantized vertical axis head rotation parameter from the 0th face frame (base image).
[0214] When head translation exists, fi_head_translation_present_flag[i] is equal to 1, and when head translation does not exist, fi_head_translation_present_flag is equal to 0.
[0215] When i of fi_translation_x[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameter corresponding to the head translation around the x-axis between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_translation_x[0] specifies the quantized x-axis head translation parameter from the 0th face frame (base image).
[0216] When i of fi_translation_y[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameter corresponding to the head translation around the y-axis between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_translation_y[0] specifies the quantized y-axis head translation parameter from the 0th face frame (base image).
[0217] When i of fi_translation_z[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameter corresponding to the head translation around the z-axis between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_translation_z[0] specifies the quantized z-axis head translation parameter from the 0th face frame (base image).
[0218] When blinking is present, fi_eye_blinking_present_flag[i] is equal to 1, and when blinking is not present, fi_eye_blinking_present_flag is equal to 0.
[0219] When i of fi_eye[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameter corresponding to the blink degree between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_eye[0] specifies the quantized blink parameter from the 0th face frame (base image).
[0220] When mouth motion is present, fi_mouth_motion_present_flag[i] is equal to 1, and when mouth motion is not present, fi_mouth_motion_present_flag is equal to 0.
[0221] When i of fi_mouth_paral[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameters corresponding to the mouth motion between the i-th face frame and the (i-1)-th face frame. When i is equal to 0, fi_mouth_paral[0] specifies the quantized mouth motion parameters from the 0th face frame (base image).
[0222] When i of fi_mouth_para2[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameters corresponding to the mouth motion between the i-th facial frame and the (i-1)-th facial frame. When i is equal to 0, fi_mouth_para2[0] specifies the quantized mouth motion parameters from the 0th facial frame (base image).
[0223] When i of fi_mouth_para3[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameters corresponding to the mouth motion between the i-th facial frame and the (i-1)-th facial frame. When i is equal to 0, fi_mouth_para3[0] specifies the quantized mouth motion parameters from the 0th facial frame (base image).
[0224] When i of fi_mouth_para4[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameters corresponding to the mouth motion between the i-th facial frame and the (i-1)-th facial frame. When i is equal to 0, fi_mouth_para4[0] specifies the quantized mouth motion parameters from the 0th facial frame (base image).
[0225] When i of fi_mouth_para5[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameters corresponding to the mouth motion between the i-th facial frame and the (i-1)-th facial frame. When i is equal to 0, fi_mouth_para5[0] specifies the quantized mouth motion parameters from the 0th facial frame (base image).
[0226] When i of fi_mouth_para6[i] is not equal to 0, fi_quantization_factor specifies the quantized residual parameters corresponding to the mouth motion between the i-th facial frame and the (i-1)-th facial frame. When i is equal to 0, fi_mouth_para6[0] specifies the quantized mouth motion parameters from the 0th facial frame (base image).
[0227] Figure 15 1 is a flow chart of an exemplary method 1500 for processing a video by generating a compressed supplemental enhancement information (SEI) message based on a facial video according to some embodiments of the present disclosure. The method 1500 describes the general syntax structure and syntax element order of the facial video generated compressed SEI message. The method 1500 may be performed by a decoder (e.g., by Figure 3A Process 300A or Figure 3B 300B) or by a device (e.g., Figure 4 For example, a processor (e.g., Figure 4 The method 1500 may be performed by a processor 402 of the computer readable medium. In some embodiments, the method 1500 may be implemented by a computer program product contained in a computer-readable medium, the computer program product including a program executed by a computer (e.g., Figure 4 Computer executable instructions, such as program codes, executed by the device 400. Figure 15 , method 1500 may include the following steps 1502 to 1506.
[0228] In step 1502, it is determined whether a facial video generation compression scheme is used based on an identification number.
[0229] In response to the determination that the facial video generation compression scheme is used, a supplemental enhancement information (SEI) message is decoded at step 1504. The SEI message includes facial information, such as Figure 10 、 Figure 12 or Figure 14 The facial information shown.
[0230] In some embodiments, decoding the SEI message further includes decoding a flag indicating whether a syntax element associated with facial information is present in the SEI message, and, in response to the presence of the syntax element associated with the facial information, decoding the facial information based on the syntax element associated with the facial information.
[0231] In some embodiments, decoding the SEI message further comprises decoding a syntax element indicating a number of facial frames generated using the facial video compression scheme, and decoding, for each facial frame, a flag indicating whether a syntax element associated with facial information is present in the SEI message.
[0232] In some embodiments, decoding the SEI message further includes: decoding a syntax element, the syntax element indicating the number of facial frames using the facial video generation compression scheme, and respectively decoding a plurality of facial presence flags for each facial frame, the facial presence flags indicating the presence and length of certain syntax elements associated with the facial video generation compression scheme.
[0233] In some embodiments, decoding the SEI message further includes: decoding a factor for each facial frame respectively, where the factor is used to indicate a quantization factor for processing facial information.
[0234] At step 1506 , a facial image is reconstructed based on the facial information and a base image associated with the SEI message.
[0235] In some embodiments, the facial information of the current facial frame (i.e., facial frame [i]) is copied from information corresponding to the previous facial frame (i.e., facial frame [i-1]). For the first facial frame ([i=0]), the facial information is copied from the base image.
[0236] In some embodiments, the facial information of the current facial frame (ie, facial frame [i]) is copied from the base image.
[0237] In some embodiments, a decoding device is provided, comprising: a receiving module configured to receive a bit stream; and The decoding module is configured to decode one or more images using the encoding information of the bit stream. The decoding module is configured to: determining whether a facial video generation compression scheme is used based on an identification number; Responsive to a determination that the facial video generation compression scheme was used, decoding a supplemental enhancement information (SEI) message, the SEI message including facial information; and A facial image is reconstructed based on the facial information and a base image associated with the SEI message.
[0238] In one embodiment, the SEI message further includes: a flag bit, the flag bit indicating whether there is a syntax element associated with facial information in the SEI message, and the decoding module is configured to: In response to the presence of the syntax element associated with facial information, the facial information is decoded based on the syntax element associated with the facial information.
[0239] In one embodiment, the SEI message further includes: a syntax element indicating the number of facial frames generated using the facial video compression scheme, and the decoding module is configured to: The flag bit of each facial frame is decoded respectively, where the flag bit indicates whether a syntax element associated with facial information exists in the SEI message.
[0240] In one embodiment, the SEI message further includes: a syntax element indicating the number of facial frames generated using the facial video compression scheme, and the decoding module is configured to: The facial information of each facial frame is decoded based on the syntax elements associated with the facial information.
[0241] In one embodiment, the decoding device further comprises: The copy module is configured to copy corresponding facial information from a previous facial frame in response to a flag indicating that a syntax element associated with the facial information does not exist in the current facial frame.
[0242] In one embodiment, the SEI message further includes: a base picture as a reference, and configuring the replication module to: In response to a flag indicating that syntax elements associated with facial information are absent from the first facial frame, corresponding facial information is copied from the base image.
[0243] In one implementation, the SEI message further includes: a base image as a reference, and the decoding apparatus further includes: The copy module is configured to copy corresponding facial information from the base image in response to a flag indicating that a syntax element associated with facial information does not exist in the current facial frame.
[0244] In one embodiment, the SEI message further includes: a factor indicating a quantization factor for processing the facial information; and configuring the decoding module to: decoding the factors; and The decoding device further comprises: A processing module is configured to process the facial information based on the factors.
[0245] In one implementation, the SEI message further includes: a syntax element indicating the number of facial frames generated using the facial video compression scheme; and configuring the decoding module to: Decoding the factors of each facial frame separately; and The processing module is configured to: For each facial frame, the facial information is processed based on the factors.
[0246] In some embodiments, an encoding device is provided, comprising: A receiving module configured to receive a video sequence; an encoding module configured to encode one or more images of the video sequence; and a generating module configured to generate a bitstream; The encoding module is configured to: An identification number is sent, the identification number indicating whether a facial video generation compression scheme is used.
[0247] In one embodiment, when using the facial video generation compression scheme, the encoding module is configured to: Sending a flag indicating whether a syntax element associated with the facial information exists; and When the flag indicates that the syntax element associated with the facial information exists, the corresponding syntax element associated with the facial information is sent.
[0248] In one implementation, the encoding module is configured to: sending a syntax element indicating a number of facial frames generated using the facial video compression scheme; and For each facial frame, a flag is sent, indicating whether there is a syntax element related to facial information.
[0249] In one implementation, the encoding module is configured to: sending a syntax element indicating a number of facial frames generated using the facial video compression scheme; and For each face frame, a factor is sent, indicating a quantization factor for processing the face information.
[0250] In some embodiments, an electronic device is provided, comprising: a memory storing an instruction set; and one or more processors configured to execute the instruction set so that the one or more processors perform a method of decoding a bit stream to output one or more images of a video stream according to the multiple method embodiments of decoding a bit stream to output one or more images of a video stream.
[0251] In some embodiments, an electronic device is provided, comprising: a memory storing an instruction set; and one or more processors configured to execute the instruction set so that the one or more processors perform a method for encoding a video sequence into a bit stream according to the above-mentioned multiple method embodiments for encoding a video sequence into a bit stream.
[0252] In some embodiments, a non-transitory computer-readable storage medium is provided that stores a video bitstream. When the bitstream is decoded by a processor, the processor is caused to execute the method for decoding a bitstream to output one or more images of a video stream according to the above-mentioned multiple method embodiments for decoding a bitstream to output one or more images of a video stream.
[0253] In some embodiments, a non-transitory computer-readable storage medium is provided, storing a video sequence of a video. When the video sequence is encoded by a processor, the processor is caused to execute the method for encoding the video sequence into a bitstream according to the above-mentioned multiple method embodiments for encoding the video sequence into a bitstream.
[0254] In some embodiments, a computer program product is provided, comprising: a plurality of computer program instructions, wherein the plurality of computer program instructions enable a computer to execute a method of decoding a bitstream to output one or more images of a video stream according to the plurality of method embodiments of decoding a bitstream to output one or more images of a video stream.
[0255] In some embodiments, a computer program product is provided, comprising: a plurality of computer program instructions, wherein the plurality of computer program instructions enable a computer to execute a method for encoding a video sequence into a bit stream according to the plurality of method embodiments for encoding a video sequence into a bit stream described above.
[0256] In some embodiments, a computer program is provided, and the computer program enables a computer to perform the method of decoding a bitstream to output one or more images of a video stream according to the multiple method embodiments of decoding a bitstream to output one or more images of a video stream.
[0257] In some embodiments, a computer program is provided, and the computer program enables a computer to execute the method of encoding a video sequence into a bitstream according to the above embodiments of the method of encoding a video sequence into a bitstream.
[0258] The embodiments may be further described using the following terms: 1. A method for decoding a bitstream to output one or more pictures of a video stream, the method comprising: receiving a bit stream; and decoding one or more images using the encoding information of the bitstream; The decoding includes: determining whether a facial video generation compression scheme is used based on an identification number; In response to a determination that the facial video generation compression scheme was used, decoding a supplemental enhancement information (SEI) message, the SEI message including facial information; and A facial image is reconstructed based on the facial information and a base image associated with the SEI message. 2. A method as described in clause 1, wherein the SEI message further comprises: a flag bit indicating whether syntax elements associated with facial information are present in the SEI message, and the method further comprising: In response to the presence of the syntax element associated with facial information, the facial information is decoded based on the syntax element associated with the facial information. 3. The method of clause 2, wherein the SEI message further comprises: a syntax element indicating a number of facial frames generated using the facial video compression scheme, and the method further comprising: The flag indicating whether a syntax element associated with facial information exists in the SEI message is decoded for each facial frame respectively. 4. The method of clause 2, wherein the SEI message further comprises: a syntax element indicating a number of facial frames generated using the facial video compression scheme, and the method further comprising: The facial information of each facial frame is decoded based on the syntax elements associated with the facial information. 5. The method according to clause 4, further comprising: In response to a flag indicating that a syntax element associated with facial information is absent in the current facial frame, corresponding facial information is copied from a previous facial frame. 6. The method of clause 5, wherein the SEI message further includes a base picture as a reference, and the method further comprises: In response to a flag indicating that syntax elements associated with facial information are absent from the first facial frame, corresponding facial information is copied from the base image. 7. The method of clause 4, wherein the SEI message further includes a base picture as a reference, and the method further comprises: In response to a flag indicating that a syntax element associated with facial information does not exist in the current facial frame, corresponding facial information is copied from the base image. 8. A method according to any of clauses 1 to 7, wherein the SEI message further comprises: a factor indicating a quantization factor for processing the facial information, and the method further comprising: decoding the factors; and The facial information is processed based on the factors. 9. The method of clause 8, wherein the SEI message further comprises: a syntax element indicating a number of facial frames generated using the facial video compression scheme, and the method further comprising: Decoding the factors of each facial frame separately; and For each facial frame, the facial information is processed based on the factors. 10. A method for encoding a video sequence into a bitstream, the method comprising: receiving a video sequence; encoding one or more images of the video sequence; and Generate a bit stream; The encoding includes: An identification number is sent, the identification number indicating whether a facial video generation compression scheme is used. 11. The method of clause 10, wherein, when the facial video generation compression scheme is used, the method further comprises: Sending a flag indicating whether a syntax element associated with the facial information exists; and When the flag indicates that the syntax element associated with the facial information exists, the corresponding syntax element associated with the facial information is sent. 12. The method according to clause 11, further comprising: sending a syntax element indicating a number of facial frames generated using the facial video compression scheme; and For each facial frame, a flag is sent, indicating whether there is a syntax element related to facial information. 13. The method according to clause 11, further comprising: sending a syntax element indicating a number of facial frames generated using the facial video compression scheme; and For each face frame, a factor is sent, indicating a quantization factor for processing the face information. 14. A non-transitory computer-readable storage medium storing a bitstream of a video, the bitstream comprising: a Supplemental Enhancement Information (SEI) message, the SEI message including facial information, Wherein, the facial image is reconstructed using the facial information based on a base image associated with the facial image. 15. The non-transitory computer-readable storage medium of clause 14, wherein the SEI message further comprises: a flag bit indicating whether a syntax element associated with facial information is present in the SEI message. 16. The non-transitory computer-readable storage medium of clause 15, wherein the SEI message further comprises: a syntax element indicating a number of facial frames generated using a facial video compression scheme. 17. The non-transitory computer-readable storage medium of clause 16, wherein the SEI message further comprises one or more syntax elements respectively associated with facial information of each facial frame. 18. The non-transitory computer-readable storage medium of clause 14, wherein the SEI message further comprises: a syntax element indicating a number of facial frames generated using a facial video compression scheme, and a flag bit indicating, for each facial frame, whether a syntax element associated with facial information is present in the SEI message. 19. The non-transitory computer-readable storage medium of clause 14, wherein the SEI message further comprises: a factor indicating a quantization factor for processing the facial information. 20. The non-transitory computer-readable storage medium of clause 14, wherein the SEI message further comprises: a syntax element indicating a number of facial frames to be generated using a facial video compression scheme, and a factor indicating a quantization factor to process the facial information of each frame.
[0259] In some embodiments, a non-transitory computer-readable storage medium comprising a plurality of instructions is further provided, and the plurality of instructions can be executed by a device for performing the above method (e.g., the disclosed encoder and decoder). In some embodiments, a non-transitory computer-readable storage medium storing a bitstream or SEI message is further provided. A compressed supplemental enhancement information (SEI) message (e.g., Figure 10 、 Figure 12 and Figure 14) encodes and decodes the bit stream. Common forms of non-transitory media include, for example, floppy disks, flexible magnetic disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage medium, CD-ROMs, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROM and EPROM, FLASH-EPROM or any other flash memory, NVRAM, cache, registers, any other memory chip or cassette and networked versions thereof. The device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memories.
[0260] It should be noted that relational terms such as "first" and "second" herein are used only to distinguish one entity or operation from another entity or operation and do not require or imply any actual relationship or order between these entities or operations. In addition, the words "comprise," "have," "contain," and "include" and other similar forms are intended to be synonymous and open-ended, in that one or more items following any of these words are not intended to be an exhaustive list of such one or more items or to be limited to the listed one or more items.
[0261] As used herein, unless otherwise specifically stated, the term "or" encompasses all possible combinations unless not feasible. For example, if a database is specified to include either A or B, then unless otherwise specifically stated or not feasible, the database may include A, or B, or A and B. As a second example, if a database is specified to include either A, B, or C, then unless otherwise specifically stated or not feasible, the database may include A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C.
[0262] It should be understood that the above embodiments can be implemented by hardware, or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-mentioned computer-readable medium. When executed by a processor, the software can execute the disclosed method. The computer unit and other functional units described in the present disclosure can be implemented by hardware, or software, or a combination of hardware and software. Those of ordinary skill in the art will also understand that the above-mentioned multiple modules / units can be combined into one module / unit, and each of the above-mentioned modules / units can be further divided into multiple sub-modules / sub-units.
[0263] In the foregoing description, embodiments have been described with reference to many specific details, which may vary depending on the implementation. Certain adaptations and modifications may be made to the described embodiments. Other embodiments will be apparent to those skilled in the art by considering the description and practice disclosed herein. The description and examples are intended to be considered merely exemplary, with the true scope and nature of the present disclosure being indicated by the appended claims. The order of steps shown in the figures is for illustrative purposes only and is not intended to be limited to any particular order of steps. Therefore, it will be understood by those skilled in the art that these steps may be performed in different orders while achieving the same method.
[0264] In the drawings and the specification, exemplary embodiments have been disclosed. However, many variations and modifications may be made to these embodiments. Accordingly, although specific terms are employed, they are used in a general and descriptive sense only and not for the purpose of limitation.
Claims
1. A method for decoding a bitstream to output one or more images of a video stream, the method comprising: receiving a bit stream; as well as decoding one or more images using the encoding information of the bitstream; The decoding includes: determining whether a facial video generation compression scheme is used based on an identification number; In response to a determination that the facial video generation compression scheme was used, decoding a supplemental enhancement information SIE message, the SIE message including facial information; and A facial image is reconstructed based on the facial information and a base image associated with the SE I message.
2. The method according to claim 1, wherein The SE I message further includes: a flag bit, the flag bit indicating whether there is a syntax element associated with facial information in the SE I message, and the method further includes: In response to the presence of the syntax element associated with facial information, the facial information is decoded based on the syntax element associated with the facial information.
3. The method according to claim 2, wherein: The SIE message further includes: a syntax element indicating a number of facial frames generated using the facial video compression scheme, and the method further includes: The flag indicating whether a syntax element associated with facial information exists in the SEl message is decoded for each facial frame respectively.
4. The method according to claim 2, wherein The SIE message further includes: a syntax element indicating a number of facial frames generated using the facial video compression scheme, and the method further includes: The facial information of each facial frame is decoded based on the syntax elements associated with the facial information.
5. The method according to claim 4, further comprising: In response to a flag indicating that a syntax element associated with facial information is absent in the current facial frame, corresponding facial information is copied from a previous facial frame.
6. The method according to claim 5, wherein: The SE I message further includes: a base image as a reference, and the method further includes: In response to a flag indicating that syntax elements associated with facial information are absent from the first facial frame, corresponding facial information is copied from the base image.
7. The method according to claim 4, wherein: The SE I message further includes: a base image as a reference, and the method further includes: In response to a flag indicating that a syntax element associated with facial information does not exist in the current facial frame, corresponding facial information is copied from the base image.
8. The method according to claim 1 or 2, wherein: The SE I message further includes: a factor indicating a quantization factor for processing the facial information, and the method further includes: decoding the factors; and The facial information is processed based on the factors.
9. The method according to claim 8, wherein The SIE message further includes: a syntax element indicating a number of facial frames generated using the facial video compression scheme, and the method further includes: Decoding the factors of each facial frame separately; and For each facial frame, the facial information is processed based on the factors.
10. A method of encoding a video sequence into a bitstream, the method comprising: receiving a video sequence; encoding one or more images of the video sequence; as well as Generate a bit stream; The encoding includes: An identification number is sent, the identification number indicating whether a facial video generation compression scheme is used.
11. The method according to claim 10, wherein: When the facial video generation compression scheme is used, the method further comprises: Sending a flag indicating whether a syntax element associated with the facial information exists; and When the flag indicates that the syntax element associated with the facial information exists, the corresponding syntax element associated with the facial information is sent.
12. The method according to claim 11, further comprising: sending a syntax element indicating a number of facial frames generated using the facial video compression scheme; as well as For each facial frame, a flag is sent, indicating whether there is a syntax element associated with the facial information.
13. The method according to claim 11, further comprising: sending a syntax element indicating a number of facial frames generated using the facial video compression scheme; as well as For each face frame, a factor is sent, indicating a quantization factor for processing the face information.
14. A decoding device comprising: a receiving module configured to receive a bit stream; as well as The decoding module is configured to decode one or more images using the encoding information of the bit stream. The decoding module is configured to: determining whether a facial video generation compression scheme is used based on an identification number; In response to a determination that the facial video generation compression scheme was used, decoding a supplemental enhancement information (SE I) message, the SE I message including facial information; and A facial image is reconstructed based on the facial information and a base image associated with the SE I message.
15. The decoding device according to claim 14, wherein: The SE I message further includes: a flag bit, the flag bit indicating whether a syntax element associated with facial information is present in the SE I message; and configuring the decoding module to: In response to the presence of the syntax element associated with facial information, the facial information is decoded based on the syntax element associated with the facial information.
16. The decoding device according to claim 15, wherein: The SE I message further includes: a syntax element indicating the number of facial frames generated using the facial video compression scheme, and configuring the decoding module to: The flag bit of each facial frame is decoded respectively, where the flag bit indicates whether there is a syntax element associated with facial information in the SIE message.
17. The decoding device according to claim 15, wherein: The SE I message further includes: a syntax element indicating the number of facial frames generated using the facial video compression scheme, and configuring the decoding module to: The facial information of each facial frame is decoded based on the syntax elements associated with the facial information.
18. The decoding device according to claim 17, further comprising: The copy module is configured to copy corresponding facial information from a previous facial frame in response to a flag indicating that a syntax element associated with the facial information does not exist in the current facial frame.
19. The decoding device according to claim 18, wherein: The SE I message further includes: a base image as a reference, and configuring the replication module to: In response to a flag indicating that syntax elements associated with facial information are absent from the first facial frame, corresponding facial information is copied from the base image.
20. The decoding device according to claim 17, wherein: The SE I message further includes: a base image as a reference, and the decoding device further includes: The copy module is configured to copy corresponding facial information from the base image in response to a flag indicating that a syntax element associated with facial information does not exist in the current facial frame.
21. The decoding device according to claim 14 or 15, wherein: The SE I message further includes: a factor indicating a quantization factor for processing the facial information; and configuring the decoding module to: decoding the factors; and The decoding device further comprises: A processing module is configured to process the facial information based on the factors.
22. The decoding device according to claim 21, wherein: The SE I message further includes: a syntax element indicating the number of facial frames generated using the facial video compression scheme, and configuring the decoding module to: Decoding the factors of each facial frame separately; and The processing module is configured to: For each facial frame, the facial information is processed based on the factors.
23. An encoding device comprising: A receiving module configured to receive a video sequence; an encoding module configured to encode one or more images of the video sequence; as well as A generating module configured to generate a bit stream; The encoding module is configured to: An identification number is sent, the identification number indicating whether a facial video generation compression scheme is used.
24. The encoding device according to claim 23, wherein: When using the facial video to generate the compression scheme, the encoding module is configured to: Sending a flag indicating whether a syntax element associated with the facial information exists; and When the flag indicates that the syntax element associated with the facial information exists, the corresponding syntax element associated with the facial information is sent.
25. The encoding device according to claim 24, wherein: The encoding module is configured to: sending a syntax element indicating a number of facial frames generated using the facial video compression scheme; and For each facial frame, a flag is sent, indicating whether there is a syntax element related to facial information.
26. The encoding device according to claim 24, wherein The encoding module is configured to: sending a syntax element indicating a number of facial frames generated using the facial video compression scheme; and For each face frame, a factor is sent, indicating a quantization factor for processing the face information.
27. An electronic device comprising: Memory, which stores instruction sets; and one or more processors configured to execute the instruction set so that the one or more processors perform the method of decoding a bit stream to output one or more images of a video stream according to any one of claims 1 to 9.
28. An electronic device comprising: Memory, which stores instruction sets; and one or more processors configured to execute the instruction set, so that the one or more processors perform the method for encoding a video sequence into a bit stream according to any one of claims 10 to 13.
29. A non-transitory computer-readable storage medium storing a bit stream of a video, which, when decoded by a processor, causes the processor to execute the method of decoding a bit stream to output one or more images of a video stream according to any one of claims 1 to 9.
30. A non-transitory computer-readable storage medium storing a video sequence of a video, which, when encoded by a processor, causes the processor to execute the method of encoding a video sequence into a bit stream according to any one of claims 10 to 13.
31. A computer program product comprising a plurality of computer program instructions, wherein: The plurality of computer program instructions enables a computer to perform the method of decoding a bit stream to output one or more images of a video stream according to any one of claims 1 to 9.
32. A computer program product comprising a plurality of computer program instructions, wherein: The plurality of computer program instructions causes a computer to perform the method of encoding a video sequence into a bitstream according to any one of claims 10 to 13.
33. A computer program, wherein The computer program causes a computer to execute the method of decoding a bit stream to output one or more images of a video stream according to any one of claims 1 to 9.
34. A computer program, wherein The computer program causes a computer to execute the method of encoding a video sequence into a bit stream according to any one of claims 10 to 13.