Method and apparatus for facial video compression
By encoding and decoding video sequences through facial semantics and 3D mesh reconstruction, the method addresses the challenge of high compression efficiency in advanced video coding standards, improving video transmission and storage efficiency.
Patent Information
- Application Number
- JP2025542088
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-09
- Filing Date
- 2024-01-15
- Publication Date
- 2026-02-10
AI Technical Summary
Existing video coding standards face challenges in achieving high compression efficiency, particularly with the development of advanced standards like VVC/H.266, which require improved techniques to maintain subjective quality at reduced bandwidth.
The method involves encoding video sequences by compressing reference pictures and converting inter-pictures into facial semantics, using a 3D mesh reconstruction based on facial semantics for decoding, and employing a processor and memory system to execute these processes.
This approach enhances video compression efficiency, allowing for higher quality video transmission and storage with reduced bandwidth requirements, particularly in applications involving facial video content.
Smart Images

Figure 2026504934000001_ABST
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This disclosure claims the benefit of priority to U.S. Provisional Application No. 63 / 480,568, filed January 19, 2023, and U.S. Patent Application No. 18 / 408,100, filed January 9, 2024. All of the above applications are expressly incorporated herein by reference in their entirety.
[0002] FIELD OF THE DISCLOSURE The present disclosure relates generally to video processing, and more particularly to methods and apparatus for facial video compression. [Background technology]
[0003] Video is a set of static pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and decompressed before display. The compression process is typically called encoding, and the decompression process is typically called decoding. There are a variety of video coding formats that use standardized video coding techniques, most commonly based on prediction, transform, quantization, entropy coding, and in-loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, which specify specific video coding formats, are developed by standardization organizations. As video standards adopt increasingly advanced video coding techniques, the coding efficiency of new video coding standards becomes increasingly higher. Summary of the Invention
[0004] An embodiment of the present disclosure provides a method for encoding a video sequence into a bitstream, the method including: receiving a video sequence; encoding one or more pictures of the video sequence; and generating a bitstream. The encoding includes compressing a reference picture; converting, based on the reference picture, a plurality of inter-pictures associated with the reference picture into facial semantics; and encoding the facial semantics.
[0005] An embodiment of the present disclosure provides a method for decoding a bitstream to output one or more pictures of a video stream, the method including receiving a bitstream and decoding one or more pictures using facial semantics in the bitstream, the decoding including reconstructing a three-dimensional (3D) mesh based on the facial semantics and generating the one or more pictures based on the reconstructed reference frame and the facial semantics.
[0006] An embodiment of the present disclosure provides a non-transitory computer-readable storage medium for storing a video bitstream, the bitstream including coded reference pictures and coded facial semantics for a plurality of inter-frames, the facial semantics being determined based on the reference frames and the plurality of inter-frames.
[0007] An embodiment of the present disclosure provides an apparatus, the apparatus including a processor and a memory configured to store executable instructions for the processor, the processor configured to perform a method according to the above embodiment by reading and executing the executable instructions from the memory.
[0008] An embodiment of the present disclosure provides a computer program product including computer program instructions, said computer program instructions enabling a computer to perform the method according to the above embodiment.
[0009] An embodiment of the present disclosure provides a computer program, which enables a computer to perform the method according to the above embodiment.
[0010] Embodiments and various aspects of the present disclosure are illustrated in the following detailed description and accompanying drawings, in which various features illustrated in the drawings are not drawn to scale. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a schematic diagram illustrating an exemplary system for encoding image data, according to some embodiments of the present disclosure. [Figure 2] FIG. 1 is a schematic diagram illustrating the architecture of a block-based video compression framework according to some embodiments of the present disclosure. [Figure 3] 1 is a schematic diagram illustrating the structure of an exemplary video sequence, according to some embodiments of the present disclosure. [Figure 4A] FIG. 2 is a schematic diagram illustrating an exemplary block-based encoding process according to some embodiments of the present disclosure. [Figure 4B] FIG. 10 is a schematic diagram illustrating another exemplary block-based encoding process according to some embodiments of the present disclosure. [Figure 5A] FIG. 2 is a schematic diagram illustrating an exemplary block-based decoding process according to some embodiments of the present disclosure. [Figure 5B] FIG. 10 is a schematic diagram illustrating another exemplary block-based decoding process according to some embodiments of the present disclosure. [Figure 6] FIG. 1 is a schematic diagram illustrating an example architecture of an end-to-end deep learning-based video compression framework, according to some embodiments of the present disclosure. [Figure 7]FIG. 1 is a schematic diagram illustrating an example architecture of a deep learning-based video generative compression framework, according to some embodiments of the present disclosure. [Figure 8] FIG. 1 is a schematic diagram illustrating an exemplary encoder-decoder coding framework with a compact feature size of 1x4x4 for talking face video, according to some embodiments of the present disclosure. [Figure 9] FIG. 1 is a schematic diagram illustrating a general encoder-decoder generated compression framework for 3DMM-assisted talking face video, according to some embodiments of the present disclosure. [Figure 10] 1 illustrates an exemplary 3DMM-assisted face interactive coding framework according to some embodiments of the present disclosure. [Figure 11] 1 is a flowchart of an exemplary method for face interactive coding according to some embodiments of the present disclosure. [Figure 12] 10 is a flowchart of another exemplary method for face interactive coding according to some embodiments of the present disclosure. [Figure 13] 1 illustrates a flowchart of an example process for context-based entropy coding according to some embodiments of the present disclosure. [Figure 14] 1 illustrates an exemplary decoder framework according to some embodiments of the present disclosure. [Figure 15] 1 shows a flowchart of an exemplary decoding process according to some embodiments of the present disclosure. [Figure 16] 1 illustrates a flowchart of an exemplary process for 3D face mesh reconstruction, according to some embodiments of the present disclosure. [Figure 17] 1 illustrates a flowchart of an exemplary process for mesh-based motion estimation according to some embodiments of the present disclosure. [Figure 18] 1 illustrates a flowchart of an exemplary process for frame generation according to some embodiments of the present disclosure. [Figure 19]1 is a block diagram of an exemplary apparatus for encoding image data according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0012] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description will refer to the accompanying drawings, in which like numbers in different drawings represent the same or similar elements unless otherwise stated. The implementations set forth in the following description of exemplary embodiments do not represent all implementations in accordance with the present invention. Instead, they are merely examples of apparatus and methods according to inventive aspects as recited in the claims. Specific aspects of the present disclosure are described in more detail below. In the event of a conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein shall control.
[0013] The ITU-T Video Coding Expert Group (ITU-T VCEG) and the ISO / IEC Moving Picture Expert Group (ISO / IEC MPEG) Joint Video Experts Team (JVET) are currently developing the General Purpose Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 at half the bandwidth.
[0014] To achieve the same subjective quality as HEVC / H.265 at half the bandwidth, JVET has developed technology beyond HEVC using the Joint Exploration Model (JEM) reference software. As coding techniques are incorporated into JEM, JEM achieves substantially higher coding performance than HEVC.
[0015] The VVC standard has evolved recently and continues to include more coding techniques that provide better compression performance. VVC is based on the same hybrid video coding system used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.
[0016] Video is a set of static pictures (or "frames") arranged in a temporal sequence to store visual information. A video capture device (e.g., a camera) can be used to capture and store the pictures in a temporal sequence, and a video playback device (e.g., a television, a computer, a smartphone, a tablet computer, a video player, or any end-user terminal with display capabilities) can be used to display the pictures in a temporal sequence. In some applications, the video capture device can also transmit the captured video in real time to a video playback device (e.g., a computer with a monitor) for purposes such as surveillance, conferencing, or live broadcasting.
[0017] To reduce the storage space and transmission bandwidth required for such applications, video can be compressed before storage and transmission and decompressed before display. Compression and decompression can be achieved by software executed by a processor (e.g., a general-purpose computer processor) or dedicated hardware. A module for compression is typically called an "encoder," and a module for decompression is typically called a "decoder." Encoders and decoders may be collectively referred to as a "codec." Encoders and decoders can be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of encoders and decoders may include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of encoders and decoders may include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed on a computer-readable medium. Video compression and decompression can be achieved by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, and the H.26x series. In some applications, a codec can decompress video from a first coding standard and recompress the decompressed video in a second coding standard, in which case the codec can be called a "transcoder."
[0018] A video coding process can identify and retain useful information that can be used to reconstruct an image and ignore information that is not important for reconstruction. If the ignored, unimportant information cannot be perfectly reconstructed, such a coding process can be called "lossy." Otherwise, it can be called "lossless." Most coding processes are lossy, which is a tradeoff to reduce the required storage space and transmission bandwidth.
[0019] Useful information about the picture being coded (called the "current picture") includes changes relative to a reference picture (e.g., a previously coded and reconstructed picture). Such changes can include pixel position changes, luminance changes, or color changes, of which position changes are of primary concern. Position changes of pixels representing an object can reflect the motion of the object between the reference picture and the current picture.
[0020] A picture coded without reference to another picture (i.e., its own reference picture) is called an "I-picture." If some or all of the blocks in a picture (e.g., blocks that generally refer to portions of a video picture) are predicted with one reference picture (e.g., uniprediction) by intra-prediction or inter-prediction, the picture is called a "P-picture." If at least one block in a picture is predicted with two reference pictures (e.g., biprediction), the picture is called a "B-picture."
[0021] 1 is a schematic diagram illustrating a system 100 for encoding image data according to some disclosed embodiments. Image data may include an image (also called a "picture" or "frame"), multiple images, or a video. An image is a static picture. Multiple images may be spatially or temporally related or unrelated. A video is a set of images arranged in a temporal sequence.
[0022] 1 , system 100 includes a source device 120 that provides encoded video data that is subsequently decoded by a destination device 140. Consistent with disclosed embodiments, each of source device 120 and destination device 140 may include any of a wide range of devices, such as a desktop computer, a notebook (e.g., laptop) computer, a server, a tablet computer, a set-top box, a mobile phone, a vehicle, a camera, an image sensor, a robot, a television, a camera, a wearable device (e.g., a smart watch, a wearable camera), a display device, a digital media player, a video game console, a video streaming device, etc. Source device 120 and destination device 140 may be configured for wireless or wired communication.
[0023] 1, source device 120 may include image / video preprocessor 122, image / video encoder 124, and output interface 126. Destination device 140 may include input interface 142, image / video decoder 144, and machine vision application 146. Image / video encoder 124 encodes an input bitstream and outputs it via output interface 126 as encoded bitstream 162. Encoded bitstream 162 is transmitted via communication medium 160 and received at input interface 142. Image / video decoder 144 decodes encoded bitstream 162 to generate decoded data.
[0024] More specifically, source device 120 may further include various devices (not shown) for providing source image data that is processed by image / video encoder 124. Devices for providing source image data may include an image / video capture device such as a camera, an image / video archive or storage device containing previously captured images / video, or an image / video feed interface that receives images / video from an image / video content provider.
[0025] Image / video encoder 124 and image / video decoder 144 may each be implemented as any of a variety of suitable encoder or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. If encoding or decoding is partially realized in software, image / video encoder 124 or image / video decoder 144 may perform techniques according to this disclosure by storing software instructions on a suitable non-transitory computer-readable medium and executing them in hardware by one or more processors. Each of image / video encoder 124 or image / video decoder 144 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (codec) in the respective device.
[0026] Image / video encoder 124 and image / video decoder 144 may operate according to any video coding standard, such as Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), AOMedia Video 1 (AV1), Joint Photographic Experts Group (JPEG), Moving Picture Experts Group (MPEG), etc. Alternatively, image / video encoder 124 and image / video decoder 144 may be customized devices that do not conform to existing standards. Although not shown in FIG. 1 , in some embodiments, image / video encoder 124 and image / video decoder 144 may be integrated with an audio encoder and decoder, respectively, and may include appropriate MUX-DEMUX units or other hardware and software to handle the encoding of both audio and video in a common data stream or separate data streams.
[0027] Output interface 126 may include any type of medium or device capable of transmitting encoded bitstream 162 from source device 120 to destination device 140. For example, output interface 126 may include a transmitter or transceiver configured to transmit encoded bitstream 162 in real time from source device 120 directly to destination device 140. Encoded bitstream 162 may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device 140.
[0028] Communication medium 160 may include a transient medium, such as a wireless broadcast or a wired network transmission. For example, communication medium 160 may include a radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). Communication medium 160 may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. In some embodiments, communication medium 160 may include routers, switches, base stations, or other equipment useful for facilitating communication from source device 120 to destination device 140. For example, a network server (not shown) may receive encoded bitstream 162 from source device 120, e.g., via a network transmission, and provide encoded bitstream 162 to destination device 140.
[0029] Communication medium 160 may be in the form of a storage medium (e.g., a non-transitory storage medium), such as a hard disk, a flash drive, a compact disc, a digital video disc, a Blu-ray disc, volatile or non-volatile memory, or other suitable digital storage medium for storing encoded image data. In some embodiments, a computing device at a media production facility, such as a disc stamping facility, may receive the encoded image data from source device 120 and create a disc containing the encoded video data.
[0030] Input interface 142 may include any type of medium or device capable of receiving information from communication medium 160. The received information includes encoded bitstream 162. For example, input interface 142 may include a receiver or transceiver configured to receive encoded bitstream 162 in real time.
[0031] The system 100 can be configured to perform video encoding and decoding based on block-based video compression techniques, deep learning-based video compression techniques, talking face video compression techniques, and the like.
[0032] Block-based video compression techniques use a block-based hybrid video coding framework to exploit spatial redundancy, temporal redundancy, and information entropy redundancy in video. This hybrid video coding framework includes motion compensation (e.g., intra / inter prediction), transforms (e.g., discrete cosine transforms), quantization, and entropy coding. Block-based video compression techniques can comply with various image / video coding standards, such as JPEG, JPEG2000, H.264 / MPEG4 Part 10, Audio Visual Coding Standard (AVS), H.265 / HEVC, and Versatile Video Coding (VVC).
[0033] 2 is a schematic diagram illustrating a block-based video compression framework 200 according to some embodiments of the present disclosure. The block-based video compression framework 200 may include an encoder configured to generate a bitstream based on input video frames and a decoder configured to reconstruct the video frames based on the bitstream. For simplicity, FIG. 2 shows only the encoder side of the block-based video compression framework 200. The decoder side of the block-based video compression framework 200 can be thought of as inverting the operations of the encoder side.
[0034] Specifically, as shown in Figure 2, the input frame xt is divided into a set of blocks, such as square regions of the same size (e.g., 8x8). The block-based video compression framework 200 includes the following steps.
[0035] The block-based video compression framework 200 performs motion estimation using a block-based motion estimation module 201. The motion estimation module 201 estimates the motion of a current frame x t and the previous reconstructed frame
number
[0036] The block-based video compression framework 200 performs motion compensation using a motion compensation module 202. The motion vector v determined by the motion estimation module 201 is t The predicted frame is calculated by copying the corresponding pixels in the previous reconstructed frame to the current frame based on
number
number
number
[0037] The block-based video compression framework 200 performs the transform and quantization using a transform module 203 and a Q module 204, respectively. t is provided by the Q module 204.
number
[0038] The block-based video compression framework 200 performs the inverse transform using the inverse transform module 205.
number
number
[0039] The block-based video compression framework 200 performs entropy coding using the entropy coding module 206. The motion vector v t and the quantization result
number
[0040] The block-based video compression framework 200 performs frame reconstruction using a reconstruction module 207.
number
number
number
number
[0041] The bitstream generated by the entropy coding module 206 can be decoded at the decoder side (not shown in FIG. 2) by performing motion compensation, inverse quantization, and frame reconstruction to generate reconstructed frames.
number
[0042] Details of the block-based video compression framework 200 are further described with reference to FIGS. 3, 4A, 4B, 5A, and 5B. Specifically, FIG. 3 illustrates the structure of an exemplary video sequence 300, according to some embodiments of the present disclosure. The video sequence 300 may be live video or captured and archived video. The video 300 may be real-life video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real-life video with augmented reality effects). The video sequence 300 may be input from a video capture device (e.g., a camera), a video archive containing pre-captured video (e.g., video files stored on a storage device), or a video feed interface (e.g., a video broadcast transceiver) for receiving video from a video content provider.
[0043] As shown in FIG. 3, video sequence 300 may include a series of pictures arranged temporally along a timeline, including pictures 302, 304, 306, and 308. Pictures 302-306 are consecutive, with more pictures between pictures 306 and 308. In FIG. 3, picture 302 is an I-picture, and its reference picture is picture 302 itself. Picture 304 is a P-picture, and its reference picture is picture 302, as indicated by the arrow. Picture 306 is a B-picture, and its reference pictures are pictures 304 and 308, as indicated by the arrows. In some embodiments, the reference picture of a picture (e.g., picture 304) may not be immediately preceding or following the picture. For example, the reference picture of picture 304 may be a picture before picture 302. It should be noted that the reference pictures of pictures 302 to 306 are merely examples, and the present disclosure does not limit the embodiment of the reference pictures to the example shown in FIG.
[0044] Typically, video codecs do not encode or decode an entire picture at once due to the computational complexity of such a task. Rather, they divide a picture into elementary segments and can encode or decode the picture segment by segment. Such elementary segments are referred to as basic processing units ("BPUs") in this disclosure. For example, structure 310 in FIG. 3 illustrates an exemplary structure of a picture (e.g., any of pictures 302-308) of video sequence 300. In structure 310, the picture is divided into 4x4 basic processing units, the boundaries of which are indicated by dashed lines. In some embodiments, the basic processing units may be referred to as "macroblocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) or "coding tree units" ("CTUs") in some other video coding standards (e.g., H.265 / HEVC, H.266 / VVC, or AVS). The basic processing units may have variable sizes in the picture, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, etc., or may have pixels of any shape or size. The size and shape of the basic processing unit can be selected based on a balance between coding efficiency and the level of detail retained in the basic processing unit for the picture.
[0045] A basic processing unit may be a logic unit that may contain groups of different types of video data stored in a computer memory (e.g., a video frame buffer). For example, a basic processing unit for a color picture may include a luma component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb, Cr) representing color information, and related syntax elements, among which the luma component and the chroma component may have the same size of the basic processing unit. The luma component and the chroma component may be referred to as a "coding tree block" ("CTB") in some video coding standards (e.g., H.265 / HEVC, H.266 / VVC, or AVS). Any operation performed on a basic processing unit can be repeatedly performed on each of its luma and chroma components.
[0046] Video coding involves multiple operational stages, examples of which are shown in Figures 4A-4B and 5A-5B. At each stage, even the size of a basic processing unit may be too large for processing, so in this disclosure, it can be further divided into segments called "basic processing subunits." In some embodiments, a basic processing subunit may be called a "block" in some video coding standards (e.g., MPEG family, H.261, H.263, H.264 / AVC, or AVS) or a "coding unit" ("CU") in some other video coding standards (e.g., H.265 / HEVC, H.266 / VVC, or AVS). A basic processing subunit may be the same size as a basic processing unit or smaller. Similar to a basic processing unit, a basic processing subunit is also a logic unit that may contain groups of different types of video data (e.g., Y, Cb, Cr, and associated syntax elements) stored in computer memory (e.g., a video frame buffer). Any operation performed on a basic processing sub-unit can be performed repeatedly on each of its luma and chroma components. It should be noted that such division can be carried out to further levels depending on the processing needs. It should also be noted that different schemes can be used to divide the basic processing units at different stages.
[0047] For example, in a mode decision stage (an example of which is shown in FIG. 4B), an encoder can decide what prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for a basic processing unit, which may be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing sub-units (e.g., CUs in H.265 / HEVC, H.266 / VVC, AVS) and decide the prediction type for each individual basic processing sub-unit.
[0048] As another example, in the prediction stage (an example of which is shown in FIGS. 4A-4B), the encoder may perform prediction operations at the level of basic processing subunits (e.g., CUs). However, in some cases, even basic processing subunits may be too large to process. The encoder may further divide the basic processing subunits into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC, H.266 / VVC, and AVS) at a level at which prediction operations can be performed.
[0049] As another example, in the transform stage (one example of which is shown in FIGS. 4A-4B), the encoder can perform transform operations on the remaining basic processing subunits (e.g., CUs). However, in some cases, even the basic processing subunits may be too large to process. The encoder can further divide the basic processing subunits into smaller segments (e.g., called "transform blocks" or "TBs" in H.265 / HEVC, H.266 / VVC, and AVS) at a level at which the transform operations can be performed. Note that the division scheme of the same basic processing subunit may differ between the prediction stage and the transform stage. For example, in H.265 / HEVC, H.266 / VVC, or AVS, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.
[0050] 3, the basic processing unit 312 is further divided into 3x3 basic processing sub-units, the boundaries of which are indicated by dotted lines. Different basic processing units of the same picture can be divided into basic processing sub-units in different schemes.
[0051] In some implementations, to provide parallel processing capabilities and error resilience for video encoding and decoding, a picture can be divided into processing regions such that the encoding or decoding process for a region of the picture does not depend on information from other regions of the picture. In other words, each region of the picture can be processed independently. This allows different regions of the picture to be processed in parallel, thereby improving coding efficiency. Also, if data for a region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thereby providing error resilience. Some video coding standards allow pictures to be divided into different types of regions. For example, H.265 / HEVC, H.266 / VVC, and AVS provide two types of regions: "slices" and "tiles." Note that different pictures in the video sequence 300 may have different partitioning schemes for dividing the picture into regions.
[0052] For example, in Figure 3, structure 310 is divided into three regions 314, 316, and 318, the boundaries of which are indicated by solid lines within structure 310. Region 314 includes four basic processing units. Regions 316 and 318 each include six basic processing units. Note that the basic processing units, basic processing subunits, and regions of structure 310 in Figure 3 are merely examples, and the present disclosure is not limited to such embodiments.
[0053] FIG. 4A shows a schematic diagram of an exemplary encoding process 400A according to an embodiment of the present disclosure. For example, encoding process 400A may be performed by an encoder. As shown in FIG. 4A, the encoder may encode a video sequence 402 into a video bitstream 428 via process 400A. Similar to video sequence 300 in FIG. 3, video sequence 402 may include a set of pictures (referred to as "original pictures") arranged in a temporal order. Similar to structure 310 in FIG. 3, each original picture in video sequence 402 may be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder may perform process 400A at the level of basic processing units for each original picture in video sequence 402. For example, the encoder may perform process 400A iteratively, in which the encoder may encode a basic processing unit in one iteration of process 400A. In some embodiments, the encoder may perform process 400A in parallel for a region of each original picture of video sequence 402 (eg, regions 314-318).
[0054] In Figure 4A, an encoder may generate prediction data 406 and prediction BPU 408 by providing a fundamental processing unit (referred to as an "original BPU") of an original picture of a video sequence 402 to a prediction stage 404. The encoder may generate a residual BPU 410 by subtracting the prediction BPU 408 from the original BPU. The encoder may generate quantized transform coefficients 416 by providing the residual BPU 410 to a transform stage 412 and a quantization stage 414. The encoder may generate a video bitstream 428 by providing the prediction data 406 and the quantized transform coefficients 416 to a binary coding stage 426. The components 402, 404, 406, 408, 410, 412, 414, 416, 426, and 428 may be referred to as a "forward pass." In process 400A, after quantization stage 414, the encoder may generate a reconstructed residual BPU 422 by providing the quantized transform coefficients 416 to an inverse quantization stage 418 and an inverse transform stage 420. The encoder may generate a predicted reference 424 used for the next iteration of process 400A in prediction stage 404 by adding the reconstructed residual BPU 422 to the prediction BPU 408. The components 418, 420, 422, and 424 of process 400A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that both the encoder and decoder use the same reference data for prediction.
[0055] The encoder can iteratively perform process 400A to encode each original BPU of the original picture (in the forward pass) and generate a prediction reference 424 for encoding the next original BPU of the original picture (in the reconstruction pass). After encoding all original BPUs of the original picture, the encoder can proceed to encode the next picture in the video sequence 402.
[0056] Referring to process 400A, an encoder may receive a video sequence 402 generated by a video capture device (e.g., a camera). As used herein, "receive" may refer to receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or operating in any manner to input data.
[0057] In the prediction stage 404, in the current iteration, the encoder may receive the original BPU and a predicted reference 424 and perform a prediction operation to generate predicted data 406 and a predicted BPU 408. The predicted reference 424 may be generated from a reconstruction pass of a previous iteration of the process 400A. The purpose of the prediction stage 404 is to reduce information redundancy by extracting predicted data 406 from the predicted data 406 and the predicted reference 424 that can be used to reconstruct the original BPU as a predicted BPU 408.
[0058] Ideally, the predicted BPU 408 would be identical to the original BPU. However, because the prediction and reconstruction operations are not ideal, the predicted BPU 408 typically differs slightly from the original BPU. To record such differences, the encoder can generate the predicted BPU 408 and then subtract it from the original BPU to generate the residual BPU 410. For example, the encoder can subtract pixel values (e.g., grayscale or RGB values) of the predicted BPU 408 from corresponding pixel values of the original BPU. Each pixel of the residual BPU 410 may have a residual value as a result of the subtraction between corresponding pixels of the original BPU and the predicted BPU 408. Compared to the original BPU, the predicted data 406 and the residual BPU 410 may have fewer bits, which can be used to reconstruct the original BPU without significant quality degradation. This compresses the original BPU.
[0059] To further compress the residual BPU 410, in the transform stage 412, the encoder can reduce spatial redundancy of the residual BPU 410 by decomposing it into a set of two-dimensional "base patterns," each associated with a "transform coefficient." The base patterns may have the same size (e.g., the size of the residual BPU 410). Each base pattern may represent a variation frequency (e.g., frequency of luminance variation) component of the residual BPU 410. None of the base patterns can be reproduced from any combination (e.g., a linear combination) of any other base patterns. In other words, the decomposition can decompose the variation of the residual BPU 410 into the frequency domain. Such a decomposition is analogous to a discrete Fourier transform of a function, in which the base patterns are analogous to the basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are analogous to the coefficients associated with the basis functions.
[0060] Different transform algorithms can use different base patterns. The transform stage 412 can use various transform algorithms, such as a discrete cosine transform or a discrete sine transform. The transform of the transform stage 412 is reversible. That is, the encoder can reconstruct the residual BPU 410 by inversely operating the transform (called an "inverse transform"). For example, to reconstruct a pixel of the residual BPU 410, the inverse transform can multiply each coefficient by the value of the corresponding pixel in the base pattern and add the products to generate a weighted sum. For a video coding standard, both the encoder and the decoder can use the same transform algorithm (and thus the same base pattern). Therefore, the encoder can record only the transform coefficients, from which the decoder can reconstruct the residual BPU 410 without receiving the base pattern from the encoder. Compared to the residual BPU 410, the transform coefficients may have fewer bits, which can be used to reconstruct the residual BPU 410 without significant quality degradation. This further compresses the residual BPU 410.
[0061] The encoder can further compress the transform coefficients in the quantization stage 414. In the transform process, different base patterns can represent different fluctuation frequencies (e.g., luminance fluctuation frequencies). Because the human eye is generally better at recognizing low-frequency fluctuations, the encoder can ignore high-frequency fluctuation information without significantly degrading the decoding quality. For example, in the quantization stage 414, the encoder can generate quantized transform coefficients 416 by dividing each transform coefficient by an integer value (referred to as a "quantization scale factor") and rounding the quotient to the nearest integer. After such an operation, some transform coefficients of the high-frequency base pattern can be converted to zero, and the transform coefficients of the low-frequency base pattern can be converted to smaller integers. The encoder can ignore the zero-valued quantized transform coefficients 416, thereby further compressing the transform coefficients. The quantization process is also lossless, in that the quantized transform coefficients 416 can be reconstructed into transform coefficients in the inverse operation of quantization (referred to as "dequantization").
[0062] The quantization stage 414 can be lossy because the encoder ignores any remainder of such divisions in its rounding operation. Typically, the quantization stage 414 can result in the most information loss in the process 400A. The more information loss, the fewer bits the quantized transform coefficients 416 require. To achieve different levels of information loss, the encoder can use different values of the quantization parameter or any other parameter of the quantization process.
[0063] In binary coding stage 426, the encoder may encode the prediction data 406 and the quantized transform coefficients 416 using a binary coding technique, such as entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or other lossless or lossy compression algorithm. In some embodiments, the encoder may encode other information in addition to the prediction data 406 and the quantized transform coefficients 416 in binary coding stage 426, such as the prediction mode used in prediction stage 404, parameters of the prediction operation, the transform type in transform stage 412, parameters of the quantization process (e.g., quantization parameters), and encoder control parameters (e.g., bitrate control parameters). The encoder may generate a video bitstream 428 using the output data of binary coding stage 426. In some embodiments, the video bitstream 428 may be further packetized for network transmission.
[0064] Referring to the reconstruction path of process 400A, in an inverse quantization stage 418, the encoder may generate reconstructed transform coefficients by performing inverse quantization on the quantized transform coefficients 416. In an inverse transform stage 420, the encoder may generate a reconstructed residual BPU 422 based on the reconstructed transform coefficients. The encoder may generate a predicted reference 424 to be used in the next iteration of process 400A by adding the reconstructed residual BPU 422 to the predicted BPU 408.
[0065] It should be noted that other variations of process 400A can be used to encode video sequence 402. In some embodiments, the stages of process 400A may be performed in a different order by the encoder. In some embodiments, one or more stages of process 400A may be combined into a single stage. In some embodiments, a single stage of process 400A may be split into multiple stages. For example, transform stage 412 and quantization stage 414 may be combined into a single stage. In some embodiments, process 400A may include additional stages. In some embodiments, process 400A may omit one or more stages in FIG. 4A.
[0066] 4B shows a schematic diagram of another exemplary encoding process 400B according to an embodiment of the present disclosure. Process 400B can be modified from process 400A. For example, process 400B can be used by an encoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 400A, the forward path of process 400B further includes a mode decision stage 430 and divides the prediction stage 404 into a spatial prediction stage 4042 and a temporal prediction stage 4044. The reconstruction path of process 400B further includes a loop filter stage 432 and a buffer 434.
[0067] Generally, prediction techniques can be categorized into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-prediction") can predict a current BPU using pixels from one or more already-coded neighboring BPUs in the same picture. That is, the prediction reference 424 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter-prediction") can predict a current BPU using regions from one or more already-coded pictures. That is, the prediction reference 424 in temporal prediction can include coded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.
[0068] Referring to process 400B, in the forward pass, the encoder performs prediction operations in a spatial prediction stage 4042 and a temporal prediction stage 4044. For example, in the spatial prediction stage 4042, the encoder may perform intra prediction. For an original BPU of a picture being encoded, the prediction reference 424 may include one or more neighboring BPUs of the same picture that were coded (in the forward pass) and reconstructed (in the reconstruction pass). The encoder may generate the predicted BPU 408 by extrapolating the neighboring BPUs. Extrapolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, the encoder may perform extrapolation at the pixel level, such as by extrapolating the value of a corresponding pixel for each pixel of the predicted BPU 408. The neighboring BPUs used for extrapolation can be positioned relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., bottom left, bottom right, top left, or top right of the original BPU), or from any direction defined by the video coding standard being used. For intra prediction, the prediction data 406 may include, for example, the positions (e.g., coordinates) of the neighboring BPUs used, the sizes of the neighboring BPUs used, parameters of the extrapolation, the orientation of the neighboring BPUs used relative to the original BPU, etc.
[0069] Also for example, in the temporal prediction stage 4044, the encoder may perform inter-prediction. For an original BPU of a current picture, the prediction reference 424 may include one or more pictures (called "reference pictures") that have been coded (in the forward pass) and reconstructed (in the reconstruction pass). In some embodiments, the reference pictures may be coded and reconstructed for each BPU. For example, the encoder may generate a reconstructed BPU by adding the reconstructed residual BPU 422 to the predicted BPU 408. Once all reconstructed BPUs of the same picture have been generated, the encoder may generate the reconstructed picture as a reference picture. The encoder may perform a "motion estimation" operation to search for a matching region within a certain range (called a "search window") of the reference picture. The position of the search window in the reference picture may be determined based on the position of the original BPU in the current picture. For example, the search window may be centered at a location in the reference picture that has the same coordinates as the original BPU in the current picture and may extend to a predetermined distance. If the encoder identifies a region in the search window that is similar to the original BPU (e.g., by a pel recursion algorithm, a block matching algorithm, etc.), the encoder can determine such a region as a matching region. The matching region may have different dimensions than the original BPU (e.g., smaller than the original BPU, equal to the original BPU, larger than the original BPU, or a different shape). Because the reference picture and the current picture are temporally separated on a timeline (e.g., as shown in FIG. 3), the matching region can be considered to "move" to the position of the original BPU over time. The encoder can record the direction and distance of the motion as a "motion vector." If multiple reference pictures are used (e.g., as in picture 306 in FIG. 3), the encoder can search for a matching region and determine its associated motion vector for each reference picture.In some embodiments, the encoder may assign weights to pixel values in the matching region of each matching reference picture.
[0070] Motion estimation can be used to identify various types of motion, such as, for example, translation, rotation, zoom, etc. For inter prediction, prediction data 406 may include, for example, the location (e.g., coordinates) of the matching region, a motion vector associated with the matching region, the number of reference pictures, weights associated with the reference pictures, etc.
[0071] To generate the predicted BPU 408, the encoder may perform a "motion compensation" operation. Motion compensation can be used to reconstruct the predicted BPU 408 based on the prediction data 406 (e.g., motion vectors) and the prediction reference 424. For example, the encoder may shift the matching region of the reference picture by the motion vector, within which the encoder can predict the original BPU of the current picture. If multiple reference pictures are used (e.g., as in picture 306 in FIG. 3), the encoder may shift the matching region of the reference picture by the average pixel value of the matching region and each motion vector. In some embodiments, if the encoder has assigned weights to the pixel values of the matching region of each matching reference picture, it may add a weighted sum of the pixel values of the shifted matching region.
[0072] In some embodiments, inter-prediction may be unidirectional or bidirectional. Unidirectional inter-prediction can use one or more reference pictures in the same temporal direction relative to the current picture. For example, picture 304 in FIG. 3 is a unidirectional inter-predicted picture in which a reference picture (e.g., picture 302) precedes picture 304. Bidirectional inter-prediction can use one or more reference pictures in both temporal directions relative to the current picture. For example, picture 306 in FIG. 3 is a bidirectional inter-predicted picture in which reference pictures (e.g., pictures 304, 308) are in both temporal directions relative to picture 304.
[0073] Still referring to the forward pass of process 400B, after spatial prediction 4042 and temporal prediction stage 4044, in mode decision stage 430, the encoder can select a prediction mode (e.g., one of intra-prediction or inter-prediction) for the current iteration of process 400B. For example, the encoder can perform a rate-distortion optimization technique, in which the encoder can select a prediction mode to minimize the value of a cost function that depends on the distortion of a reconstructed reference picture in a candidate prediction mode and the bitrate of the candidate prediction mode. Depending on the selected prediction mode, the encoder can generate a corresponding predicted BPU 408 and predicted data 406.
[0074] In the reconstruction path of process 400B, if intra-prediction mode is selected in the forward path, after generating the prediction reference 424 (e.g., the encoded and reconstructed current BPU in the current picture), the encoder can provide the prediction reference 424 directly to the spatial prediction stage 4042 for later use (e.g., for extrapolation of the next BPU of the current picture). The encoder can provide the prediction reference 424 to the loop filter stage 432, where the encoder can apply a loop filter to the prediction reference 424 to reduce or eliminate distortions (e.g., blocking artifacts) introduced during coding of the prediction reference 424. The encoder can apply various loop filter techniques to the loop filter stage 432, such as deblocking, sample adaptive offset, adaptive loop filtering, etc. The loop-filtered reference picture can be stored in a buffer 434 (or a “decoded picture buffer”) for later use (e.g., for use as an inter-prediction reference picture for a future picture in the video sequence 402). The encoder may store one or more reference pictures in a buffer 434 for use in the temporal prediction stage 4044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) along with the quantized transform coefficients 416, the prediction data 406, and other information in the binary coding stage 426.
[0075] FIG. 5A shows a schematic diagram of an exemplary decoding process 500A according to an embodiment of the present disclosure. Process 500A may be a decompression process corresponding to compression process 400A in FIG. 4A. In some embodiments, process 500A may be similar to the reconstruction path of process 400A. A decoder can decode video bitstream 428 into video stream 504 based on process 500A. Video stream 504 may be very similar to video sequence 402. However, due to information loss in the compression and decompression processes (e.g., quantization stage 414 in FIGS. 4A-4B), video stream 504 is generally not identical to video sequence 402. Similar to processes 400A and 400B in FIGS. 4A-4B, a decoder can perform process 500A at the basic processing unit (BPU) level for each coded picture in video bitstream 428. For example, an encoder may perform process 500A iteratively, in which the encoder decodes a fundamental processing unit in one iteration of process 500A. In some embodiments, a decoder may perform process 500A in parallel for a region (e.g., regions 314-318) of each coded picture in video bitstream 428.
[0076] In FIG. 5A , a decoder may provide a portion of a video bitstream 428 associated with a basic processing unit (referred to as a “coding BPU”) of a coded picture to a binary decoding stage 502. In the binary decoding stage 502, the decoder may decode the portion into prediction data 406 and quantized transform coefficients 416. The decoder may generate a reconstructed residual BPU 422 by providing the quantized transform coefficients 416 to an inverse quantization stage 418 and an inverse transform stage 420. The decoder may generate a prediction BPU 408 by providing the prediction data 406 to a prediction stage 404. The decoder may generate a prediction reference 424 by adding the reconstructed residual BPU 422 to the prediction BPU 408. In some embodiments, the prediction reference 424 may be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder may provide the prediction reference 424 to the prediction stage 404 for performing a prediction operation in a next iteration of the process 500A.
[0077] The decoder may iteratively perform process 500A to decode each of the coded BPUs of the coded picture and generate the predicted reference 424 for coding the next coded BPU of the coded picture. After decoding all of the coded BPUs of the coded picture, the decoder may output the picture to the video stream 504 for display and proceed to decode the next coded picture in the video bitstream 428.
[0078] In binary decoding stage 502, the decoder may perform the inverse of the binary coding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or other lossless compression algorithm). In some embodiments, the decoder may decode other information in binary decoding stage 502, in addition to prediction data 406 and quantized transform coefficients 416, such as, for example, a prediction mode, parameters of the prediction operation, a transform type, parameters of the quantization process (e.g., quantization parameters), and encoder control parameters (e.g., bitrate control parameters). In some embodiments, if video bitstream 428 is transmitted in packets over a network, the decoder may depacketize video bitstream 428 before providing it to binary decoding stage 502.
[0079] 5B shows a schematic diagram of another exemplary decoding process 500B according to an embodiment of the present disclosure. Process 500B can be modified from process 500A. For example, process 500B can be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 500A, process 500B further divides prediction stage 404 into a spatial prediction stage 4042 and a temporal prediction stage 4044, and further includes a loop filter stage 432 and a buffer 434.
[0080] In process 500B, for a coding basic processing unit (referred to as a "current BPU") of a coding picture being decoded (referred to as a "current picture"), prediction data 406 decoded by the decoder from binary decoding stage 502 may include various data depending on the prediction mode used by the encoder to encode the current BPU. For example, if intra prediction is used by the encoder to encode the current BPU, prediction data 406 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, etc. The parameters of the intra prediction operation may include, for example, the positions (e.g., coordinates) of one or more reference neighboring BPUs, the size of the neighboring BPUs, extrapolation parameters, directions of the neighboring BPUs relative to the original BPU, etc. Also, for example, if inter prediction is used by the encoder to encode the current BPU, prediction data 406 may include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, etc. Parameters for the inter-prediction operation may include, for example, the number of reference pictures currently associated with the BPU, weights associated with each of the reference pictures, the locations (e.g., coordinates) of one or more matching regions in each reference picture, one or more motion vectors associated with each of the matching regions, etc.
[0081] Based on the prediction mode indicator, the decoder may determine whether to perform spatial prediction (e.g., intra prediction) in spatial prediction stage 4042 or temporal prediction (e.g., inter prediction) in temporal prediction stage 4044. Details of performing such spatial or temporal prediction have been described in FIG. 4B and will not be repeated below. After performing such spatial or temporal prediction, the decoder may generate a predicted BPU 408. The decoder may generate a predicted reference 424 by adding the predicted BPU 408 and a reconstructed residual BPU 422, as shown in FIG. 5A.
[0082] In process 500B, the decoder may supply the prediction reference 424 to the spatial prediction stage 4042 or the temporal prediction stage 4044 for performing a prediction operation in the next iteration of process 500B. For example, if the current BPU is decoded by intra prediction in the spatial prediction stage 4042, after generating the prediction reference 424 (e.g., the decoded current BPU), the decoder may supply the prediction reference 424 directly to the spatial prediction stage 4042 for later use (e.g., for extrapolation of the next BPU of the current picture). If the current BPU is decoded by inter prediction in the temporal prediction stage 4044, after generating the prediction reference 424 (e.g., the reference picture decoded by all BPUs), the decoder may reduce or remove distortion (e.g., blocking artifacts) by supplying the prediction reference 424 to the loop filter stage 432. The decoder may apply a loop filter to the prediction reference 424 as described in FIG. 4B. The loop-filtered reference picture may be stored in a buffer 434 (e.g., a decoded picture buffer in computer memory) for later use (e.g., for use as an inter-predicted reference picture for a future coded picture of the video bitstream 428). The decoder may store one or more reference pictures in the buffer 434 for use in the temporal prediction stage 4044. In some embodiments, the prediction data may further include loop filter parameters (e.g., loop filter strength). In some embodiments, if the prediction mode indicator of the prediction data 406 indicates that inter-prediction was used to encode the current BPU, the prediction data includes loop filter parameters.
[0083] In addition to block-based video compression techniques, deep learning can be used in video compression to achieve competitive performance compared to traditional compression schemes. For example, end-to-end image compression algorithms, with end-to-end training and nonlinear transformations, have shown better rate-distortion (RD) performance than JPEG, JPEG2000, and even HEVC. Furthermore, video compression algorithms based on deep neural networks (DNNs), such as the Deep Video Compression Model (DVC), can achieve promising RD performance. These schemes can function without prior knowledge of the video content. For video conferencing / telephony applications, deep generative models such as First Order Motion Model (FOMM) and Face Video-to-Video Synthesis (Face_vid2vid) can achieve promising performance at ultra-low bitrates. In particular, these models leverage the fact that human motion information typically resides in these video variations, providing powerful priors that can be used for frame synthesis. These features are described by variations in human structures, such as landmarks and keypoints, and can be further propagated to animate reference frames and generate human motion video.
[0084] Deep learning-based algorithms can be used to replace or enhance some operations or functions of block-based video coding tools, such as intra / inter prediction, entropy coding, and in-loop filtering. Instead of designing specific modules, end-to-end image / video compression algorithms can be used for joint optimization of the entire image / video compression framework. For example, the DVC scheme, an end-to-end video coding scheme that jointly optimizes all components for video compression, can be used. Furthermore, to address issues of content adaptation and error propagation awareness, an online encoder update scheme can be used to improve video compression performance. Furthermore, FVC can be used by developing all key modules of the end-to-end compression framework in feature space. Based on a recurrent probability model and a weighted recurrent quality enhancement network, recurrent learning for video compression (RLVC) and HLVC can be used to exploit temporal correlation between video frames. Four effective modules in multi-frame prediction for learning-based video compression (M-LVC) can be used. However, like traditional video coding tools, these learning-based video compression methods target general natural scenes without specific consideration of human content such as faces, bodies, or other parts of the body.
[0085] FIG. 6 is a schematic diagram illustrating an example architecture of an end-to-end deep learning-based video compression framework 600 according to some embodiments of the present disclosure. The framework 600 uses various deep learning models to jointly optimize video compression components, such as motion estimation, motion compression, and residual compression. Specifically, it utilizes learning-based optical flow estimation to obtain motion information and reconstruct the current frame. It then uses two autoencoder-style neural networks to compress the corresponding motion and residual information. The modules in the framework 600 are jointly trained by a single loss function, in which they cooperate by considering the trade-off between reducing the number of compression bits and improving the quality of the decoded video. There is a one-to-one correspondence between the block-based video compression framework 200 shown in FIG. 2 and the end-to-end deep learning-based video compression framework 600 shown in FIG. 6. A brief summary of the relationship and differences is provided below. The end-to-end deep learning-based video compression framework 600 may include an encoder configured to generate a bitstream based on input video frames and a decoder configured to reconstruct the video frames based on the bitstream. For simplicity, FIG. 6 shows only the encoder side of the end-to-end deep learning-based video compression framework 600.
[0086] As shown in Figure 6, the framework 600 can perform motion estimation and compression. The optical flow network module 601 uses a CNN (convolutional neural network) model to estimate motion information v t Instead of directly encoding the raw optical flow values, we use an MV encoder-decoder network to compress and decode the optical flow values. First, we use the MV encoder net module 602 to generate motion information v t The motion information v can be encoded. t The coded motion representation of m tand by the Q module 603,
number
number
[0087] The framework 600 can also perform motion compensation. A motion compensation network, or motion compensation net module 605, calculates a predicted frame based on the obtained optical flow.
number
number
number
[0088] The framework 600 can also perform transforms, quantization, and inverse transforms. Linear transforms are replaced by using a highly nonlinear residual encoder-decoder network, such as the residual encoder net module 606 shown in Figure 6, which generates the residual r t is the expression y t Then, y t is calculated by the Q module 607.
number
number
number
[0089] The framework 600 can also perform entropy coding. In the test stage, the quantized motion representation
number
number
number
number
[0090] Additionally, the original frame, the reconstructed frame, and the encoded frame can be used to determine the loss of the framework 600. The determined loss can also be used to refine the kernel workings within the framework 600 to achieve better performance.
[0091] Framework 600 may also perform frame reconstruction (not shown in FIG. 6) similar to the frame reconstruction described in connection with framework 200.
[0092] The end-to-end deep learning-based video compression framework 600 can be used for facial video compression, such as talking face generative video coding. For example, end-to-end deep learning-based talking face generative video coding can use generative models such as variational autoencoding (VAE) and generative adversarial networks (GAN). Promising performance improvements can be achieved with facial video compression. For example, X2Face can be used to control face generation via image, audio, and pose codes. Furthermore, a realistic neural talking head model can be used via multi-shot adversarial learning. For video-to-video synthesis tasks, Face-vid2vid can be used. A scheme that leverages compact 3D keypoint representations to drive a generative model for rendering target frames can also be used. Furthermore, a mobile-enabled video chat system based on FOMM can be used. VSBNet, which uses adversarial learning to reconstruct original frames from landmarks, can also be used. Furthermore, an end-to-end talking head video compression framework based on compact feature learning (CFTE), designed for highly efficient talking face video compression for ultra-low bandwidth scenes, can be used. The CFTE scheme leverages compact feature representation to compensate for temporal evolution and reconstruct target facial video frames in an end-to-end manner. Furthermore, the CFTE scheme can be incorporated into video coding frameworks for rate-distortion supervision. Although these algorithms achieve frame reconstruction with a small number of facial parameters due to the powerful rendering capabilities of deep generative models, some head pose movements and facial expression movements are still not accurately rendered compared to the original video.
[0093] FIG. 7 is a schematic diagram illustrating an exemplary deep learning-based video generation and compression framework 700 according to some embodiments of the present disclosure. The framework 700 is suitable for compressing and generating talking face videos. For example, the framework 700 can be based on a first-order motion model (FOMM). The FOMM deforms a reference source frame to follow the motion of the driving video. This method works for various types of video (e.g., motion pictures, cartoons), but can also be used for facial animation applications. The FOMM follows an encoder-decoder architecture with a motion transfer component that includes the following steps:
[0094] First, the keypoint extractor (also called the motion module) is trained using equivariant loss without explicit labels. This keypoint extractor calculates two sets of 10 trained keypoints for the source frame and the driving frame. The trained keypoints are transformed from a channel size × 64 × 64 feature map through a Gaussian map function, so that each corresponding keypoint can represent different channel feature information. Note that every keypoint is an (x,y) point that can represent the most important information in the feature map.
[0095] A dense motion network then uses the landmarks and source frames to generate a dense motion field and occlusion map.
[0096] The encoder 710 then encodes the source frames using a conventional image / video compression scheme such as HEVC / VVC or JPEG / BPG, where VVC is used to compress the source frames.
[0097] In a later stage, the resulting feature map is warped by a dense motion field (by a differentiable grid sample operation) and multiplied with the occlusion map.
[0098] Finally, the decoder 720 generates the image from the warped map.
[0099] Figure 8 is a schematic diagram illustrating an example encoder-decoder coding framework 600 with a compact feature size of 1x4x4 for talking face video, according to some embodiments of the present disclosure. Figure 8 provides another basic framework for a deep-based video generative compression scheme based on compact feature representation, i.e., CFTE, which follows an encoder-decoder architecture that applies a context-based coding scheme.
[0100] On the encoder 810 side, the compression framework includes three modules: an encoder (also called a VVC encoding module) for compressing key frames, a feature extractor for extracting compact human features for other inter frames, and a feature coding module for compressing inter-predicted residuals of the compact human features. First, a key frame representing human texture is compressed by the VVC encoder. Through the compact feature extractor, each subsequent inter frame is represented by a compact feature matrix with a size of 1x4x4. Note that the size of the compact feature matrix is not fixed, and the number of feature parameters can be increased or decreased depending on the specific requirements for bit consumption. These extracted features are then inter-predicted and quantized, and the residuals are finally entropy coded as the final bitstream.
[0101] On the decoder 820 side, this compression framework also includes three main modules: decoding to reconstruct keyframes, reconstructing compact features through entropy decoding and compensation, and generating final videos using the reconstructed features and decoded keyframes. More specifically, during final video generation, compact feature extraction can represent decoded keyframes from a VVC bitstream as features. Next, considering features from keyframes and interframes, associated sparse motion fields are calculated to facilitate the generation of pixel-wise dense motion maps and occlusion maps. Finally, based on a deep generative model, the decoded keyframes, pixel-wise dense motion fields, and occlusion maps with implicit motion field characterization are used to generate final videos with accurate appearance, pose, and expression.
[0102] To further improve coding performance, numerous studies have focused on 3D faces. They employ a 3D head model to encode only the pose parameters for the task of face-specific video compression. Subsequently, both eigenspace and principal component analysis (PCA) models have been used for this task. However, based on these traditional 3D techniques, the visual quality of the reconstructed images is unacceptable. With the development of deep generative models, this 3D MMM-assisted face video generation task can yield promising results.
[0103] 9 is a schematic diagram illustrating a general encoder-decoder generation compression framework 900 for 3DMM-assisted talking face video, according to some embodiments of the present disclosure. Generally, 3DMM-assisted face video generation involves the use of shape
number
number
number
number
number
number
[0104] Although conventional or learning-based end-to-end video compression methods can achieve relatively high-efficiency compression performance for talking face video, directly applying common compression algorithms to ultra-low bitrate talking face video compression systems has some drawbacks.
[0105] First, block-based hybrid coding schemes or learning-based end-to-end video compression methods designed for universal video scenes cannot reduce semantic redundancy in terms of specific scenes that exhibit strong prior knowledge and statistical regularity. As a result, such algorithms are not suitable for ultra-low bitrate human video compression scenes.
[0106] Second, traditional or learning-based end-to-end video compression methods target generic natural scenes without specifically considering human motion information. In particular, these features from talking faces and moving bodies are described by the variation of feature structures with strong prior feature structures, such as landmarks and keypoints, which greatly help in reconstructing higher quality videos.
[0107] Furthermore, although generative compression algorithms such as FOMM and Face_vid2vid fully realize frame reconstruction with a small number of parameters due to the powerful rendering capabilities of deep generative models, some head pose movements and facial expression movements are still not accurately rendered compared to the original talking face or moving body footage. That is, most of the movements based on 2D face representations (i.e., 2D landmarks and 2D keypoints) either perform poorly in terms of photorealism, do not satisfy the identity preservation problem, or do not fully transcribe driving poses and expressions.
[0108] Furthermore, existing 2D generative compression algorithms do not support including semantic information for controlling head movement pose in the compressed codestream, significantly limiting their application in human communication in the meta-universe. Meanwhile, in most 3D MMM-assisted generative models, the transmitted facial parameters are still complex, requiring more compression bits and unable to satisfy ultra-low bandwidth communication scenes.
[0109] To overcome the above challenges, this disclosure provides a face-interactive coding framework for ultra-low-bitrate, highly controllable, and privacy-preserving facial communication. The disclosed face-interactive coding paradigm follows the statistical regularities and semantic meaning of talking faces, which are successfully projected into a low-dimensional representation with high independence in terms of mouth motion, eye blinks, head pose, head translation, and head position. The compact facial semantic characterization can significantly remove redundancy and achieve high compression efficiency, while at the same time enabling the advancement of talking face video reconstruction toward controllable synthesis and friendly interaction.
[0110] 10 illustrates an exemplary 3DMM-aided face-interactive coding framework 1000 according to some embodiments of the present disclosure. As shown in FIG. 10, the 3DMM-aided face-interactive coding framework 1000 includes an encoder 1020 and a decoder 1040. Typically, an input video signal 1010 is encoded by the encoder 1020 to generate a coded bitstream 1030, and the coded bitstream 1030 is decoded by the decoder 1040 to obtain an output video signal 1050. The output video signal 1050 obtained by the proposed 3DMM-aided face-interactive coding framework 1000 may include three types of outputs: an ultra-low bitrate face communication 1051, a highly controllable face communication 1052, and a privacy-preserving face communication 1053.
[0111] Specifically, the 3DMM-assisted facial interactive coding framework 1000 includes three processes for ultra-low bitrate, high controllability, and privacy protection, respectively. An input video signal 1010 is compressed by a VVC encoding module 1021 (e.g., a conventional VVC codec) with a 3DMM-assisted facial semantics extraction model 1022 and a blink intensity prediction module 1023, and then encoded by an entropy coding module 1024. As shown in FIG. 10 , a coded bitstream 1030 obtained by the encoding process by the encoder 1020 may include a VVC bitstream 1031 and a compact facial semantics bitstream 1032. Then, a corresponding decoding process is performed by a decoder 1040 to obtain high-quality talking face video reconstruction at a very low bitrate. The decoder 1040 includes a VVC decoding module 1041, a 3DMM-assisted facial semantics extraction module 1042, and a blink intensity prediction module 1043. The decoder 104 further includes a decoded facial semantics buffer 1045, a 3D face mesh reconstruction module 1046, a mesh-based motion estimation module 1047, and a frame generation module 1048. In a first process, an ultra-low bitrate facial communication 1051 can be obtained by generating a frame with the decoded VVC frame (e.g., by the frame generation module 1048) and the motion information obtained by the mesh-based motion estimation module 1047. In a second process, the facial video can be further manipulated by a controllable semantic editing module 1061 for friendly interactivity. That is, facial semantics can be edited / modified before generating the frame. A final facial video is generated from the decoded VVC frame based on the edited / modified semantics. Therefore, a highly controllable facial communication 1052 can be obtained by the second process.In the third process, when generating frames with facial semantics, a virtual character image can be simulated by animating the virtual character reference module 1062, which results in privacy-preserving face communication 1053. It can be understood that quantization 1025 and inverse quantization 1044 can also be applied during the coding process.
[0112] More specifically, the encoder 1020 of the 3DMM-assisted face-interactive coding framework 1000 includes four sub-processing schemes: an intra-coding scheme based on a conventional hybrid coding framework (e.g., a VVC encoding module 1021) for compressing key reference frames; a 3DMM-assisted compact facial semantic representation module 1022 for characterizing facial semantic meaning for inter-frames; a blink prediction module 1023 for describing eye motion states for inter-frames; and a context-based encoding module 1024 for high-efficiency compression through inter-prediction. FIG. 11 is a flowchart of an exemplary method 1100 for face-interactive coding according to some embodiments of the present disclosure. The method 1100 may be performed by an encoder (e.g., by process 400A of FIG. 4A or 400B of FIG. 4B) or by one or more software or hardware components of an apparatus. In some embodiments, the method 1100 may be realized by a computer program product embodied in a computer-readable medium that includes computer-executable instructions, such as program code, executed by a computer. 10 and 11, the method 1100 may include the following steps 1102-1106.
[0113] In step 1102, a reference frame is compressed. For example, a key reference frame (i.e., the first frame) of the input video signal 1010 is compressed by a block-based hybrid coding framework (i.e., the VVC codec module 1021), which facilitates the creation of rich texture representation for successive frames.
[0114] In step 1104, based on the reference frame, a number of inter-frames associated with the reference frame are converted into facial semantics. For example, the input video signal 1010 is processed by a learning-based 3DMM-aided compact facial semantic representation module 1022 and a blink prediction module 1023 to project subsequent inter-frames into a set of facial semantics. In some embodiments, the facial semantics include head pose, face position, head translation, mouth motion, and blinks.
[0115] In step 1106, the facial semantics are encoded. Specifically, the transformed compact facial semantics are effectively inter-predicted, quantized, and entropy coded with a context-based entropy coding algorithm (e.g., by entropy coding module 1024).
[0116] The decoder 1040 of the 3DMM-assisted face interactive coding framework 1000 also includes four sub-process schemes: decoding to reconstruct key reference frames, reconstructing compact facial semantics through entropy decoding and compensation, reconstructing 3D face meshes through the decoded controllable facial semantics, and generating a final face image.
[0117] 12 is a flowchart of another exemplary method 1200 for face interactive coding according to some embodiments of the present disclosure. Method 1200 may be performed by a decoder (e.g., by process 500A of FIG. 5A or 500B of FIG. 5B) or by one or more software or hardware components of an apparatus. In some embodiments, method 1200 may be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions, such as program code, for execution by a computer. Referring to FIGS. 10 and 12, method 1200 may include the following steps 1202-1206:
[0118] In step 1202, the reference frame is reconstructed. The key reference frame is reconstructed by the VVC decoding model 1041, and the facial semantics of the key reference frame is extracted as a benchmark.
[0119] In step 1204, inter-frame facial semantics are decoded, for example, by context-based entropy decoding and compensation.
[0120] In step 1206, a 3D mesh is reconstructed based on the decoded facial semantics. Using the semantics of the inter-frames and key reference frames, a corresponding 3D face mesh can be reconstructed by the 3D face mesh reconstruction module 1046 and further sent to a mesh-based motion estimation module 1047 (e.g., a designed coarse-to-dense motion field reconstruction module) for estimating a dense motion field and facial attention map. In some embodiments, the reconstruction of the 3D face mesh can be controllable by modifying the facial semantics; for example, semantic parameters related to the facial semantics can be modified by the controllable semantic editing module 1061 so that the pose and expression of the 3D mesh can be changed toward personalized characterization.
[0121] In step 1208, one or more frames of the final talking face video are reconstructed. Using the reconstructed reference frames and explicit facial motion guidance (i.e., reconstructed facial semantics), the final talking face video can be reconstructed with high quality and personalized control due to the strong inference capabilities of the deep generative model. In some embodiments, a virtual character can be applied to one or more frames, and a 3D face mesh of the virtual character can also be reconstructed based on the semantics.
[0122] The encoding and decoding processes used in the proposed 3DMM-assisted face interactive coding framework are described in detail as follows.
[0123] In some embodiments, a process for facial semantics extraction in the 3DMM-assisted face interactive coding framework 1000 is described below. A talking face with a clear structure can be further developed from the statistical rules of PCA-based face scanning into a set of facial semantic representations (i.e., identity, expression, texture, etc.). Existing learning-based 3D face reconstruction models share an encoder-decoder architecture. The encoder typically uses a CNN backbone as a regressor to characterize a set of 3D face semantics, and at the same time, a 3DMM template is treated as a decoder to reconstruct a face mesh. Learning-based 3D face reconstruction can economically represent face images and significantly reduce coding bits for the task of talking face video compression. In some embodiments, facial semantics are characterized using multiple 3D regression parameters. Based on a pre-trained 3D face reconstruction model (i.e., WM3DR), a VVC reconstruction key reference frame is generated.
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0124] In addition, in 3D MMM-assisted face reconstruction (e.g., 3D face mesh reconstruction 1046), the expression coefficient
number
number
number
[0125] In some embodiments, facial behavior analysis is used to measure blink intensity (i.e.,
number
number
[0126] In some embodiments, the encoding process includes facial semantics encoding. These compact facial semantics δ compress To achieve high compression efficiency, a context-based entropy coding module (e.g., entropy coding module 1024 in FIG. 10) is employed. FIG. 13 shows a flowchart of an example process 1300 for context-based entropy coding according to some embodiments of the present disclosure. As shown in FIG. 13, process 1300 includes steps 1302 and 1304.
[0127] In step 1302, the residual is obtained by inter-predicting the facial semantics. For example, these compact facial semantics characterized from talking face frames are inter-predicted to remove redundancies. Il is given as follows:
number
number
number
number
number
[0128] In step 1304, the residuals are coded into a bitstream. For example, these inter-predicted residuals Res Il is quantized and further converted into a bitstream. When the coded bitstream is transmitted over a communication network and received at the decoder side, entropy decoding, dequantization, and frame compensation are performed serially to reconstruct the facial semantics.
[0129] FIG. 14 illustrates an exemplary decoder framework 1400 according to some embodiments of the present disclosure. As shown in FIG. 14, the decoder framework 1400 includes a facial mesh reconstruction 1410 (e.g., 3D face mesh reconstruction 1046 in FIG. 10 ), a mesh-based motion estimation 1420 (e.g., mesh-based motion estimation 1047 in FIG. 10 ), and a frame generation 1430 (e.g., frame generation 1047 in FIG. 10 ) based on a reconstructed key reference frame 1401 and a decoded compact facial semantics 1402. The decoder framework 1400 is designed to accurately reconstruct a face mesh from compact facial semantics characterizations and also explicitly depict a dense motion field and facial attention map. In this manner, talking face footage can be realistically reconstructed and perceptually compensated. FIG. 15 illustrates a flowchart of an exemplary decoding process 1500 according to some embodiments of the present disclosure. 14 and 15, a decoding process 1500 includes steps 1502-1506.
[0130] In step 1502, a facial mesh is reconstructed from the facial semantics, for example, by 3D facial mesh construction 1410.
[0131] In step 1504, a dense motion field and facial attention map for the face mesh is obtained, for example, by mesh-based motion estimation 1420.
[0132] In step 1506, one or more pictures (e.g., reconstructed face image 1403) are reconstructed and compensated based on the dense motion field and facial attention map, for example by frame generation 1430. In some embodiments, the face image is a talking face image.
[0133] In some embodiments, the process of 3D facial mesh reconstruction 1410 of the decoder framework 1400 is described below. Figure 16 shows a flowchart of an example process 1600 for 3D face mesh reconstruction, according to some embodiments of the present disclosure. With reference to Figures 14 and 16, the process 1600 includes steps 1602-1606.
[0134] In step 1602, a 3D face mesh of the VVC reconstruction key reference frame and the subsequent inter-frame is reconstructed. Taking the WM3DR model as an example, the decoder of the WM3DR model provides a parametric 3D MMM template 1416 to generate decoded semantic coefficients (i.e., σ) extracted (e.g., by facial semantics extraction 1411) from the VVC reconstruction key reference frame 1401 (i.e., the first frame).
number
number
number
number
number
number
number
number
number
number
number
number
number
[0135] In step 1604, the corresponding 2D face mesh is obtained. Based on the parameterized 3D face shape S and texture T, the 3D face vertices V can be synthesized. Then, the corresponding face mesh in the 2D image plane is
number
number
number
number
[0136] In step 1606, the motion of the eye region is recalibrated based on the 2D face mesh.
number
number
number
number
[0137] In conclusion, the decoded semantic parameter set
number
number
[0138] In some embodiments, a process for mesh-based motion estimation 1420 of decoder framework 1400 is described below. Figure 17 shows a flowchart of an example process 1700 for mesh-based motion estimation according to some embodiments of the present disclosure. With reference to Figures 14 and 17, process 1700 includes steps 1702 and 1704.
[0139] In step 1702, a coarse motion field (e.g., coarse mesh-based motion flow 1422) is obtained based on the motion of each vertex in the 2D face mesh from the VVC reconstructed key reference frame and the current inter frame.
number
number
number
number
number
number
number
number
number
number
[0140] In step 1704, a coarsely deformed frame is obtained based on the coarse motion field and the VVC reconstructed key reference frame. For example, the approximated coarse motion field
number
number
number
number
number
number
[0141] Rough deformation frame
number
[0142] In step 1706, a dense motion field (e.g., dense motion flow 1426A) and a facial attention map (e.g., facial attention map 1426B) are obtained based on the coarse motion field (e.g., 1421), the coarse deformation picture (e.g., 1424), and the blink motion map (e.g., 1417B). In this example, the mesh-approximated coarse motion field
number
number
number
number
number
number
[0143] In some embodiments, the process for frame generation 1430 of decoder framework 1400 is described below. The powerful inference capabilities of generative adversarial networks bring great benefits to the new paradigm of face generative compression. Generative adversarial networks are used to reconstruct high-fidelity talking face video through motion guidance information. A channel-split spatial feature transformation (CSSFT) mechanism can preserve face fidelity for reconstruction, and the CSSFT-GAN-based face reconstruction module utilizes a dense motion field
number
number
number
[0144] In step 1802, multi-scale spatial features are obtained, e.g., VVC reconstruction key reference frames
number
number
[0145] In step 1804, an attention-based feature warping operation on the multi-scale spatial features is performed to obtain warped facial spatial features. Specifically, these multi-scale spatial features
number
number
number
number
number
number
[0146] In step 1806, transformed facial features are obtained based on the warped facial spatial features. Specifically, the warped results are transmitted to a CSSFT-based face generation module (e.g., Frame Generation 1430), where the transformed facial features are generated through a convolutional layer for each resolution scale spatial feature.
number
number
number
[0147] In step 1808, a talking face frame (e.g., reconstructed face image 1403) is generated by concatenating the warped facial spatial features with the transformed facial features. For example,
number
number
number
number
number
[0148] We further describe the model supervision and loss function used in the disclosed 3DMM-aided face interactive coding framework. In the disclosed 3DMM-aided face interactive coding framework, a perceptual loss L is used to supervise the end-to-end training process. per , Hostile Ross L adv , id preservation loss L id , the reconstruction texture loss L tex Note that these related loss functions do not need to be used together and can be combined according to the actual task needs.
[0149] In summary, the overall end-to-end training loss is given by:
number
[0150] The disclosed 3DMM-assisted face interactive coding framework has various extended applications. The disclosed 3DMM-assisted face interactive coding framework can control head movement pose and facial expression in the compressed code stream, which can be further applied to ultra-low bandwidth face video conferencing, human communication in the meta-universe, virtual uploaders for live commerce, etc. Semantic parameter sets including mouth motion, eye blink, head pose, and head translation are also available.
number
number
[0151] "Ultra-Low Bandwidth Facial Video Conferencing": Recently, the demand for video conferencing / chat has increased dramatically. The disclosed 3DMM-assisted generative compression framework fully utilizes the strong statistical regularity of facial images and achieves facial image reconstruction for end users at ultra-low bitrates by simply compressing 3D compact parameters.
[0152] "Facial Communication in the Metaverse": The Metaverse is a virtual world and digital living space constructed by humans using digital technology, which is mapped by or transcends the real world and can interact with it. Human-character communication is crucial to this new social system. The disclosed 3DMM-assisted generative compression framework can effectively realize human parameter transfer and human character reconstruction. Specifically, by using the pose information and facial parameters in the extracted 3DMM template, corresponding face meshes can be rendered, and related motion information can be learned from these meshes and directly transferred to specific human characters. This ensures stable communication bitrates, similar to face video conferencing.
[0153] "Virtual Characters for Live Entertainment": With the growing demand for live streaming in Japan, virtual commerce characters are developing rapidly and appearing to be a new trend in live streaming. Virtual characters for live commerce maintain the interactivity of real people while providing better freshness and attracting more customers to live rooms. With the rise of young people, the two-dimensional world is also becoming more attractive. Therefore, the disclosed Steam can also be applied in this field, reducing the burden on network bandwidth caused by a large number of customers in a live room at the same time.
[0154] 19 is a block diagram of an exemplary apparatus 1900 for encoding image data, according to some embodiments of the present disclosure. The apparatus 1900 can be used to perform the video compression methods described above. As shown in FIG. 19, the apparatus 1900 can include a processor 1902. When the processor 1902 executes the instructions described herein, the apparatus 1900 can become a dedicated machine for video encoding or decoding. The processor 1902 can be any type of circuitry capable of manipulating or processing information. For example, processor 1902 may include any combination of any number of central processing units (“CPUs”), graphics processing units (“GPUs”), neural processing units (“NPUs”), microcontroller units (“MCUs”), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, IP cores, programmable logic arrays (PLAs), programmable array logic (PALs), generic array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems on chips (SoCs), or application specific integrated circuits (ASICs), etc. In some embodiments, processor 1902 may be a set of processors grouped as a single logical component. For example, as shown in FIG. 19, processor 1902 may include multiple processors, including processor 1902a, processor 1902b, and processor 1902n.
[0155] The device 1900 may include a memory 1904 configured to store data (e.g., a set of instructions, computer code, intermediate data, etc.). For example, as shown in FIG. 19, the stored data may include program instructions (e.g., program instructions for implementing the methods described herein). The processor 1902 may access the program instructions and data for processing (e.g., via a bus 1910) and perform operations or manipulations on the data for processing by executing the program instructions. The memory 1904 may include a high-speed random access storage device or a non-volatile storage device. In some embodiments, the memory 1904 may include any combination of any number of random access memories (RAMs), read-only memories (ROMs), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, compact flash (CF) cards, etc. The memory 1904 may also be a group of memories grouped as a single logical component (not shown in FIG. 19).
[0156] Bus 1910 may be a communication device that transfers data between components within apparatus 1900, such as an internal bus (e.g., a CPU memory bus) or an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Express port).
[0157] For ease of explanation and to avoid ambiguity, the processor 1902 and other data processing circuitry will be collectively referred to as "data processing circuitry" in this disclosure. The data processing circuitry may be implemented entirely as hardware or as a combination of software, hardware, or firmware. Additionally, the data processing circuitry may be a single, independent module or may be combined in whole or in part with other components of the device 1900.
[0158] The device 1900 may further include a network interface 1906 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.) In some embodiments, the network interface 1906 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communications ("NFC") adapters, or cellular network chips, etc.
[0159] In some embodiments, apparatus 1900 may further comprise a peripheral interface 1908 to provide connection with one or more peripheral devices. As shown in Figure 19, the peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, touchpad, or touchscreen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera, or an input interface coupled to a video archive), etc.
[0160] It should be noted that a video codec according to the present disclosure may be implemented as any combination of software or hardware modules in the device 1900. For example, some or all of the stages of the disclosed methods may be implemented as one or more software modules in the device 1900, such as program instructions loadable into the memory 1904. Also, for example, some or all of the stages of the disclosed methods may be implemented as one or more hardware modules in the device 1900, such as dedicated data processing circuitry (e.g., FPGA, ASIC, NPU, etc.).
[0161] In some embodiments, a non-transitory computer-readable storage medium storing a bitstream is also provided. The bitstream can be encoded and decoded according to the facial interactive codec method described above. For example, the bitstream can include an encoded reference frame and encoded facial semantics for multiple inter-frames, where the facial semantics are determined based on the reference frame and the multiple inter-frames. The encoded reference frame and the encoded facial semantics can be decoded according to the method described above, e.g., method 1200 (FIG. 12), and used to generate one or more pictures.
[0162] In some embodiments, a non-transitory computer-readable storage medium containing instructions is also provided, which may be executed by a device (e.g., the disclosed encoders and decoders) to perform the above-described methods. Common forms of non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape or other magnetic data storage media, CD-ROMs, any other optical data storage media, physical media with patterns of holes, RAM, PROMs, and EPROMs, flash EPROMs or other flash memory, NVRAM, cache, registers, other memory chips or cartridges, and networked versions thereof. A device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.
[0163] The embodiments described in this disclosure may be combined in any manner or used alone.
[0164] In summary, the face interactive coding framework has the following technical features:
[0165] The disclosed interactive coding framework for talking face video can ensure the service of immersive video conferencing and metaverse-related activities. Unlike existing controllable facial manipulation before video encoding or after video decoding, which has high complexity and latency, the disclosed framework is unified and flexible so as not to introduce additional complexity and latency. Specifically, the disclosed scheme can directly characterize talking face frames with highly disentangled facial semantics and manipulate them into the coding bitstream. In this way, the disclosed scheme enjoys the advantages of low complexity without extra manipulation, promising rate-distortion performance, and vivid facial character animation.
[0166] The disclosed transmission coding bitstream provides improved compact representation and semantic interpretation. First, compared to block-based video compression and feature-based end-to-end video compression, the present disclosure characterizes only talking face frames with 14-dimensional facial semantic parameters, significantly facilitating ultra-low bitrate (e.g., 2-4 kbps) facial video communication. Furthermore, the disclosed coding bitstream clarifies semantic meaning in terms of mouth motion, eye blinks, head pose, and head translation, providing benefits for editing facial semantics for interactivity and transferring facial semantics to virtual characters for user privacy.
[0167] Based on the above-mentioned mesh-based motion estimation module (e.g., mesh-based motion estimation module 1047 shown in FIG. 10) and frame generation module (e.g., frame generation module 1048 shown in FIG. 10), the face mesh reconstructed by the 3DMM template can be better evolved into a dense motion field and facial guidance map for pixel-wise face generation.
[0168] The proposed framework differs from existing face generative compression algorithms such as FOMM and Face_vid2vid, which describe motion from a first-order Taylor expansion in a neighborhood of learned sparse keypoints. In contrast, the disclosed motion representation Steam is designed for dense face meshes with strong geometry, directly mapping each vertex position to a 2D plane with its corresponding location. In this way, the corresponding facial semantics can be easily edited, thereby enabling the evolution of 3D face meshes toward personalized characterization.
[0169] Furthermore, compared to these generative algorithms that only employ hourglass networks to achieve feature warping, the disclosed interactive coding framework provides a GAN-based face reconstruction module with a channel-splitting spatial feature transformation mechanism, thus preserving the fidelity of the face during face reconstruction.
[0170] Finally, the disclosed framework can support ultra-low bandwidth face video conferencing, face interactive communication in the meta-universe, and virtual characters for live entertainment.
[0171] In some embodiments, an apparatus is provided, the apparatus including a processor and a memory configured to store executable instructions for the processor, the processor being configured to implement a method according to the above embodiments by reading and executing the executable instructions from the memory.
[0172] In some embodiments, a computer program product is provided comprising computer program instructions, said computer program instructions enabling a computer to perform a method according to the above embodiments.
[0173] In some embodiments, a computer program is provided, said computer program enabling a computer to carry out a method according to the above embodiments.
[0174] It should be noted that, in this specification, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another and do not require or imply an actual relationship or order between those entities or operations. Furthermore, the words "comprising," "having," "containing," "containing," and other similar forms are intended to be equivalent in meaning and open-ended in that they do not imply that the item or items following any of these words are an exhaustive list of the item or items, or that the items are limited to only the listed item or items.
[0175] As used herein, unless otherwise stated, the term "or" includes all possible combinations unless impossible. For example, if it is stated that a database may include A or B, then the database may also include A, or B, or A and B, unless otherwise stated or impossible. As a second example, if it is stated that a database may include A, B, or C, then the database may also include A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C, unless otherwise stated or impossible.
[0176] It is understood that the above embodiments can be realized by hardware, or software (program code), or a combination of hardware and software. If realized by software, it may be stored in the computer-readable medium. When executed by a processor, the software can perform the disclosed methods. The computing units and other functional units described in this disclosure can be implemented by hardware, or software, or a combination of hardware and software. Those skilled in the art can understand that multiple of the above modules / units may be combined into one module / unit, or that each of the above modules / units may be further divided into multiple sub-modules / sub-units.
[0177] In the foregoing specification, embodiments have been described with reference to numerous specific details that vary from embodiment to embodiment. Certain adaptations and variations can be made to the above-described embodiments. Other embodiments will be apparent to those skilled in the art from consideration of the detailed description and practice of the invention disclosed herein. It is intended that the specification and examples be considered exemplary, with the true scope and spirit of the invention being indicated by the following claims. It is also intended that the order of steps depicted in the figures is for illustrative purposes only and is not intended to be limited to the particular order of steps. Thus, one skilled in the art will recognize that steps can be performed in different orders when performing the same method.
[0178] In the drawings and specification, illustrative embodiments are disclosed. However, these embodiments are susceptible to many variations and modifications. Thus, although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation.
Claims
1. 1. A method for encoding a video sequence into a bitstream, comprising: receiving a video sequence; encoding one or more pictures of the video sequence; generating a bitstream; The step of encoding comprises: compressing the reference picture; converting, based on the reference picture, a plurality of inter-pictures associated with the reference picture into facial semantics; encoding the facial semantics. method.
2. The method of claim 1 , wherein the video sequence includes a talking face video.
3. The step of converting the plurality of inter-pictures associated with the reference picture into facial semantics based on a reference frame includes: The method of claim 1 , further comprising characterizing the facial semantics using a plurality of three-dimensional (3D) regression parameters.
4. The method of claim 3 , wherein the plurality of 3D regression parameters comprises an identity coefficient, an albedo coefficient, a scene lighting coefficient, an expression coefficient, a rotation coefficient, a translation coefficient, and a position coefficient.
5. The method of claim 4 , wherein the dimensionality of the expression coefficients is less than 64.
6. The method of claim 5 , wherein the dimension of the expression coefficients is six, and the expression coefficients represent a motion of a mouth region.
7. The facial semantics are The method of claim 1 , wherein one or more of the following are described: head pose, face position, head translation, mouth motion, or eye blinking.
8. The method of claim 7 , wherein blink intensity is predicted by facial behavior analysis.
9. The step of encoding the facial semantics comprises: obtaining a residual by inter-predicting the facial semantics; The method of claim 1 , further comprising the step of: encoding the residual into the bitstream.
10. 1. A method for decoding a bitstream to output one or more pictures of a video stream, comprising: receiving a bitstream; and decoding one or more pictures using facial semantics in the bitstream; The step of decoding comprises: reconstructing a reference picture; decoding interpicture facial semantics; reconstructing a three-dimensional (3D) mesh based on the facial semantics; generating the one or more pictures based on the reconstructed reference frames and facial semantics.
11. generating the one or more pictures based on the 3D mesh, obtaining a dense motion field and facial attention map of the 3D mesh; The method of claim 10 , further comprising: reconstructing and compensating the one or more pictures based on the dense motion field and the facial attention map.
12. reconstructing the 3D mesh based on the reconstructed reference frame and facial semantics, reconstructing a 3D face mesh for the reconstructed reference picture or the interpicture; obtaining a corresponding 2D face mesh of the 3D face mesh; The method of claim 10 , further comprising: recalibrating the motion of the eye region based on the 2D face mesh.
13. The step of obtaining the dense motion field and the facial attention map comprises: obtaining a coarse motion field based on the motion of each vertex in the 2D face mesh from the reconstructed reference picture and the current interpicture; obtaining a coarse distorted picture based on the coarse motion field and the reconstructed reference picture; The method of claim 11 , further comprising: obtaining the fine motion field and the facial attention map based on the coarse motion field, the coarsely deformed picture, and an eye blink motion map.
14. Obtaining multi-scale spatial features; obtaining warped facial spatial features by an attention-based feature warping operation on the multi-scale spatial features; obtaining transformed facial features based on the warped facial spatial features; 12. The method of claim 11, further comprising generating the one or more pictures by concatenating the warped facial spatial features and the transformed facial features.
15. The method of any one of claims 10 to 14, wherein the video stream includes a talking face video.
16. The facial semantics are 16. The method of any one of claims 10 to 15, describing one or more of head pose, face position, head translation, mouth motion, or eye blinking.
17. The step of reconstructing the 3D mesh based on the facial semantics comprises: modifying the facial semantics; The method of claim 10 , further comprising: constructing the 3D mesh based on the modified facial semantics.
18. prior to the step of generating the one or more pictures, The method of claim 10 , comprising applying a virtual character to the one or more frames.
19. 1. A non-transitory computer-readable storage medium for storing a video bitstream, the bitstream comprising: a coded reference picture; a plurality of inter-frame encoded facial semantics; The facial semantics are determined based on a reference frame and the plurality of inter-frames.
20. The facial semantics are 20. The non-transitory computer-readable storage medium of claim 19, describing one or more of head pose, face position, head translation, mouth motion, or eye blinking.
21. 1. An apparatus for encoding a video sequence into a bitstream, comprising: a processor; a memory configured to store executable instructions for the processor; 10. An apparatus, wherein the processor is configured to implement the method for encoding a video sequence into a bitstream according to claim 1 by reading and executing the executable instructions from the memory.
22. 1. An apparatus for decoding a bitstream and outputting one or more pictures of a video stream, comprising: a processor; a memory configured to store executable instructions for the processor; 19. An apparatus, wherein the processor is configured to implement the method of decoding a bitstream and outputting one or more pictures of a video stream according to any one of claims 10 to 18 by reading and executing the executable instructions from the memory.
23. A computer program product comprising computer program instructions that enable a computer to carry out the method of any one of claims 1 to 18.
24. A computer program enabling a computer to carry out the method according to any one of claims 1 to 18.