SEI message for generated facial image

By incorporating temporal upsampling for machine vision in SEI messages, the video coding efficiency is enhanced, addressing inefficiencies in existing video coding standards and improving performance in machine vision applications.

JP2026515597APending Publication Date: 2026-05-19ALIBABA INNOVATION PRIVATE LIMITED
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP Β· JP
Patent Type
Applications
Current Assignee / Owner
ALIBABA INNOVATION PRIVATE LIMITED
Filing Date
2024-04-03
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing video coding techniques, such as VVC/H.266, do not adequately consider temporal upsampling for machine vision in Additional Enhancement Information (SEI) messages, leading to inefficiencies in video compression and decompression processes.

Method used

Implementing temporal upsampling for machine vision based on SEI messages to enhance video coding efficiency by generating and utilizing SEI messages that include information for neural network post-processing filters.

Benefits of technology

Improves video coding efficiency by enabling better compression and decompression of video data, particularly in applications requiring machine vision tasks like object recognition and face recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026515597000001_ABST
    Figure 2026515597000001_ABST
Patent Text Reader

Abstract

The present invention provides a method and apparatus for processing video data using generated facial image supplemental enhancement information (SEI) messages. An exemplary method for generating a facial image includes receiving a bitstream, decoding the encoded information of the bitstream to obtain a base image and supplemental enhancement information (SEI) messages, determining whether the SEI messages are applicable to a neural network for generating a facial image, determining, in response that the SEI messages are applicable to a neural network for generating a facial image, the mode used to encode the facial image and the corresponding facial information parameters based on the SEI messages, and generating a facial image by the neural network based on the base image and facial information parameters.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Cross-reference of related applications) This disclosure claims priority to U.S. Provisional Application No. 63 / 494,493 filed on April 6, 2023, U.S. Provisional Application No. 63 / 511,200 filed on June 30, 2023, U.S. Provisional Application No. 63 / 587,763 filed on October 4, 2023, U.S. Provisional Application No. 63 / 618,387 filed on January 8, 2024, and U.S. Provisional Application No. 18 / 622,621 filed on March 29, 2024, all of which are incorporated herein by reference in their entirety.

[0002] This disclosure relates to video processing in general, and more specifically to a method and apparatus for performing facial image generation and compression using enhanced information (SEI) messages. [Background technology]

[0003] Video is a collection of still images (or "frames") that capture visual information. To reduce memory usage and transmission bandwidth, video can be compressed before being stored or transmitted and decompressed before being displayed. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, the most common of which are based on prediction, transformation, quantization, entropy coding, and in-loop filtering. Video coding standards that define specific video coding formats, such as High Efficiency Video Coding (HEVC / H.265), Multipurpose Video Coding (VVC / H.266), and AVS standards, are developed by standardization bodies. As more and more advanced video coding techniques are adopted in video standards, the coding efficiency of new video coding standards is increasing. [Overview of the project]

[0004] Embodiments of this disclosure provide a method and apparatus for processing video data using generated facial image enhancement information (SEI) messages.

[0005] According to some exemplary embodiments, a method for generating a face image is provided. This method includes receiving a bitstream, decoding the encoded information of the bitstream to obtain a base image and an additional enhancement information (SEI) message, determining whether the SEI message is applicable to a neural network for generating a face image, determining, in response that the SEI message is applicable to a neural network for generating a face image, a mode used to encode the face image and corresponding face information parameters based on the SEI message, and generating a face image by the neural network based on the base image and face information parameters.

[0006] According to some exemplary embodiments, a method is provided for encoding a video sequence into a bitstream. This method includes receiving a video sequence and generating a bitstream by encoding one or more images of the video sequence, including encoding a base image and an additional enhancement information (SEI) message of one or more images, the SEI message indicating a mode used for encoding a face image and corresponding face information parameters, the bitstream being used to generate a face image by a neural network based on the base image and face information parameters.

[0007] According to some exemplary embodiments, a non-temporary, computer-readable storage medium is provided for storing a bitstream of video. The bitstream includes a base image and an Additional Enhancement Information (SEI) message indicating a mode used to encode a face image and corresponding face information parameters, and the bitstream is used to generate a face image by a neural network based on the base image and face information parameters.

[0008] Embodiments and various aspects of the present disclosure are shown in the following detailed description and the accompanying drawings. The various features shown in the drawings are not drawn to scale.

Brief Description of the Drawings

[0009] [Figure 1] FIG. is a schematic diagram showing an exemplary system for encoding image data according to some embodiments of the present disclosure.

[0010] [Figure 2A] FIG. is a schematic diagram showing an exemplary encoding process of a hybrid video coding system according to an embodiment of the present disclosure.

[0011] [Figure 2B] FIG. is a schematic diagram showing another exemplary encoding process of a hybrid video coding system according to an embodiment of the present disclosure.

[0012] [Figure 3A] FIG. is a schematic diagram showing an exemplary decoding process of a hybrid video coding system according to an embodiment of the present disclosure.

[0013] [Figure 3B] FIG. is a schematic diagram showing another exemplary decoding process of a hybrid video coding system according to an embodiment of the present disclosure.

[0014] [Figure 4] FIG. is a block diagram of an exemplary apparatus for encoding image data according to some embodiments of the present disclosure.

[0015] [Figure 5] FIG. is a flowchart of an exemplary method for processing video based on a generated face video SEI message according to some embodiments of the present disclosure.

[0016] [Figure 6]This is another flowchart of an exemplary method for processing video based on generated face video SEI messages, according to some embodiments of the present disclosure.

[0017] [Figure 7A] This is a flowchart illustrating an exemplary method for generating a facial image according to some embodiments of the present disclosure.

[0018] [Figure 7B] This is a schematic diagram showing a neural network for generating facial images according to some embodiments of the present disclosure.

[0019] [Figure 8] This is a schematic diagram showing exemplary SEI messages according to some embodiments of the present disclosure.

[0020] [Figure 9] This is a schematic diagram showing another exemplary SEI message according to some embodiments of the present disclosure.

[0021] [Figure 10] This is a schematic diagram showing another exemplary SEI message according to some embodiments of the present disclosure.

[0022] [Figure 11] This is a schematic diagram showing another exemplary SEI message according to some embodiments of the present disclosure.

[0023] [Figure 12] This is a schematic diagram showing another exemplary SEI message according to some embodiments of the present disclosure.

[0024] [Figure 13] This is a flowchart of an exemplary method for encoding a video sequence into a bitstream, according to some embodiments of the present disclosure. [Modes for carrying out the invention]

[0025] Hereinafter, exemplary embodiments illustrated in the accompanying drawings will be referred to in detail. The following description refers to the accompanying drawings, and the same numbers in different drawings represent identical or similar elements unless otherwise indicated. The implementations described in the following description of exemplary embodiments do not represent all implementations of the present invention. Rather, they are merely examples of apparatus and methods relating to aspects of the invention described in the accompanying claims. Specific aspects of this disclosure will be described in more detail below. In the event of any conflict between terms and definitions provided herein and terms and / or definitions incorporated by reference, the terms and definitions provided herein shall prevail.

[0026] The Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (ITU-TVCEG) and the ISO / IEC Video Experts Group (ISO / IECMPEG) is currently developing the Multipurpose Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, VVC's goal is to achieve the same subjective quality as HEVC / H.265 using half the bandwidth. The VVC standard has been progressing smoothly since April 2018, with further coding techniques being added to provide even better compression performance. VVC is based on the same hybrid video coding system used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.

[0027] Neural network post-processing filters are used in image processing. In some existing techniques, Additional Enhancement Information (SEI) messages are used to specify the characteristics of the neural network post-processing filter. However, temporal upsampling for machine vision is not considered when generating SEI messages. Therefore, there is a need to implement temporal upsampling for machine vision based on SEI messages.

[0028] Figure 1 is a block diagram showing a system 100 for image data preprocessing and coding according to some disclosed embodiments. Image data may include images (also called β€œpictures” or β€œframes”), multiple images, or video. Images are still images. Multiple images may or may not be spatially or temporally related. Video is a collection of images arranged in a temporal sequence.

[0029] As shown in Figure 1, the system 100 includes a source device 120 that provides encoded video data to be later decoded by a destination device 140. According to the disclosed embodiments, each of the source device 120 and the destination device 140 may include any of a wide range of devices, including desktop computers, notebook (e.g., laptop) computers, servers, tablet computers, set-top boxes, mobile phones, vehicles, cameras, image sensors, robots, televisions, wearable devices (e.g., smartwatches or wearable cameras), display devices, digital media players, video game consoles, video streaming devices, and the like. The source device 120 and the destination device 140 may be equipped for wireless or wired communication.

[0030] As shown in Figure 1, the source device 120 may include an image / video preprocessor 122, an image / video encoder 124, and an output interface 126. The destination device 140 may include an input interface 142, an image / video decoder 144, and one or more machine vision applications 146. The image / video preprocessor 122 preprocesses image data, i.e., images or videos, to generate an input bitstream for the image / video encoder 124. The image / video encoder 124 encodes the input bitstream and outputs the encoded bitstream 162 via the output interface 126. The encoded bitstream 162 is transmitted via the communication medium 160 and received by the input interface 142. The image / video decoder 144 then decodes the encoded bitstream 162 to generate decoded data available for use by the machine vision application 146.

[0031] More specifically, the source device 120 may further include various devices (not shown) for providing source image data to be preprocessed by the image / video preprocessor 122. Devices for providing source image data may include image / video capture devices such as cameras, image / video archives or storage devices containing previously captured images / videos, or image / video feed interfaces for receiving images / videos from image / video content providers.

[0032] The image / video encoder 124 and the image / video decoder 144 may each be implemented as one or more suitable encoder or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. If encoding or decoding is partially implemented in software, the image / video encoder 124 or the image / video decoder 144 may implement the techniques of this disclosure by storing instructions for the software in a suitable non-temporary computer-readable medium and executing the instructions in hardware using one or more processors. Each of the image / video encoder 124 or the image / video decoder 144 may be included in one or more encoders or decoders, which may be integrated as part of a combined encoder / decoder (CODEC) in separate devices.

[0033] The image / video encoder 124 and image / video decoder 144 may operate according to any video coding standard such as Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), Multipurpose Video Coding (VVC), AO Media Video 1 (AV1), Joint Photographic Professional (JPEG), or Video Professional (MPEG). Alternatively, the image / video encoder 124 and image / video decoder 144 may be customized devices that do not conform to existing standards. Although not shown in Figure 1, in some embodiments, the image / video encoder 124 and image / video decoder 144 may be integrated with an audio encoder and decoder, respectively, to handle the encoding of both audio and video within a common data stream or separate data streams, or may include an appropriate MUX-DEMUX unit or other hardware and software.

[0034] The output interface 126 may include any type of medium or device capable of transmitting the encoded bitstream 162 from the source device 120 to the destination device 140. An example of the output interface 126 may be a transmitter or transceiver configured to transmit the encoded bitstream 162 directly from the source device 120 to the destination device 140 in real time. The encoded bitstream 162 may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to the destination device 140.

[0035] The communication medium 160 may include transient media such as wireless broadcast or wired network transmission. Examples of the communication medium 160 may include a radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). The communication medium 160 may form part of a packet-based network such as a local area network, a wide area network, or a global network (such as the Internet). In some embodiments, the communication medium 160 may include routers, switches, base stations, or any other equipment useful for facilitating communication from the source device 120 to the destination device 140. For example, a network server (not shown) may receive the encoded bitstream 162 from the source device 120 and provide the encoded bitstream 162 to the destination device 140, for example, by network transmission.

[0036] The communication medium 160 may be in the form of a storage medium (e.g., a non-temporary storage medium) such as a hard disk, flash drive, compact disc, digital video disc, Blu-ray disc, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded image data. In some embodiments, a computing device of a media manufacturing facility, such as a disc stamping machine, may receive encoded image data from the source device 120 and produce a disc containing the encoded video data.

[0037] The input interface 142 may include any type of medium or device capable of receiving information from the communication medium 160. The received information includes the encoded bitstream 162. An example of the input interface 142 may be a receiver or transceiver configured to receive the encoded bitstream 162 in real time.

[0038] The machine vision application 146 includes various hardware and / or software for utilizing the decoded image data generated by the image / video decoder 144. Examples of the machine vision application 146 include a display device for displaying the decoded image data to a user, and one of various display devices such as a cathode ray tube (CRT), liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, or any other type of display device. As another example, the machine vision application 146 may include one or more processors configured to perform various machine vision applications using the decoded image data, such as object recognition / tracking, face recognition, image matching, image / video search, augmented reality, robot vision / navigation, autonomous driving, 3D structure construction, stereo-enabled, and motion tracking.

[0039] Next, with reference to Figures 2A-2B and 3A-3B, exemplary image data encoding and decoding techniques (such as those implemented by encoder 124 and decoder 144 in Figure 1) will be described.

[0040] Figure 2A shows a schematic diagram of an exemplary encoding process 200A according to an embodiment of the present disclosure. For example, the encoding process 200A may be performed by an encoder, such as the image / video encoder 124 in Figure 1. As shown in Figure 2A, the encoder may encode a video sequence 202 into a video bitstream 228 according to process 200A. The video sequence 202 may include a set of images arranged in chronological order (referred to as β€œoriginal images”). Each original image in the video sequence 202 may be divided by the encoder into a basic processing unit, a basic processing subunit, or a region for processing. In some embodiments, the encoder may perform process 200A at the level of a basic processing unit for each original image in the video sequence 202. For example, the encoder may perform process 200A in an iterative manner, in which case the encoder may encode a basic processing unit in a single iteration of process 200A. In some embodiments, the encoder may perform process 200A in parallel for each region of the original images in the video sequence 202.

[0041] As shown in Figure 2A, the encoder can supply the basic processing unit (called the "original BPU") of the original image of the video sequence 202 to the prediction stage 204 to generate prediction data 206 and prediction BPU 208. The encoder can subtract the prediction BPU 208 from the original BPU to generate residual BPU 210. The encoder can supply the residual BPU 210 to the conversion stage 212 and the quantization stage 214 to generate quantization conversion coefficients 216. The encoder can supply the prediction data 206 and quantization conversion coefficients 216 to the binary coding stage 226 to generate video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be called the "forward path". During process 200A, after the quantization stage 214, the encoder can supply the quantization conversion coefficients 216 to the inverse quantization stage 218 and the inverse conversion stage 220 to generate the reconstructed residual BPU 222. The encoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate the prediction criterion 224, which will be used in the prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as the β€œreconstruction path”. The reconstruction path may be used to ensure that both the encoder and decoder use the same reference data for prediction.

[0042] The encoder can iteratively perform process 200A to encode each original BPU of the original image (in the forward path) and generate a prediction criterion 224 for encoding the next original BPU of the original image (in the reconstruction path). After encoding all original BPUs of the original image, the encoder can proceed to encode the next image in the video sequence 202.

[0043] Referring to process 200A, the encoder can receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term β€œreceive” can mean any action by any means for receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or inputting data.

[0044] In prediction stage 204, in the current iteration, the encoder receives the original BPU and prediction criterion 224 and can perform prediction calculations to generate prediction data 206 and prediction BPU 208. The prediction criterion 224 may be generated from the reconstruction path of the previous iteration in process 200A. The objective of prediction stage 204 is to reduce information redundancy by extracting prediction data 206 that can be used to reconstruct the original BPU as prediction BPU 208 from the prediction data 206 and prediction criterion 224.

[0045] Ideally, the predicted BPU 208 can be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 is generally slightly different from the original BPU. To record such differences, the encoder can generate the predicted BPU 208 and then subtract it from the original BPU to generate the residual BPU 210. For example, the encoder can subtract the pixel values ​​(e.g., grayscale values ​​or RGB values) of the predicted BPU 208 from the corresponding pixel values ​​of the original BPU. Each pixel of the residual BPU 210 can have a residual value as a result of such subtraction between the original BPU and the corresponding pixels of the predicted BPU 208. Although the predicted data 206 and residual BPU 210 may have fewer bits compared to the original BPU, they can be used to reconstruct the original BPU without significant quality degradation. Thus, the original BPU is compressed.

[0046] To further compress the residual BPU210, in the transformation stage 212, the encoder can reduce the spatial redundancy of the residual BPU210 by decomposing it into a set of two-dimensional "basis patterns," each associated with a "transformation coefficient." The basis patterns can have the same size (e.g., the size of the residual BPU210). Each basis pattern can represent the fluctuating frequency components (e.g., the frequency of the luminance fluctuations) of the residual BPU210. No basis pattern can be reconstructed from any combination of any other basis patterns (e.g., a linear combination). In other words, this decomposition allows the fluctuations of the residual BPU210 to be decomposed into the frequency domain. Such a decomposition is analogous to the discrete Fourier transform of a function, where the basis patterns are analogous to the base functions of the discrete Fourier transform (e.g., trigonometric functions), and the transformation coefficients are analogous to the coefficients associated with the base functions.

[0047] Different transformation algorithms can use different basis patterns. Various transformation algorithms can be used in transformation stage 212, such as discrete cosine transform and discrete sine transform. The transformation in transformation stage 212 is reversible; that is, the encoder can reconstruct the residual BPU210 by performing the inverse operation of the transformation (called the "inverse transformation"). For example, to reconstruct the pixels of the residual BPU210, the inverse transformation can be used to multiply the values ​​of the corresponding pixels in the basis pattern by the coefficients associated with each pixel, and then add the products to generate a weighted sum. In video coding standards, both the encoder and decoder can use the same transformation algorithm (i.e., the same basis pattern). Therefore, the encoder can record only the transformation coefficients that can reconstruct the residual BPU210 without receiving the basis pattern from the encoder. While the transformation coefficients have fewer bits compared to the residual BPU210, they can be used to reconstruct the residual BPU210 without significant quality degradation. Thus, the residual BPU210 is further compressed.

[0048] The encoder can further compress the conversion coefficients in the quantization stage 214. In the conversion process, different basis patterns may represent different fluctuation frequencies (e.g., luminance fluctuation frequencies). Because the human eye is generally more perceptible to low-frequency fluctuations, the encoder can ignore information about high-frequency fluctuations without causing significant quality degradation during decoding. For example, in the quantization stage 214, the encoder can generate quantization conversion coefficients 216 by dividing each conversion coefficient by an integer value (called a "quantization parameter") and rounding the quotient to the nearest integer. After such an operation, some conversion coefficients for high-frequency basis patterns are converted to zero, and conversion coefficients for low-frequency basis patterns are converted to smaller integers. The conversion coefficients are further compressed because the encoder can ignore the zero-value quantization conversion coefficients 216. The quantization process is reversible, and in this quantization process, the quantization conversion coefficients 216 can be reconstructed into conversion coefficients by the inverse operation of quantization (called "inverse quantization").

[0049] Because the encoder ignores the remainder of such division in rounding operations, the quantization stage 214 can be irreversible. Typically, the quantization stage 214 can result in the greatest loss of information in process 200A. The greater the loss of information, the fewer bits are required for the quantization conversion coefficient 216. To obtain different levels of loss of information, the encoder can use different values ​​for the quantization parameter or any other parameter of the quantization process.

[0050] In the binary coding stage 226, the encoder can encode the prediction data 206 and quantization conversion coefficients 216 using binary coding techniques such as entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other reversible or reversible compression algorithm. In some embodiments, in addition to the prediction data 206 and quantization conversion coefficients 216, the encoder can encode other information in the binary coding stage 226, such as the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the conversion type in the conversion stage 212, the parameters of the quantization process (e.g., quantization parameters), and encoder control parameters (e.g., bitrate control parameters). The encoder can generate a video bitstream 228 using the output data from the binary coding stage 226. In some embodiments, the video bitstream 228 may be further packetized for network transmission.

[0051] Referring to the reconstruction path of process 200A, in the inverse quantization stage 218, the encoder can perform inverse quantization on the quantization transformation coefficients 216 to generate reconstruction transformation coefficients. In the inverse transformation stage 220, the encoder can generate reconstruction residual BPU 222 based on the reconstruction transformation coefficients. The encoder can add the reconstruction residual BPU 222 to the prediction BPU 208 to generate a prediction criterion 224 to be used in the next iteration of process 200A.

[0052] It should be noted that the video sequence 202 can be encoded using other variations of process 200A. In some embodiments, the stages of process 200A may be performed in a different order by the encoder. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be divided into multiple stages. For example, the conversion stage 212 and the quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, one or more stages in Figure 2A may be omitted from process 200A.

[0053] Figure 2B shows a schematic diagram of another exemplary coding process 200B according to an embodiment of the present disclosure. For example, coding process 200B may be performed by an encoder such as the image / video encoder 124 in Figure 1. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B further includes a mode determination stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B further includes a loop filter stage 232 and a buffer 234.

[0054] Generally, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-image prediction or "intra-prediction") can predict the current BPU using pixels from one or more already coded neighboring BPUs within the same image. That is, the prediction criterion 224 in spatial prediction may include neighboring BPUs. Spatial prediction can reduce spatial redundancy inherent to an image. Temporal prediction (e.g., inter-image prediction or "inter-prediction") can predict the current BPU using regions from one or more already coded images. That is, the prediction criterion 224 in temporal prediction may include coded images. Temporal prediction can reduce temporal redundancy inherent to an image.

[0055] Referring to process 200B, in the forward path, the encoder performs prediction calculations in the spatial prediction stage 2042 and the temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder can perform intra-prediction. For the original BPU of the image being encoded, the prediction criterion 224 may include one or more neighboring BPUs that have been encoded (in the forward path) and reconstructed (in the reconstruction path) within the same image. The encoder can generate a prediction BPU 208 by extrapolating neighboring BPUs. Examples of extrapolation techniques include linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, the encoder can perform extrapolation at the pixel level, for example, by extrapolating the values ​​of the corresponding pixels for each pixel of the prediction BPU 208. The adjacent BPU used for extrapolation may be located in various directions relative to the original BPU, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., below left, below right, above left, or above right of the original BPU), or in any direction defined by the video coding standard used. For intra-prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the adjacent BPU used, the size of the adjacent BPU used, the extrapolation parameters, and the orientation of the adjacent BPU used relative to the original BPU.

[0056] As another example, in the time prediction stage 2044, the encoder can perform interpretation. For the original BPU of the current image, the prediction criterion 224 may include one or more images (referred to as "reference images") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference images may be encoded and reconstructed for each BPU. For example, the encoder may generate a reconstructed BPU by adding the reconstructed residual BPU 222 to the prediction BPU 208. Once all the reconstructed BPUs for the same image have been generated, the encoder can generate the reconstructed image as the reference image. The encoder can perform a "motion estimation" operation to search for a matching region within the range of the reference image (referred to as the "search window"). The position of the search window in the reference image may be determined based on the position of the original BPU in the current image. For example, the search window may be centered in the reference image at a position having the same coordinates as the original BPU in the current image and extended outward by a predetermined distance. When the encoder identifies a region similar to the original BPU within the search window (e.g., using a Pell recursive algorithm, a block matching algorithm, etc.), the encoder can determine such a region as a matching region. The matching region may have different dimensions from the original BPU (e.g., smaller than, equal to, larger than, or different in shape from the original BPU). Because the reference image and the current image are temporally separated in the timeline, the matching region can be considered to "move" to the position of the original BPU over time. The encoder may record the direction and distance of such movement as a "motion vector". If multiple reference images are used, the encoder can search for matching regions and determine the associated motion vector for each reference image. In some embodiments, the encoder can assign weights to the pixel values ​​of the matching regions in separate matching reference images.

[0057] Motion estimation can be used to identify various types of motion, such as translation, rotation, and zoom. In the case of interpretation, the prediction data 206 may include, for example, the location of the matching region (e.g., coordinates), the motion vector associated with the matching region, the number of reference images, and the weights associated with the reference images.

[0058] To generate a predicted BPU 208, the encoder can perform a β€œmotion compensation” operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on prediction data 206 (e.g., motion vectors) and prediction criteria 224. For example, the encoder can move the matching region of a reference image according to the motion vector. In the matching region, the encoder can now predict the original BPU of the image. If multiple reference images are used, the encoder can move the matching region of the reference images according to separate motion vectors and the average pixel values ​​of the matching regions. In some embodiments, if the encoder has assigned weights to the pixel values ​​of the matching regions of separate matching reference images, the encoder can add the weighted sum of the pixel values ​​of the moved matching regions.

[0059] In some embodiments, interpretation can be unidirectional or bidirectional. In unidirectional interpretation, one or more reference images in the same time direction relative to the current image can be used. In unidirectional interpretation, a reference image prior to the current image can be used. In bidirectional interpretation, one or more reference images in both time directions relative to the current image can be used.

[0060] Further reference to the forward path of process 200B, after the spatial prediction stage 2042 and the temporal prediction stage 2044, in the mode determination stage 230, the encoder can select a prediction mode for the current iteration of process 200B (e.g., either intra-prediction or inter-prediction). For example, the encoder can perform a rate-distortion optimization technique that selects a prediction mode that minimizes the value of the cost function, depending on the bitrate of the candidate prediction mode and the distortion of the reconstructed reference image under the candidate prediction mode. Depending on the selected prediction mode, the encoder can generate the corresponding prediction BPU 208 and prediction data 206.

[0061] In the reconstruction path of process 200B, if intra-prediction mode is selected in the forward path, after generating the prediction criterion 224 (e.g., the current BPU encoded and reconstructed within the current image), the encoder can directly supply the prediction criterion 224 to the spatial prediction stage 2042 for later use (e.g., extrapolation of the next BPU of the current image). If inter-prediction mode is selected in the forward path, after generating the prediction criterion 224 (e.g., the current image with all BPUs encoded and reconstructed), the encoder can supply the prediction criterion 224 to the loop filter stage 232, where the encoder can apply loop filters to the prediction criterion 224 to reduce or eliminate distortions (e.g., blocking artifacts) introduced by inter-prediction. In the loop filter stage 232, the encoder can apply various loop filtering techniques, such as deblocking, sample-adaptive offset, and adaptive loop filtering. The loop-filtered reference image may be stored in buffer 234 (or β€œDecoded Image Buffer”) for later use (e.g., as an interpretation reference image for future images in video sequence 202). The encoder may store one or more reference images used in the time prediction stage 2044 in buffer 234. In some embodiments, the encoder may encode the loop filter parameters (e.g., the loop filter intensity) along with the quantization transformation coefficients 216, prediction data 206, and other information in the binary coding stage 226.

[0062] In some embodiments, the input video sequence 202 is processed block by block according to the encoding process 200B. In VVC, the coding tree unit (CTU) is the largest block unit and can be up to 128 Γ— 128 chroma samples (and corresponding chroma samples depending on the chroma format). The CTU can be further partitioned into coding units (CU) using a quadtree, binary tree, or ternary tree. At the leaf nodes of the partition structure, coding information is sent, such as the coding mode (intra-mode or inter-mode), motion information if intercoded (reference index, motion vector difference, etc.), and quantization conversion coefficients 216. When intra-prediction (also called spatial prediction) is used, spatially adjacent samples are used to predict the current block. When inter-prediction (also called temporal prediction or motion-compensated prediction) is used, samples from an already coded image called a reference image are used to predict the current block. Inter-prediction may use single or bi-prediction. In single prediction, only one motion vector pointing to one reference image is used to generate the prediction signal for the current block. In dual prediction, two motion vectors, each pointing to its own reference image, are used to generate a prediction signal for the current block. The motion vectors and reference indices are sent to the decoder to identify where the prediction signal for the current block came from. After intra or interpretation, in the mode determination stage 230, the optimal prediction mode for the current block is selected, for example, based on a rate-distortion optimization method. Based on the optimal prediction mode, a prediction BPU 208 is generated and subtracted from the input video block.

[0063] Referring further to Figure 2B, the predicted residual BPU 210 is sent to the transformation stage 212 and quantization stage 214 to generate quantization transformation coefficients 216. The quantization transformation coefficients 216 are then dequantized in the dequantization stage 218 and inverse transformed in the inverse transformation stage 220 to obtain the reconstructed residual BPU 222. The predicted BPU 208 and the reconstructed residual BPU 222 are added together before loop filtering to form the prediction criterion 224. Loop filtering is used to provide reference samples for intra-prediction. In the loop filtering stage 232, loop filtering such as deblocking, sample adaptive offset (SAO), and adaptive loop filtering (ALF) may be applied to the prediction criterion 224 to form reconstructed blocks stored in buffer 234 and used to provide reference samples for intra-prediction. The coding information generated in the mode determination stage 230 (coding mode (intra or inter predictive), intra predictive mode, motion information, quantization residual coefficients, etc.) is sent to the binary coding stage 226 to further reduce the bitrate before being packed into the output video bitstream 228.

[0064] Figure 3A shows a schematic diagram of an exemplary decoding process 300A according to an embodiment of the present disclosure. For example, the decoding process 300A may be carried out by a decoder such as the image / video decoder 144 in Figure 1. Process 300A may be a decompression process corresponding to the compression process 200A in Figure 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. The decoder (e.g., the image / video decoder 144 in Figure 1) can decode the video bitstream 228 into a video stream 304 according to process 300A. The video stream 304 may be very similar to the video sequence 202. However, due to loss of information in the compression and decompression processes (e.g., the quantization stage 214 in Figures 2A-2B), the video stream 304 is generally not identical to the video sequence 202. Similar to processes 200A and 200B in Figures 2A-2B, the decoder can perform process 300A at the level of the basic processing unit (BPU) for each pixel encoded in the video bitstream 228. For example, the decoder can perform process 300A in an iterative manner, in which case the decoder can decode the BPU in one iteration of process 300A. In some embodiments, the decoder can perform process 300A in parallel for each region of the image encoded in the video bitstream 228.

[0065] As shown in Figure 3A, the decoder can supply a portion of the video bitstream 228 associated with the basic processing unit of the encoded image (called the "encoded BPU") to the binary decoding stage 302. In the binary decoding stage 302, the decoder can decode that portion into prediction data 206 and quantization conversion coefficients 216. The decoder can supply the quantization conversion coefficients 216 to the inverse quantization stage 218 and the inverse transformation stage 220 to generate the reconstructed residual BPU 222. The decoder can supply the prediction data 206 to the prediction stage 204 to generate the prediction BPU 208. The decoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate the prediction criterion 224. In some embodiments, the prediction criterion 224 may be stored in a buffer (e.g., a decoded image buffer in computer memory). The decoder can supply the prediction criterion 224 to the prediction stage 204 for performing the prediction calculation in the next iteration of process 300A.

[0066] The decoder can iteratively perform process 300A to decode each encoding BPU of the encoded image and generate a prediction criterion 224 for encoding the next encoding BPU of the encoded image. After decoding all encoding BPUs of the encoded image, the decoder can output the image to the video stream 304 for display and proceed to decode the next encoded image in the video bitstream 228.

[0067] In the binary decoding stage 302, the decoder can perform the inverse operation of the binary coding technique used by the encoder (e.g., entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or other lossless compression algorithms). In some embodiments, in addition to the predicted data 206 and quantization conversion coefficients 216, the decoder can decode other information in the binary decoding stage 302, such as the prediction mode, parameters of the prediction operation, conversion type, parameters of the quantization process (e.g., quantization parameters), and encoder control parameters (e.g., bitrate control parameters). In some embodiments, if the video bitstream 228 is transmitted in packets over the network, the decoder can depacketize the video bitstream 228 before supplying it to the binary decoding stage 302.

[0068] Figure 3B shows a schematic diagram of another exemplary decoding process 300B according to an embodiment of the present disclosure. For example, decoding process 300B may be carried out by a decoder such as the image / video decoder 144 in Figure 1. Process 300B may be modified from process 300A. For example, process 300B may be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 300A, process 300B further divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and further includes a loop filter stage 232 and a buffer 234.

[0069] In process 300B, the prediction data 206 decoded by the decoder from the binary decoding stage 302 for the encoding base processing unit ("current BPU") of the encoded image being decoded ("current image") may contain various types of data depending on the prediction mode used by the encoder to encode the current BPU. For example, if intra-prediction is used by the encoder to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra-prediction, parameters of the intra-prediction operation, etc. Parameters of the intra-prediction operation may include, for example, the location (e.g., coordinates) of one or more adjacent BPUs used as reference, the size of the adjacent BPUs, extrapolation parameters, the orientation of the adjacent BPUs relative to the original BPU, etc. As another example, if inter-prediction is used by the encoder to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter-prediction, parameters of the inter-prediction operation, etc. Parameters for the interpretation calculation may include, for example, the number of reference images currently associated with the BPU, the weights associated with each reference image, the locations (e.g., coordinates) of one or more matching regions within separate reference images, and one or more motion vectors associated with each matching region.

[0070] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. Details of performing such spatial or temporal predictions are shown in Figure 2B and will not be repeated below. After performing such spatial or temporal predictions, the decoder can generate a prediction BPU 208. The decoder can then generate a prediction criterion 224 by adding the prediction BPU 208 and the reconstructed residual BPU 222, as shown in Figure 3A.

[0071] In process 300B, the decoder can supply the prediction criterion 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing prediction calculations in the next iteration of process 300B. For example, if the current BPU is decoded using intra-prediction in the spatial prediction stage 2042, after generating the prediction criterion 224 (e.g., decoded current BPU), the decoder can supply the prediction criterion 224 directly to the spatial prediction stage 2042 for later use (e.g., extrapolation of the next BPU of the current image). If the current BPU is decoded using inter-prediction in the temporal prediction stage 2044, after generating the prediction criterion 224 (e.g., reference image decoded by all BPUs), the encoder can supply the prediction criterion 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can apply a loop filter to the prediction criterion 224 as shown in Figure 2B. The loop-filtered reference image may be stored in a buffer 234 (e.g., a "decoded image buffer") in preparation for later use (e.g., use as an inter-prediction reference image for a future encoded image of the video bitstream 228). The decoder may store one or more reference images used in the time prediction stage 2044 in buffer 234. In some embodiments, if the prediction mode indicator of the prediction data 206 indicates that inter-prediction was used to encode the BPU, the prediction data may further include loop filter parameters (e.g., loop filter strength).

[0072] Referring back to Figure 1, the image / video preprocessor 122, image / video encoder 124, and image / video decoder 144 may each be implemented as any suitable hardware, software, or combination thereof. Figure 4 is a block diagram of an exemplary apparatus 400 for processing image data according to an embodiment of the present disclosure. For example, apparatus 400 may be a preprocessor, an encoder, or a decoder. As shown in Figure 4, apparatus 400 may include a processor 402. When the processor 402 executes instructions described herein, apparatus 400 may become a dedicated machine for preprocessing, encoding, and / or decoding image data. The processor 402 may be any type of circuit capable of manipulating or processing information. For example, the processor 402 may include any combination of any number of central processing units (or "CPUs"), graphics processing units (or "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), generic array logic (GALs), composite programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), systems on a chip (SoCs), application-specific integrated circuits (ASICs), and so on. In some embodiments, the processor 402 may also be a set of processors grouped as a single logical component. For example, as shown in Figure 4, the processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.

[0073] The device 400 may also include a memory 404 configured to store data (e.g., a collection of instructions, computer code, intermediate data, etc.). For example, as shown in Figure 4, the stored data may include program instructions (e.g., program instructions for implementing each stage in processes 200A, 200B, 300A, or 300B) and processing data (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 can access the processing program instructions and data (e.g., via the bus 410) and execute the program instructions to perform arithmetic or operations on the processing data. The memory 404 may include a high-speed random-access storage device or a non-volatile storage device. In some embodiments, the memory 404 may include any number of random-access memories (RAM), read-only memories (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, compact flash (CF) cards, and any combination thereof. Memory 404 may also be a memory group (not shown in Figure 4) that is grouped as a single logical component.

[0074] Bus 410 may be a communication device that transfers data between components within the device 400, such as an internal bus (e.g., a CPU-memory bus) or an external bus (e.g., a universal serial bus port, a peripheral component interconnection express port).

[0075] To facilitate explanation without creating ambiguity, the processor 402 and other data processing circuits are collectively referred to as the β€œdata processing circuits” in this disclosure. The data processing circuits may be implemented as hardware as a whole, or as a combination of software, hardware, or firmware. The data processing circuits may also be a single, independent module, or may be incorporated whole or in part into any other component of the device 400.

[0076] The device 400 may further include a network interface 406 for providing wired or wireless communication to a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.). In some embodiments, the network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near-field communication (NFC) adapters, cellular network chips, and the like.

[0077] In some embodiments, the apparatus 400 may further include a peripheral interface 408 for providing connectivity to one or more peripheral devices. As shown in Figure 4, peripheral devices may include, but are not limited to, cursor control devices (e.g., mouse, touchpad, or touchscreen), keyboards, displays (e.g., cathode ray tube displays, liquid crystal displays, or light-emitting diode displays), video input devices (e.g., cameras or input interfaces coupled to a video archive), and the like.

[0078] It should be noted that the video codec (for example, the codec that performs processes 200A, 200B, 300A, or 300B) may be implemented as any combination of any software modules or hardware modules within the device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more software modules of the device 400, such as program instructions that can be loaded into memory 404. As another example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of the device 400, such as dedicated data processing circuits (e.g., FPGA, ASIC, NPU, etc.).

[0079] SEI messages are intended to be transmitted within the coded video bitstream in the manner specified by the video coding specification, or by other means defined by the specifications of the system using such coded video bitstream. SEI messages may contain various types of data indicating the timing of video images, or describing various characteristics of the coded video and how they should be used or extended. SEI messages are also defined as containing arbitrary user-defined data. SEI messages do not affect the core decoding process, but they can indicate how the video should be post-processed or displayed.

[0080] The emergence of deep generative models, including Variation Autoencoders (VAEs) and Generative Adversarial Networks (GANs), has enabled promising performance improvements in face video compression. For example, X2Face can be used to control face generation via image, audio, and pose codes. Other realistic neural talking head models can be used via fusion shot adversarial learning. For video-to-video synthesis tasks, Face-vidtovid (also known as "Face_vid2vid") can be used. Furthermore, schemes that leverage compact 3D keypoint representations to drive generative models for rendering target frames can also be employed. Additionally, motion-aware video chat systems based on FOMM can be used. VSBNet, which reconstructs the original frame from landmarks using adversarial learning, can also be used. Finally, an end-to-end talking head video compression framework based on Compact Feature Learning (CFTE), designed for highly efficient talking face video compression in ultra-low bandwidth scenarios, is also available. The CFTE scheme leverages compact feature representations to compensate for temporal changes and reconstructs target face video frames in an end-to-end manner. Furthermore, the CFTE scheme can be incorporated into a video coding framework with a rate distortion targeting function. Additionally, facial semantics can be used via 3DMM templates to characterize facial images and impart facial manipulation capabilities to facial image coding.

[0081] Figure 5 is a flowchart of an exemplary method 500 for processing video based on generated face video SEI messages, according to some embodiments of the present disclosure. Method 500 describes the general syntax structure of a generated face video SEI message and the order in which several syntax elements are generated. As shown in Figure 5, Method 500 may include steps 502 to 508.

[0082] In step 502, the encoder (e.g., the image / video encoder 124 in Figure 1 or the device 400 in Figure 4) may generate an identification number indicator within the SEI message and signal it to the decoder (e.g., the image / video decoder 144 in Figure 1 or the device 400 in Figure 4). The signaled identification number indicator can be used to identify the SEI message and also to indicate whether the current generated face video SEI message matches the generation network in the decoder.

[0083] In step 504, if the identification number indicator indicates that a face image generation compression scheme is being used, the encoder may generate multiple (e.g., three) face information type presence flags and other parameters, signaling them to indicate the presence and length of specific syntax elements associated with the face image generation compression scheme.

[0084] In step 506, the encoder may generate face parameter information corresponding to the face presence flag generated in step 504 and signal it to the decoder.

[0085] In step 508, the decoder may use the signaled corresponding face parameter information to reconstruct or retarget relevant face images based on the base image.

[0086] Figure 6 is another flowchart of an exemplary method 600 for processing video based on generated face video SEI messages, according to some embodiments of the present disclosure. Method 600 describes the general syntax structure of a generated face video SEI message and the order in which several syntax elements are generated. As shown in Figure 6, Method 600 may include steps 602-610.

[0087] In step 602, the encoder (e.g., the image / video encoder 124 in Figure 1 or the device 400 in Figure 4) may generate an identification number indicator or other indicators for identifying SEI messages. For example, a key indicator may be generated and signaled to identify the analysis network used to generate the syntax elements of the current generated face video SEI message. The value of the key indicator (e.g., gfv_key, described later) may be used to determine whether the analysis network in the encoder matches the generation network in the decoder (e.g., the image / video decoder 144 in Figure 1 or the device 400 in Figure 4). Additionally, an image sequence count specifies the display sequence count modulo 1 << 31 for the images currently generated in the SEI message. In some embodiments, the analysis network in the encoder is also expected to be the generation network in the decoder for decoding the bitstream sent by the encoder, and is therefore also called the symmetric (neural) network.

[0088] In step 604, the encoder may determine whether to generate a parameter existence flag for the current parameter group (mode) and signal it. Only if the flag is true can the encoder generate the next parameter and signal it to the decoder in a subsequent step. If the flag is false, the parameter for the current group (mode) is not generated and is not signaled. As shown in Figure 6, step 604 may be implemented in a polling manner for each mode until a true flag is found.

[0089] If, in step 606, it is determined to signal the presence of a parameter, the encoder may first generate and signal the number of parameter types when signaling the current parameter group.

[0090] In step 608, the encoder can generate and signal detailed parameters for each parameter type accordingly.

[0091] In step 610, the decoder may, after receiving the SEI message, decode the parameters and reconstruct the face image. Thus, the face image can be generated using the decoded parameters based on the base image.

[0092] In some embodiments, the base image is an image that can be coded using conventional coding methods. For example, it can be coded in H.264, H.265, or H.266. The base image provides texture information of a human face, and the generation network can use this image to generate additional face images, relying on parameters that indicate the differences between the generated image and the base image. The base image is typically the first image in a video sequence, and subsequent images can be generated by the neural network. The parameters required to generate the next image are signaled within SEI messages. A single SEI message contains the parameters required to generate one or more images.

[0093] Figure 7A is a flowchart of an exemplary method 700 for generating a facial image according to some embodiments of the present disclosure. As shown in Figure 7A, the method 700 may include steps 702-710 which can be implemented by a decoder (e.g., the image / video decoder 144 in Figure 1, or the device 400 in Figure 4).

[0094] In step 702, the decoder may receive a bitstream from an encoder (e.g., the image / video encoder 124 in Figure 1 or the device 400 in Figure 4) or a content distribution operator. As is understood, the bitstream may contain coded information for a series of images. In some embodiments, an SEI message may be signaled before a group of images (GOP) that can be decoded along with the pixel content of the images.

[0095] In step 704, the decoder may decode the encoded information of the bitstream to obtain the base image and SEI message. As described above, the base image is a reference image that can provide texture information of a human face. The generative network can refer to the base image and generate a current face image based on parameters that indicate the difference between the current face image and the base image. These parameters can be communicated in the SEI message.

[0096] In step 706, the decoder may determine whether the SEI message is applicable to the neural network for generating the face image. In some embodiments, the SEI message may be applied to the decoder's neural network for generating the face image if it satisfies two criteria: (1) the SEI message contains information about the face image to be reconstructed, and (2) the information in the SEI message is available to the decoder.

[0097] In step 708, if the decoder determines in step 706 that the SEI message is to be applied to a neural network for generating a face image, it may determine the mode used to encode the face image and the corresponding face information parameters based on the SEI message.

[0098] In step 710, the decoder may generate a face image based on the base image and face information parameters using a neural network. Figure 7B is a schematic diagram showing a neural network for generating a face image according to some embodiments of the present disclosure. As shown in Figure 7B, the base image and face information parameters (e.g., face landmarks) can be input to the decoder's generative neural network. The generative neural network produces a reconstructed face image as output. As is understood, the face image is encoded as face information parameters relative to the base image. The proposed method of exchanging only encoded face information parameters compared to transmitting the entire face image has proven advantageous in terms of bandwidth.

[0099] Table 1 below summarizes face representations for a generated face video compression algorithm. In this disclosure, the algorithm is also referred to as a mode, method, or technique. In particular, face images exhibit strong statistical regularities and can be leanly characterized using 2D landmarks, 2D keypoints, region matrices, 3D keypoints, compact feature matrices, or face semantics. Such face description strategies lead to reduced coding bitrate and improved coding efficiency, and are therefore applicable to video conferencing and live entertainment. [Table 1]

[0100] At the 29th Joint Video Expert Team (JVET) meeting, the generated facial image SEI message was proposed. The generated facial image SEI message can represent head posture and facial expression states using a set of facial semantic information. The initially proposed SEI message considered only one facial representation in facial generation compression. However, facial images can be described by variations of feature structures with strong prior distribution, such as landmarks, 2D keypoints, region matrices, 3D keypoints, compact feature matrices, facial semantics, and other formats. These facial representations can provide a high degree of freedom in the syntax design and semantic description of the SEI message for facial image compression. It is desirable for the SEI message of the VVC standard to consider various facial representations in the task of facial image compression.

[0101] Currently, in addition to facial video reconstruction, there is a demand for more common use cases in facial video communication, such as facial video retargeting or animation. For example, as metaverse activity becomes more popular, there may be a need to transfer the movement of a real-world face into a virtual metaverse world and represent it with another person. Furthermore, facial video reconstruction is expected to become more realistic, and corresponding retargeting is anticipated. Establishing SEI messages is crucial for compressing facial video, as it can incorporate various forms of data, such as specifying the timing of video frames, characterizing encoded video, and elucidating potential uses or extensibility. This ensures that reconstructed facial video can be efficiently post-processed and displayed in a user-friendly manner to meet the actual needs of the user.

[0102] To address at least one of the above problems, this disclosure proposes a novel SEI message called a generated facial video SEI message. The proposed SEI is compatible with different facial representations such as 2D keypoints, 2D landmarks, 3D keypoints, or facial semantics, and can be used to reconstruct high-quality talking face video at ultra-low bitrates or to manipulate talking face video for personalized characterization. Thus, the proposed generated facial video SEI message is applicable to video conferencing, live entertainment, facial animation, and metaverse-related functions.

[0103] Table 2 below shows an exemplary syntax for an SEI message. The process for implementing an SEI message may be carried out as follows. As shown in Table 2, an identification number indicator gfv_id may be generated by the encoder and signaled to the decoder. In some embodiments, gfv_id may be used to indicate whether the SEI is used for encoding a face image. In some embodiments, a conventional SEI message may be signaled, and a generated face image SEI message, as shown in Table 2, may also be signaled, although it is not used for encoding a face image. Therefore, gfv_id may be used to distinguish between a conventional SEI message and a generated face image SEI message. Referring back to Figure 7A, in step 706, the decoder may determine, based on the identification number indicator, whether the SEI message is used for encoding a face image. In some embodiments, the identification number indicator may be used to indicate a generated face image filter.

[0104] Referring further to Table 2, the SEI message may include the mode indicator gfv_feature_mode. The mode used for encoding the face image is determined by the decoder in step 708 based on gfv_feature_mode. Under different gfv_feature_modes (gfv_feature_mode==0, gfv_feature_mode==1, gfv_feature_mode==0, etc.), parameter indicators for conveying face information parameters may be included in the SEI message. For example, in the case of gfv_feature_mode==0, which means that 2D face landmarks are selected as the feature mode in the SEI message, the parameter indicators (1) 2d_landmark_quantization_factor, (2) 2d_landmark_num, and (3) x[i], y[i] may be included in the SEI message and can be decoded as face information parameters.

[0105] To summarize, in some embodiments, the SEI messages in Table 2 can be utilized by the following operations: First, the generated face image SEI message includes an identification number gfv_id which may be used to identify the generated face image SEI message. Second, a mode indicator gfv_feature_mode is signaled within the SEI message to determine the mode to be used. Next, depending on the mode to be used, the corresponding parameters are signaled. Third, after receiving and decoding the SEI message, the parameters signaled within the SEI message and the previously decoded base image can be used as input to reconstruct the face image in a high-quality or user-friendly manner via the generative adversarial network's generation capabilities. [Table 2-1] [Table 2-2] [Table 2-3]

[0106] The semantics of the above syntax are as follows:

[0107] In some embodiments, the SEI message may include face parameters of different feature representations, such as 2D keypoints, 2D landmarks, 3D keypoints, or face semantics, which may be used for face generation compression. These face representations may be classified into different types, and gfv_feature_mode may be signaled to determine which type of face representation is being utilized. Based on the base image, the face parameters in the SEI message may be used to reconstruct the face image, with each SEI message being used to generate one face image.

[0108] In some embodiments, gfv_id includes an identification number which may be used to identify the generated facial video SEI message. The value of gfv_id is between 0 and 2. 32 It can be within the range of -2 or less.

[0109] In some embodiments, the value of gfv_feature_mode may be in the range of 0 to 5. Feature code values ​​between 6 and 128 may be reserved for future use by ITU-T|ISO / IEC and may not be present in the bitstream. The decoder may ignore GFV SEI messages in which gfv_feature_mode is in the range of 6 to 128. Feature code values ​​greater than 1023 may not be present in the bitstream and are not reserved for future use. [Table 3]

[0110] [1] Benjamin Bross, Ye-Kui Wang, Yan Ye, Shan Liu, Jianle Chen, Gary J Sullivan, and Jens-Rainer Ohm,β€³ Overview of the versatile video coding(vvc)standard and its applications, β€³ IEEE Transactions on Circuits and Systems for Video Technology,2021。 [2]G.J.Sullivan,J.R.Ohm,W.J.Han,T.Wiegand, β€³ Overview of the high efficiency video coding (HEVC) standard, β€³ IEEE Trans.Circuits and Systems for Video Technology,vol.22,no.12,pp.1649-1668,Dec.2012。 [3]Ian Goodfellow,Jean Pouget-Abadie,Mehdi Mirza,Bing Xu,David Warde Farley,Sherjil Ozair,Aaron Courville,and Yoshua Bengio, β€³ Generative adversarial nets, β€³ Advances in neural information processing systems,vol.27,2014。 [4]Aliaksandr Siarohin,St_ephane Lathuili_ere,Sergey Tulyakov,Elisa Ricci,and Nicu Sebe, β€³ First order motion model for image animation, β€³ Advances in Neural Information Processing Systems,vol.32,pp.7137-7147,2019。 [5]Ting-Chun Wang,Arun Mallya,and Ming-Yu Liu, β€³One-shot free-view neural talking-head synthesis for video conferencing, β€³ in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 10039-10049. [6]Bolin Chen, Zhao Wang, Bin Li, Rongqun Lin, Shiqi Wang, and Yan Ye, β€³ Beyond key-point coding:Temporal evolution inference with compact feature representation for talking face video compression, β€³ in Proceedings of the IEEE Data Compression Conference,2022. Regarding the definition of gfv_feature_mode in Table 3, if there are facial feature modes other than the given six types, gfv_feature_mode can be further extended. Also, if there are other, or future, generative compression models that can show better rate distortion performance than the representative models above, these existing models can be replaced.

[0111] The definitions and possible values ​​of the parameter indicators in each mode are shown below.

[0112] In some embodiments, when a 2D facial landmark is determined as the mode used for encoding the facial image, the parameter indicator indicates that the facial representation of the 2D facial landmark can be generated and signaled by the encoder and then received by the decoder.

[0113] 2d_landmark_quantization_factor specifies the quantization factor for processing face semantic parameters (i.e., x[i] and y[i]) in mode 0 (i.e., gfv_feature_mode is equal to 0). The parameter values ​​used for face generation are equal to the values ​​of the corresponding syntax elements divided by 2d_landmark_quantization_factor.

[0114] 2d_landmark_num specifies the number of face landmarks in 2D coordinates.

[0115] x[i] specifies the quantized x-axis value of the i-th point in 2D coordinates.

[0116] y[i] specifies the quantized y-axis value of the i-th point in 2D coordinates.

[0117] In some embodiments, when 2D facial keypoints are determined as the mode used for encoding facial images, the parameter indicator indicates that the facial representation of the 2D facial keypoints can be generated and signaled by the encoder and then received by the decoder.

[0118] 2d_keypoint_quantization_factor specifies the quantization factor for processing face semantic parameters (i.e., x[i], y[i], and affine_transmation_matrix[i][j][k]) in mode 1 (i.e., gfv_feature_mode is equal to 1). The parameter values ​​used for face generation are equal to the values ​​of the corresponding syntax elements divided by 2d_keypoint_quantization_factor.

[0119] 2d_keypoint_num specifies the number of face keypoints in 2D coordinates.

[0120] If is_affine_transformation_matrix_flag is equal to 1, it indicates that the SEI message contains affine transformation parameters. If is_affine_transformation_matrix_flag is equal to 0, it indicates that the SEI message does not carry affine transformation matrix parameters.

[0121] `affine_transformation_matrix[i][j][k]` specifies the quantization element values ​​from the corresponding affine transformation matrix.

[0122] In some embodiments, when a consistency region is determined as the mode used for encoding the facial image, the parameter indicator indicates that the facial representation of the consistency region can be generated and signaled by the encoder and subsequently received by the decoder.

[0123] The `region_quantization_factor` specifies the quantization factor for processing face semantic parameters (i.e., x[i], y[i], affine_transmation_matrix[i][j][k], and covariance_matrix[i][m][n]) in mode 2 (i.e., gfv_feature_mode is equal to 2). The parameter values ​​used for face generation are equal to the values ​​of the corresponding syntax elements divided by the `region_quantization_factor`.

[0124] `region_keypoint_num` specifies the number of face keypoints in 2D coordinates.

[0125] If is_covariance_matrix_flag is equal to 1, it indicates that the SEI message includes covariance matrix parameters. If is_covariance_matrix_flag is equal to 0, it indicates that the SEI message does not carry covariance matrix parameters.

[0126] covariance_matrix[i][m][n] specifies the quantized element values ​​of the corresponding covariance matrix.

[0127] In some embodiments, when 3D facial keypoints are determined as the mode used for encoding facial images, the parameter indicator indicates that the facial representation of the 3D facial keypoints can be generated and signaled by the encoder and then received by the decoder.

[0128] 3d_keypoint_quantization_factor specifies the quantization factor for processing face semantic parameters in mode 3 (i.e., gfv_feature_mode is equal to 3). The parameter values ​​used for face generation are equal to the values ​​of the corresponding syntax elements divided by 3d_keypoint_quantization_factor.

[0129] 3d_keypoint_num specifies the number of face keypoints in 3D coordinates.

[0130] If is_rotation_matrix_flag is equal to 1, it indicates that the SEI message includes rotation matrix parameters. If is_rotation_matrix_flag is equal to 0, it indicates that the SEI message does not carry rotation matrix parameters.

[0131] If is_translation_matrix_flag is equal to 1, it indicates that the SEI message contains translation matrix parameters. If is_translation_matrix_flag is equal to 0, it indicates that the SEI message does not carry translation matrix parameters.

[0132] z[i] specifies the quantized z-axis value of the i-th point in 3D coordinates.

[0133] `rotation_matrix[j][k]` specifies the quantized element values ​​of the corresponding rotation matrix.

[0134] `translation_matrix[l]` specifies the quantized element values ​​of the corresponding translation matrix.

[0135] In some embodiments, when compact features are determined as the mode used for encoding facial images, the parameter indicator indicates that the facial representation of compact features can be generated and signaled by an encoder and then received by a decoder.

[0136] The compact_feature_quantization_factor specifies the quantization factor for processing face semantic parameters (i.e., compact_feature_matrix_element[k][m][n]) in mode 4 (i.e., gfv_feature_mode is equal to 4). The parameter values ​​used for face generation are equal to the values ​​of the corresponding syntax elements divided by the compact_feature_quantization_factor.

[0137] `matrix_channel` specifies the channel for compact features, and `matrix_channel` can be 1 or greater.

[0138] `matrix_width` specifies the width (number of rows) of the compact feature, and `matrix_width` can be 1 or greater.

[0139] `matrix_height` specifies the height (number of columns) of a compact feature, and `matrix_height` can be 1 or greater.

[0140] compact_feature_matrix_element specifies the quantized element value from the corresponding face compact feature.

[0141] In some embodiments, when facial semantics is determined as the mode used for encoding facial images, the parameter indicator indicates that the facial representation of the facial semantics can be generated and signaled by the encoder and then received by the decoder.

[0142] The semantic_quantization_factor specifies the quantization factor for processing face semantic parameters (i.e., semantic_element[l][y][z]) in mode 5 (i.e., gfv_feature_mode is equal to 5). The parameter values ​​used for face generation are equal to the values ​​of the corresponding syntax elements divided by the semantic_quantization_factor.

[0143] `semantic_type_num` specifies the number of face semantic types in the SEI message. The value of `semantic_type_num` is between 0 and 2. 6 It may be within the following range. In some embodiments, in an SEI message, the face semantic type may be classified into mouth parameters, eye parameters, head rotation parameters, head translation parameters, and head position parameters.

[0144] The semantic_idx contains an identification number related to the face semantics that may be used in SEI messages. [Table 4]

[0145] Regarding the definition of semantic_idx in Table 4, semantic_idx can be further extended if there are other facial semantic representations besides the given semantic type.

[0146] semantic_width[l] specifies the width (number of rows) of the l-th semantic type, and semantic_width[l] can be 1 or greater.

[0147] semantic_height[l] specifies the height (number of columns) of the l-th semantic type, and semantic_height[l] can be 1 or greater.

[0148] semantic_element[l][y][z] specifies the quantized element value from the corresponding face semantics.

[0149] `data_quantization_factor` specifies a quantization factor for processing face semantic parameters (i.e., `data_element[i]`) in other latent modes. The parameter values ​​used for face generation are equal to the values ​​of the corresponding syntax elements divided by `data_quantization_factor`.

[0150] `data_length` specifies the length of a feature parameter that may be used in other `gfv_feature_modes` besides the given mode.

[0151] `data_element[i]` specifies the quantized value of the i-th feature in `data_length`.

[0152] The syntax of the generated facial image SEI messages shown in Table 2 may be encoded and signaled within a bitstream and further decoded by a decoder. Figure 8 is a schematic diagram showing an exemplary SEI message 800 according to some embodiments of the present disclosure. As shown in Figure 8, the SEI message 800 may include an identification number indicator (gfv_id) 801, a mode indicator (gfv_feature_mode) 802, a parameter indicator 803, and the like. The corresponding bit lengths of these parameters are also shown in Figure 8 as an example. As is understood, the bit lengths of these parameters may be designed in other ways. The parameter indicator 803 is signaled in accordance with the mode indicator 802. Other parameters are not shown in Figure 8 to avoid ambiguity.

[0153] Table 5 shows an exemplary syntax for an SEI message. As shown in Table 5, an identification number indicator gfv_id may be generated by the encoder and signaled to the decoder. In some embodiments, gfv_id may be used to indicate whether the SEI is used to encode a face image. As mentioned above, gfv_id may be used to distinguish between a conventional SEI message and a generated face image SEI message. Referring back to Figure 7A, the decoder may, in step 706, determine based on the identification number indicator whether the SEI message is used to encode a face image. In some embodiments, the identification number indicator may be used to indicate a generated face image filter.

[0154] Referring further to Table 5, an SEI message may include at least one parameter presence indicator (e.g., coordinate_present_flag, matrix_present_flag, or semantic_present_flag). The mode used for encoding the face image is determined by the decoder in step 708 based on the parameter presence indicator. In the different modes indicated by the parameter presence indicator, parameter indicators for conveying face information parameters may be included in the SEI message. For example, in the case of coordinate_present_flag==1 (whereas matrix_present_flag==0 and semantic_present_flag==0), meaning that the SEI message includes coordinate parameters, the parameter indicators (1) coordinate_quantization_factor, (2) is_3D_coordinate_flag, (3) coordinate_point_num, and (4) x[i], y[i], z[i]) may be included in the SEI message and can be decoded as face information parameters.

[0155] In some embodiments, the SEI messages in Table 5 may be utilized by the following actions: Firstly, the generated face image SEI message includes an identification number gfv_id which may be used to identify the generated face image SEI message. Secondly, it signals parameter presence flags coordinate_present_flag, matrix_present_flag, and semantic_present_flag for different parameter types. If the parameter presence flag indicates that a parameter of the current type exists, the corresponding parameter is signaled. Otherwise, signaling for the parameter of the current type is skipped. In this example, parameters are categorized into three types: keypoint, matrix, and semantics. Semantic types include parameters that indicate the state, movement, or position of the mouth, eyes, or head. Matrix types include affine translational matrices, covariance matrices, translational matrices, rotational matrices, and compact feature matrices. Thirdly, after receiving and decoding the SEI message, the parameters signaled within the SEI and the previously decoded base image can be used as input to reconstruct the facial image in a high-quality or user-friendly manner through the generative adversarial network's generation capabilities. [Table 5-1] [Table 5-2]

[0156] The semantics of the above syntax are as follows:

[0157] In some embodiments, the SEI message may include face parameters for different feature representations, such as 2D keypoints, 2D landmarks, 3D keypoints, or face semantics, and these feature representations may be used for face generation compression. These face representations can be classified into three types: coordinate parameters, matrix parameters, and semantic parameters. Based on the base image, the face parameters in the SEI message may be used to reconstruct the face image, with each SEI message being used to generate one face image.

[0158] In some embodiments, gfv_id includes an identification number that can be used to identify the generated facial video SEI message. The value of gfv_id is between 0 and 2. 32 It can be within the range of -2 or less.

[0159] The definitions and possible values ​​of the parameter indicators in each mode are shown below.

[0160] In some embodiments, if coordinate_present_flag is equal to 1, it indicates that the SEI message contains coordinate parameters, whereas if coordinate_present_flag is equal to 0, it indicates that the SEI message does not carry coordinate parameters. If it is determined that keypoint coordinate information will be used to encode the face image, the parameter indicator indicates that the keypoint coordinate information may be generated and signaled by the encoder and received by the decoder.

[0161] `coordinate_quantization_factor` specifies the quantization factor for processing face coordinate parameters (i.e., x[i], y[i], and z[i]). The parameter values ​​used for face generation are equal to the values ​​of the corresponding syntax elements divided by `coordinate_quantization_factor`.

[0162] If is_3D_coordinate_flag is equal to 1, it indicates that the coordinate parameters belong to a 3D space and can be represented as (x[i], y[i], z[i]). If is_3D_coordinate_flag is equal to 0, it indicates that the coordinate parameters belong to a 2D space and can be represented as (x[i], y[i]).

[0163] `coordinate_point_num` specifies the number of face coordinate parameter sets included in the SEI message. Each parameter set represents a single point on the coordinate system. The value of `coordinate_point_num` is between 0 and 2. 10 It may be within the following range.

[0164] x[i] specifies the quantized x-axis value of the i-th keypoint.

[0165] y[i] specifies the quantized y-axis value of the i-th keypoint.

[0166] z[i] specifies the quantized z-axis value of the i-th keypoint.

[0167] In some embodiments, if matrix_present_flag is equal to 1, it indicates that the SEI message contains matrix parameters, whereas if matrix_present_flag is equal to 0, it indicates that the SEI message does not carry matrix parameters. If it is determined that matrix parameters are used to encode the face image, the parameter indicator indicates that the matrix parameters may be generated and signaled by the encoder and received by the decoder.

[0168] `matrix_quantization_factor` specifies the quantization factor for processing the face matrix parameters (i.e., `matrix_element[j][k][m][n]`). The parameter values ​​used for face generation are equal to the values ​​of the corresponding syntax elements divided by `matrix_quantization_factor`.

[0169] `matrix_type_num` specifies the number of face matrix types in the SEI message. The value of `matrix_type_num` is between 0 and 2. 6 The following ranges are possible. In some embodiments, face matrix types can be classified into affine translational matrices, covariance matrices, rotational matrices, translational matrices, and compact feature matrices. These matrices can also be further distinguished by whether or not they are associated with points in coordinates.

[0170] matrix_idx contains an identification number for the face matrix that may be used in SEI messages. [Table 6]

[0171] Regarding the definition of matrix_idx in Table 6, matrix_idx can be further extended if there is a face matrix representation different from the given matrix type.

[0172] If matrix_num_equal_to_coordinate_point_flag[j] is equal to 1, it indicates that the number of face matrices is equal to coordinate_point_num.

[0173] matrix_num[j] specifies the number of face matrices if coordinate_present_flag or matrix_num_equal_to_coordinate_point_flag[j] does not exist. Otherwise, if both of these flags exist, the number of face matrices is equal to coordinate_point_num.

[0174] matrix_width[j] specifies the width (number of rows) of the j-th face matrix, and matrix_width[j] can be 1 or greater.

[0175] matrix_height[j] specifies the height (number of columns) of the j-th face matrix, and matrix_height[j] can be 1 or greater.

[0176] matrix_element[j][k][m][n] specifies the quantized element values ​​of the corresponding face matrix.

[0177] In some embodiments, if semantic_present_flag is equal to 1, it indicates that the SEI message contains semantic parameters, whereas if semantic_present_flag is equal to 0, it indicates that the SEI message does not carry semantic parameters. If it is determined that semantic parameters are used to encode the face image, the parameter indicator indicates that the semantic parameters may be generated and signaled by the encoder and received by the decoder.

[0178] The semantic_quantization_factor specifies the quantization factor for processing face semantic parameters (i.e., semantic_element[l][y][z]). The parameter values ​​used for face generation are equal to the values ​​of the corresponding syntax elements divided by the semantic_quantization_factor.

[0179] `semantic_type_num` specifies the number of face semantic types in the SEI message. The value of `semantic_type_num` is between 0 and 2. 6 It may fall within the following range. In some embodiments, the facial semantic type can be classified into mouth parameters, eye parameters, head rotation parameters, head translation parameters, and head position parameters.

[0180] The semantic_idx contains an identification number related to the face semantics that may be used in SEI messages. [Table 7]

[0181] Regarding the definition of semantic_idx in Table 7, semantic_idx can be further extended if there are other facial semantic representations besides the given semantic type.

[0182] semantic_width[l] specifies the width (number of rows) of the l-th semantic type, and semantic_width[l] can be 1 or greater.

[0183] semantic_height[l] specifies the height (number of columns) of the l-th semantic type, and semantic_height[l] can be 1 or greater.

[0184] semantic_element[l][y][z] specifies the quantized element value from the corresponding face semantics.

[0185] The syntax of the generated face image SEI messages shown in Table 5 may be encoded and signaled within a bitstream and further decoded by a decoder. Figure 9 is a schematic diagram showing an exemplary SEI message 900 according to some embodiments of the present disclosure. As shown in Figure 9, the SEI message 900 may include an identification number indicator (gfv_id) 901, a coordinate indicator (coordinate_present_flag) 902, a matrix indicator (matrix_present_flag) 903, a semantic indicator (semantic_present_flag) 904, a parameter indicator 905, and the like. The corresponding bit lengths of these parameters are also shown in Figure 9 as an example. As is understood, the bit lengths of these parameters may be designed in other ways. The coordinate indicator 902, the matrix indicator 903, and the semantic indicator 904 are signaled cooperatively; that is, in a suitable SEI message 900, only one of the coordinate indicator 902, the matrix indicator 903, and the semantic indicator 904 may be activated. Furthermore, parameter indicators 905 can be generated and signaled depending on the activated mode. Other parameters are not shown in Figure 9 to avoid ambiguity.

[0186] Table 8 shows an example of the syntax for a common SEI message for generated facial images. As shown in Table 8, a key indicator gfv_key may be generated by the encoder and signaled to the decoder. In some embodiments, gfv_key may be used to indicate a neural network used on the encoder side to generate facial information parameters. In some embodiments, the neural network on the decoder side may not match the neural network on the encoder side. Therefore, gfv_key may be signaled to the decoder to determine if a match exists. Referring back to Figure 7A, the decoder may, in step 706, determine whether the SEI message matches its neural network based on the key indicator. In some embodiments, the neural network on the encoder side is also expected to be the neural network on the decoder side for decoding the bitstream, and is therefore also called the target network.

[0187] Referring further to Table 8, the SEI message may include an order indicator gfv_pic_order_cnt to indicate the display order of the face images generated in response to the SEI message.

[0188] In some embodiments, the SEI message may include a parameter presence indicator (e.g., coordinate_present_flag or matrix_present_flag). The mode used for encoding the face image is determined by the decoder in step 708 based on the parameter presence indicator. In the different modes indicated by the parameter presence indicator, parameter indicators for conveying face information parameters may be included in the SEI message. For example, in the case of coordinate_present_flag==1 (whereas matrix_present_flag==0), which means that the SEI message contains coordinate parameters, parameter indicators such as coordinate_precision_factor_minus1, num_coordinates_minus1, coordinate_z_present_flag, etc., may be included in the SEI message and can be decoded as face information parameters.

[0189] In some embodiments, the SEI messages in Table 8 may be utilized by the following operations: First, a key gfv_key is signaled, which can be used to determine whether the current generated face image SEI message matches a generation network. The decoder or post-processor may generate a face image using the received SEI message only if the received SEI message contains a known gfv_key specified by the application. Otherwise, the parameters in the received SEI message may be extracted by a network that does not match the generation network within the decoder or post-processor, in which case the decoder or post-processor ignores the received SEI message and does not generate a face image. Then, an image order count gfv_pic_order_cnt is signaled to indicate the order of the images generated by this SEI message.

[0190] Secondly, a parameter presence flag is signaled for different parameter types. If the parameter presence flag indicates the presence of a parameter of the current type, the corresponding parameter is signaled. Otherwise, signaling for the parameter of the current type is skipped. In some embodiments, parameters can be classified into two types: keypoints and matrices. Compared to some of the embodiments described above, which are accompanied by Table 5, semantic types are merged into matrix types in this specification. In some embodiments, matrix types may include affine translation matrices, covariance matrices, mouth matrices, eye matrices, head rotation matrices, head translation matrices, head position matrices, and compact feature matrices. In some embodiments, matrix types may include affine translation matrices, covariance matrices, compact feature matrices, and semantic matrices. Semantic matrix types may further include mouth parameter matrices, eye parameter matrices, head rotation parameter matrices, head translation matrices, and head position matrices.

[0191] Thirdly, after receiving and decoding the SEI message, the parameters signaled within the SEI and the previously decoded base image can be used as input to reconstruct the facial image in a high-quality or user-friendly manner through the generative adversarial network's generation capabilities. [Table 8-1] [Table 8-2]

[0192] This SEI message specifies the syntax and semantics of various face representations that can be used by a deep generative network to generate images based on a base image. The base image is an image that provides the texture information required by the deep generative network. The base image may be coded as a bitstream compliant with Rec.ITU-TH.264, Rec.ITU-TH.265, Rec.ITU-TH.266, etc. The decoder may decode the base image from the compliant bitstream as the first image in the sequence, and subsequent images may be generated by the deep generative network using the syntax elements signaled in the generated face video SEI message. Alternatively, the decoder may periodically decode the base image from the compliant bitstream. Based on the base image, the deep generative network may generate multiple subsequent images using the syntax elements signaled in the generated face video SEI message. The textures of the base image and any additionally generated images are expected to primarily contain talking faces.

[0193] To use this SEI message, you need to define the following variables. - The width and height of the cropped decoded output base image in luma sample units (referred to herein as CroppedWidth and CroppedHeight, respectively). - Cropped chroma sample arrays CroppedYPic and chroma sample arrays CroppedCbPic and CroppedCrPic of the decoded output base image. - Chroma format indicator (represented as ChromaFormatId in this specification)

[0194] The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc.

[0195] In some embodiments, gfv_key represents a key that may be used to identify the analysis network used to generate the syntax elements of the current generated facial image SEI message. The value of gfv_key may be used to determine whether the analysis network on the encoder side matches the generation network on the decoder side.

[0196] The generating network on the decoder side may generate video images using SEI messages only if it recognizes the value of gfv_key. The value of gfv_key may be specified by an application that assumes network matching between the encoder and decoder. In such an application, if the parameters in the received SEI message are extracted by an analysis network that does not match the generating network on the decoder side, the generating network may ignore the received SEI message.

[0197] In some embodiments, gfv_pic_order_cnt specifies the display order count modulo 1 << 31 for images currently generated in the SEI message.

[0198] In some embodiments, if coordinate_present_flag is equal to 1, it indicates that keypoint coordinate information exists, while if coordinate_present_flag is equal to 0, it indicates that keypoint coordinate information does not exist. If it is determined that keypoint coordinate information will be used to encode the face image, the parameter indicator indicates that the keypoint coordinate information can be generated and signaled by the encoder and received by the decoder.

[0199] As a bitstream conformance requirement, for any i from 0 to num_matrix_types_minus1, if matrix_type_idx[i] is equal to 0 or 1, then the value of coordinate_present_flag can be equal to 1.

[0200] coordinate_precision_factor_minus1+1 represents the bitwise length of coordinate_x, coordinate_y, and coordinate_z.

[0201] In some embodiments, the lower limit of precision can be greater than 1. For example, it may be k, and the syntax element may be changed to coordinate_precision_factor_minusk. coordinate_precision_factor_minusk+k represents the bitwise length of coordinate_x, coordinate_y, and coordinate_z.

[0202] num_coordinates_minus1+1 indicates the number of keypoint coordinates. The value of num_coordinates_minus1 is between 0 and 2. 10 It can be within the range of -1 or less.

[0203] If `coordinate_z_present_flag` is equal to 1, it indicates that the z-axis coordinate information for the keypoint exists. If `coordinate_z_present_flag` is equal to 0, it indicates that the z-axis coordinate information for the keypoint does not exist. If `coordinate_z_present_flag` does not exist, it is presumed to be 0.

[0204] coordinate_z_max_value_minus1+1 indicates the maximum absolute value of the z-axis coordinate of the keypoint.

[0205] In some embodiments, the lower limit of the maximum value of the z-axis coordinate may be greater than 1. For example, it may be k, and the signaled syntax element is changed to coordinate_z_max_value_minusk. coordinate_z_max_value_minusk+k represents the maximum absolute value of the z-axis coordinate of the keypoint.

[0206] coordinate_x_abs[i] specifies the normalized absolute value of the x-axis coordinate of the i-th keypoint.

[0207] `coordinate_x_sign_flag[i]` specifies the sign of the x-axis coordinate of the i-th keypoint. If `coordinate_x_sign_flag[i]` does not exist, it is presumed to be equal to 0.

[0208] coordinate_y_abs[i] specifies the normalized absolute value of the y-axis coordinate of the i-th keypoint.

[0209] `coordinate_y_sign_flag[i]` specifies the sign of the y-axis coordinate of the i-th keypoint. If `coordinate_y_sign_flag[i]` does not exist, it is assumed to be equal to 0.

[0210] coordinate_z_abs[i] specifies the normalized absolute value of the z-axis coordinate of the i-th keypoint.

[0211] `coordinate_z_sign_flag[i]` specifies the sign of the z-axis coordinate of the i-th keypoint. If `coordinate_z_sign_flag[i]` does not exist, it is assumed to be equal to 0.

[0212] The variables coordinateX[i], coordinateY[i], and coordinateZ[i], which represent the x, y, and z axis coordinates of the i-th keypoint, are derived as follows:

number

[0213] In some embodiments, a matrix_present_flag equal to 1 indicates the presence of a matrix parameter, while a matrix_present_flag equal to 0 indicates the absence of a matrix parameter. If it is determined that a matrix parameter will be used to encode a face image, the parameter indicator indicates that the matrix parameter may be generated and signaled by the encoder and received by the decoder.

[0214] matrix_element_precision_factor_minus1+1 represents the bitwise length of matrix_element_dec[i][j][k][l].

[0215] num_matrix_types_minus1+1 indicates the number of matrix types signaled in the SEI message. The value of matrix_type_num_minus1 is between 0 and 2. 6 It can be within the range of -1 or less.

[0216] matrix_type_idx[i] indicates the index of the i-th type matrix, as specified in Table 9. [Table 9]

[0217] The undefined matrix type is used to represent matrix types other than affine translation matrices, covariance matrices, rotation matrices, translation matrices, and compact feature matrices. Users may extend matrix types using this matrix type.

[0218] If num_matrices_equal_to_num_coordinates_flag[i] is equal to 1, it indicates that the number of matrices of the i-th type of matrix is ​​equal to num_coordinates_minus1+1. If num_matrices_equal_to_num_coordinates_flag[i] is equal to 0, it indicates that the number of matrices of the i-th type of matrix is ​​not equal to num_coordinates_minus1+1.

[0219] num_matrices_info[i] provides information for deriving the number of matrices of the i-th type of matrix.

[0220] matrix_width_minus1[i]+1 represents the width of a matrix of the i-th type.

[0221] If matrix_width_minus1[i] does not exist, the following can be inferred: -If matrix_type_idx[i] is equal to 0, 1, or 4, and coordinate_z_present_flag is equal to 1, then matrix_width_minus1[i] is presumed to be equal to 2. - Otherwise, if matrix_type_idx[i] is equal to 0, 1, or 4 and coordinate_z_present_flag is equal to 0, then matrix_width_minus1[i] is presumed to be equal to 1. -Otherwise (if matrix_type_idx[i] is equal to 5 or 6), matrix_width_minus1[i] is presumed to be equal to 0. matrix_height_minus1[i]+1 represents the height of the i-th type matrix. If matrix_height_minus1[i] does not exist, the following can be inferred: -If matrix_type_idx is equal to 0, 1, 4, 5, or 6, and coordinate_z_present_flag is equal to 1, then matrix_height_minus1[i] is presumed to be equal to 2. -Otherwise (if matrix_type_idx is equal to 0, 1, 4, 5, or 6 and coordinate_z_present_flag is 0), matrix_height_minus1[i] is presumed to be equal to 1.

[0222] num_matrices_minus1[i]+1 indicates the number of matrices of the i-th type.

[0223] If matrix_for_3D_space_flag[i] is equal to 1, it indicates that the i-th type of matrix is ​​a matrix in 3D space. If matrix_size_flag[i] is equal to 0, it indicates that the i-th type of matrix is ​​a matrix in 2D space.

[0224] The variable numMatrices[i], which indicates the number of matrices of type i, is derived as follows:

number

[0225] matrix_element_int[i][j][k][l] represents the integer part of the value of the matrix element at position (k,l) of the j-th matrix of the i-th type of matrix.

[0226] matrix_element_dec[i][j][k][l] represents the fractional part of the value of the matrix element at position (k,l) of the j-th matrix of the i-th type of matrix.

[0227] `matrix_element_sign_flag[i][j][k][l]` indicates the sign of the matrix element at position (k,l) of the j-th matrix of the i-th type of matrix. If `matrix_element_sign_flag[i][j][k][l]` does not exist, it is presumed to be equal to 0.

[0228] The variable MatrixElementVal[i][j][k][l], which represents the value of the matrix element at position (k,l) of the j-th matrix of type i, is derived as follows:

number

[0229] The process DeriveInputTensors(), which derives the input tensors inputTensorImgY, inputTensorImgCb, inputTensorImgCr, inputTensorKeyPoint, and inputTensorMatrix, is specified as follows: Initialize inputTensorImgY, inputTensorImgCb, inputTensorImgCr, inputTensorKeyPoint, and inputTensorMatrix to 0.

number

[0230] The process StoreOutputTensors() for deriving the output sample arrays OutputYPic, OutputCbPic, and OutputCrPic from the output tensors outputTensorY, outputTensorCb, and outputTensorCr is specified as follows:

number

[0231] The following processes are used to generate video images. PictureGeneration() is a deep generative network-based process used to take parameters signaled within face generation SEI messages and a decoded base image as input and output image sample values.

number

[0232] In some embodiments, the generative neural network may be divided into two subnetworks (e.g., a first subnetwork and a second subnetwork) to overcome interoperability issues.

[0233] In some embodiments, the first subnetwork may be a flow translator network that transfers parameters signaled in the SEI message to a flow map, while the second subnetwork may be a generative network that generates an image based on the flow map output by the first subnetwork. The flow translator network can support different types of face representations for different algorithms and transfer all of these face representation parameters to the flow map. Thus, the generative network may be fixed regardless of the type of face representation signaled in the SEI message, since the input to the generative network is always a flow map.

[0234] In some embodiments, the first subnetwork may be a parameter translator network that converts face representation parameters signaled within an SEI message into parameters of a specified type, while the second subnetwork may be a generation network that generates an image based on face representation parameters of a specified type. As a result, the generation network may also be fixed in that it supports parameters of a specified type as input.

[0235] In some embodiments, the SEI message can be used to directly signal the network or to provide a URI for indicating the type of network signaled within the SEI message to a subnetwork. For example, a key indicator (e.g., gfv_key as described herein) for indicating the type of network signaled within the SEI message can be signaled.

[0236] The network signaled or indicated within the SEI message can be updated. The first SEI message in the current CLVS may be used to signal or indicate the network as the base GFV, and subsequent SEI messages can be used to update the base GFV by signaling or indicating the updated network.

[0237] The syntax of the generated face video SEI message shown in Table 8 can be encoded and signaled within the bitstream and further decoded by the decoder. FIG. 10 is a schematic diagram showing an exemplary SEI message 1000 according to some embodiments of the present disclosure. As shown in FIG. 10, the SEI message 1000 may include a key indicator (gfv_key) 1001, an order indicator (gfv_pic_order_cnt) 1002, a coordinate indicator (coordinate_present_flag) 1003, a matrix indicator (matrix_present_flag) 1004, etc. The corresponding bit lengths of these parameters are also shown as an example in FIG. 10. As is understood, the bit lengths of these parameters can be designed in other ways. The coordinate indicator 1003 and the matrix indicator 1004 are signaled collaboratively. That is, in a proper SEI message 1000, only one of the coordinate indicator 1003 and the matrix indicator 1004 can be activated. Further, the parameter indicator 1005 can be generated and signaled according to the activated mode. Other parameters are not shown in FIG. 10 to avoid ambiguity.

[0238] An example of the syntax of the common SEI message for the generated face video is shown in Table 10. As shown in Table 10, the identification number indicator gfv_id can be generated by the encoder and signaled to the decoder. In some embodiments, gfv_id can be used to indicate whether the SEI is used for the coding of the face image.

[0239] In some embodiments, a key presence indicator (e.g., gfv_key_present_flag, described below) to indicate the presence of a key indicator (e.g., gfv_key, as described herein) may be included in the SEI message. Referring back to Figure 7A, in step 706, assuming that the key presence indicator indicates the presence of the key indicator, the decoder may determine whether the SEI message matches the decoder's neural network based on the key indicator. That is, if the key presence indicator indicates the absence of the key indicator, the decoder may skip determining whether the neural network matches.

[0240] In some embodiments, neural network indicators to indicate the configuration of the target neural network may be included in the SEI message. Thus, if a key presence indicator (e.g., gfv_key_present_flag) indicates the absence of a key indicator (e.g., gfv_key), the decoder can configure (e.g., initialize or update) the neural network based on the neural network indicators. In some embodiments, the neural network indicators may include gfv_nn_mode_idc, gfv_nn_type_idc, etc., as described below. [Table 10-1] [Table 10-2]

[0241] A Generated Face Video (GFV) SEI message indicates face parameters and specifies a neural network, denoted as Generator(), used to generate a new output image using the indicated face parameters and a previously decoded output image. The Generator() network is further divided into a translator subsystem, Translator(), and an image generation network, Decoder(). In one example, Translator() converts face parameters into a flow map, and Decoder() generates an image based on the flow map output by Translator(). In another example, Translator() converts face parameters signaled within the SEI message into parameters of a specified type, and Decoder() generates an image based on the face parameters of the specified type output by Translator().

[0242] Note that facial parameters can be determined from the source image before encoding.

[0243] It should also be noted that if the current image is not a base image, the GFV SEI message may be used to generate a new face image based on a previously decoded base image, face parameters transmitted by the GFV SEI message, or an optional fused image. The base image and fused image may be encoded in accordance with Rec.ITU-TH.264, Rec.ITU-TH.265, Rec.ITU-TH.266, etc.

[0244] To use this SEI message, you need to define the following variables. β€’ The width and height of the input image in units of luma samples (referred to as CroppedWidth and CroppedHeight, respectively, in this specification). β€’ Luma sample arrays (baseCroppedYPic) and chroma sample arrays (baseCroppedCbPic and baseCroppedCrPic) of the decoded output image (referred to as BasePicture), corresponding to the source base image. β€’ Luma sample array driveCroppedYPic and chroma sample arrays driveCroppedCbPic and driveCroppedCrPic for the decoded output image (referred to as DrivePicture) corresponding to the source-driven image. β€’ BitDepthY of the luma sample array of the input image. β€’ BitDepthC of the chroma sample array (if any) of the input image.

[0245] The gfv_id contains an identification number that identifies facial feature information and may be used to specify a neural network that can be used as a Generator(). The value of gfv_id is between 0 and 2. 32 It can be within the range of -2 or less. 256 or more and 511 or less, and 2 31 The above 2 32 gfv_id values ​​less than or equal to -2 are reserved for future use by ITU-T|ISO / IEC. gfv_id values ​​in the range of 256 to 511, or 2 31 The above 2 32 Decoders that encounter a GFV SEI message within the range of -2 or less may ignore that SEI message.

[0246] For example, if an output image contains multiple faces, note that different values ​​for gfv_id in different GFV SEI messages may be used to identify different faces.

[0247] In some embodiments, if gfv_key_present_flag is equal to 1, it indicates that the syntax element gfv_key exists, but the syntax elements gfv_nn_mode_idc, gfv_nn_reserved_zero_bit_a, gfv_nn_tag_uri, gfv_nn_uri, and gfv_nn_payload_byte[i] do not exist. If gfv_key_present_flag is equal to 0, it indicates that the syntax element gfv_key does not exist, but the syntax elements gfv_nn_mode_idc, gfv_nn_reserved_zero_bit_a, gfv_nn_tag_uri, gfv_nn_uri, and gfv_nn_payload_byte[i] may be signaled.

[0248] In some embodiments, gfv_key may represent a key used to identify the analysis network used to generate the syntax elements of the current generated facial image SEI message.

[0249] Note that the value of gfv_key may be used to determine whether the analysis network on the encoder side matches the generation network on the decoder side. The generation network on the decoder side may generate video images using SEI messages only if it recognizes the value of gfv_key. The value of gfv_key may also be specified by applications that assume network matching between the encoder and decoder. In such applications, if parameters in the received SEI message are extracted by an analysis network that does not match the generation network on the decoder side, the generation network can ignore the received SEI message.

[0250] In some embodiments, a neural network mode indicator to indicate the format of information about the neural network may be included in the SEI message. For example, if gfv_nn_mode_idc is equal to 0, it indicates that the neural network information is included in the GFV SEI message and that the neural network information is in the ISO / IEC 15938-17 bitstream format. If gfv_nn_mode_idc is equal to 1, it indicates that the neural network information is in the format identified by the URI indicated by gfv_nn_uri and identified by the tag URI gfv_nn_tag_uri.

[0251] The value of gfv_nn_mode_idc is between 0 and 255. Values ​​of gfv_nn_mode_idc between 2 and 255 are reserved for future use by ITU-T|ISO / IEC and do not exist in ITU-T|ISO / IEC compliant bitstreams. ITU-T|ISO / IEC compliant decoders can ignore GFV SEI messages where gfv_nn_mode_idc is between 2 and 255.

[0252] In some embodiments, an SEI message may include a neural network type indicator to indicate the type of neural network indicated within the SEI message. For example, gfv_nn_type_idc indicates the type of network included in or indicated in the GFV SEI message.

[0253] If gfv_nn_type_idc is equal to 0, the GFV SEI message contains or indicates Generator(). The network is a complete image generator with a motion estimation module for converting face parameters into flow maps and a frame generation module for reconstructing face images.

[0254] When gfv_nn_type_idc is equal to 1, Generator() is included or indicated in the GFV SEI message. Also, Translator() is a network that converts different types of face parameters signaled within the SEI message into a flow map for image generation.

[0255] When gfv_nn_type_idc is equal to 2, Translator() is included or indicated in the GFV SEI message. Also, Translator() is a network that converts different types of face parameters into specified types of face parameters.

[0256] When gfv_nn_type_idc is equal to 3, Decoder() is included or indicated in the GFV SEI message. Decoder() generates an image using a flow map or specified types of face parameters as input.

[0257] gfv_nn_reserved_zero_bit_a is 0.

[0258] A tag URI having the syntax and semantics defined in IETF RFC4151 is included in gfv_nn_tag_uri. The tag URI identifies the format and related information regarding the neural network used as the base GFV, or an update to the base GFV (having the same gfv_id value as specified by gfv_nn_uri).

[0259] For example, when gfv_nn_tag_uri is "tag:iso.org,2023:15938-17", it indicates that the neural network data is identified by gfv_nn_uri and conforms to ISO / IEC 15938-17.

[0260] In some embodiments, gfv_nn_uri includes a URI having the syntax and semantics defined in IETF Internet Standard 66, which identifies the neural network used as the base GFV, or an update to the base GFV having the same gfv_id value.

[0261] If `coordinate_present_flag` is equal to 1, it indicates that keypoint coordinate information exists, while if `coordinate_present_flag` is equal to 0, it indicates that keypoint coordinate information does not exist. If it is decided that keypoint coordinate information will be used to encode the face image, the parameter indicator indicates that the keypoint coordinate information can be generated and signaled by the encoder and received by the decoder.

[0262] coordinate_precision_factor_minus1+1 represents the bitwise length of coordinate_x, coordinate_y, and coordinate_z.

[0263] In some embodiments, the lower limit of precision can be greater than 1. For example, it may be k, and the syntax element may be changed to coordinate_precision_factor_minusk. coordinate_precision_factor_minusk+k represents the bitwise length of coordinate_x, coordinate_y, and coordinate_z.

[0264] num_coordinates_minus1+1 indicates the number of keypoint coordinates. The value of num_coordinates_minus1 is between 0 and 2. 10 It is within the range of -1 or less.

[0265] If `coordinate_z_present_flag` is equal to 1, it indicates that the z-axis coordinate information for the keypoint exists. If `coordinate_z_present_flag` is equal to 0, it indicates that the z-axis coordinate information for the keypoint does not exist. If `coordinate_z_present_flag` does not exist, it is presumed to be 0.

[0266] coordinate_z_max_value_minus1+1 indicates the maximum absolute value of the z-axis coordinate of the keypoint.

[0267] In some embodiments, the lower limit of the maximum value of the z-axis coordinate may be greater than 1. For example, it may be k, and the signaled syntax element is changed to coordinate_z_max_value_minusk. coordinate_z_max_value_minusk+k represents the maximum absolute value of the z-axis coordinate of the keypoint.

[0268] coordinate_x_abs[i] specifies the normalized absolute value of the x-axis coordinate of the i-th keypoint.

[0269] `coordinate_x_sign_flag[i]` specifies the sign of the x-axis coordinate of the i-th keypoint. If `coordinate_x_sign_flag[i]` does not exist, it is presumed to be equal to 0.

[0270] coordinate_y_abs[i] specifies the normalized absolute value of the y-axis coordinate of the i-th keypoint.

[0271] `coordinate_y_sign_flag[i]` specifies the sign of the y-axis coordinate of the i-th keypoint. If `coordinate_y_sign_flag[i]` does not exist, it is assumed to be equal to 0.

[0272] coordinate_z_abs[i] specifies the normalized absolute value of the z-axis coordinate of the i-th keypoint.

[0273] `coordinate_z_sign_flag[i]` specifies the sign of the z-axis coordinate of the i-th keypoint. If `coordinate_z_sign_flag[i]` does not exist, it is assumed to be equal to 0.

[0274] The variables coordinateX[i], coordinateY[i], and coordinateZ[i], which represent the x, y, and z axis coordinates of the i-th keypoint, are derived as follows:

number

[0275] In some embodiments, a matrix_present_flag equal to 1 indicates the presence of a matrix parameter, while a matrix_present_flag equal to 0 indicates the absence of a matrix parameter. If it is determined that a matrix parameter will be used to encode a face image, the parameter indicator indicates that the matrix parameter may be generated and signaled by the encoder and received by the decoder.

[0276] matrix_element_precision_factor_minus1+1 represents the bitwise length of matrix_element_dec[i][j][k][l].

[0277] num_matrix_types_minus1+1 indicates the number of matrix types signaled in the SEI message. The value of matrix_type_num_minus1 is between 0 and 2. 6 It is within the range of -1 or less.

[0278] `matrix_type_idx` indicates the index of the matrix type specified in Table 11. [Table 11]

[0279] Note that the undefined matrix type is used to represent matrix types other than affine translation matrices, covariance matrices, rotation matrices, translation matrices, and compact feature matrices. Users may extend matrix types using this matrix type.

[0280] If num_matrices_equal_to_num_coordinates_flag[i] is equal to 1, it indicates that the number of matrices of the i-th matrix type is equal to num_coordinates_minus1+1. If num_matrix_equal_to_num_coordinates_flag[i] is equal to 0, it indicates that the number of matrices of the i-th matrix type is not equal to num_coordinates_minus1+1.

[0281] num_matrices_info[i] provides information about the number of matrices of the i-th type of matrix.

[0282] matrix_width_minus1[i]+1 represents the width of a matrix of the i-th type.

[0283] If matrix_width_minus1[i] does not exist, the following can be inferred: If matrix_type_idx[i] is equal to 0, 1, or 4, and coordinate_z_present_flag is equal to 1, then matrix_width_minus1[i] is presumed to be equal to 2. Otherwise, if matrix_type_idx[i] is equal to 0, 1, or 4 and coordinate_z_present_flag is equal to 0, then matrix_width_minus1[i] is presumed to be equal to 1. Otherwise (if matrix_type_idx[i] is equal to 5 or 6), matrix_width_minus1[i] is presumed to be equal to 0.

[0284] matrix_height_minus1[i]+1 represents the height of the i-th type matrix.

[0285] If matrix_height_minus1[i] does not exist, the following can be inferred: If matrix_type_idx is equal to 0, 1, 4, 5, or 6, and coordinate_z_present_flag is equal to 1, then matrix_height_minus1[i] is presumed to be equal to 2. - Otherwise (if matrix_type_idx is equal to 0, 1, 4, 5, or 6 and coordinate_z_present_flag is 0), matrix_height_minus1[i] is presumed to be equal to 1.

[0286] num_matrices_minus1[i]+1 indicates the number of matrices of the i-th type.

[0287] The variable numMatrice[i], which indicates the number of matrices of type i, is derived as follows:

number

[0288] matrix_element_int[i][j][k][l] represents the integer part of the value of the matrix element at position (k,l) of the j-th matrix of the i-th type of matrix.

[0289] matrix_element_dec[i][j][k][l] represents the fractional part of the value of the matrix element at position (k,l) of the j-th matrix of the i-th type of matrix.

[0290] `matrix_element_sign_flag[i][j][k][l]` indicates the sign of the matrix element at position (k,l) of the j-th matrix of the i-th type of matrix. If `matrix_element_sign_flag[i][j][k][l]` does not exist, it is presumed to be equal to 0.

[0291] The variable MatrixElementVal[i][j][k][l], which represents the value of the matrix element at position (k,l) of the j-th matrix of type i, is derived as follows:

number

[0292] The process DeriveInputTensors(), which derives the input tensors inputTensorImgY, inputTensorImgCb, inputTensorImgCr, inputTensorKeyPoint, and inputTensorMatrix, is specified as follows: Initialize inputTensorImgY, inputTensorImgCb, inputTensorImgCr, inputTensorKeyPoint, and inputTensorMatrix to 0.

number

[0293] If gfv_nn_type_idc is equal to 1, the process TranslateTensors() is used to derive the input sensor TransTensorFlow for Decoder() from the output of Translator(). If gfv_nn_type_idc is equal to 2, the process TranslateTensors() derives the input sensor TransTensorKeyPoint and TransTensorMatrix for Decoder() from the output of Translator(). This process can be described as follows:

number

[0294] The StoreOutputTensors() process for deriving sample values ​​in the output sample arrays OutputYPic, OutputCbPic, and OutputCrPic, which are generated from the output tensors outputTensorY, outputTensorCb, and outputTensorCr, is specified as follows:

number

[0295] The following processes are used to generate video images: Generator() is used to take parameters signaled in GFV SEI messages and a decoded base image as input and output image sample values. Translator() is used to convert face parameters signaled in SEI messages into flow maps or to convert face parameters signaled in SEI messages into face parameters of a specified type.

number

[0296] The disclosed facial video compression and generation methods, while described above in relation to SEI messages, are not limited to SEI messages. Rather, they may be implemented in other ways. For example, an extension of an existing video coding standard or a new standard may be defined to support the proposed generated video compression. The video may be encoded in the base layer or upper layer using the proposed method. Furthermore, all the syntax elements, semantics, and decoding methods described above are applicable to extensions or new standards.

[0297] In some embodiments, the first subnetwork is a parameter translator network that converts face representation parameters signaled within an SEI message into parameters in a fixed format, and the second subnetwork is a generation network that generates an image based on the parameters in the fixed format.

[0298] The first subnetwork, denoted as TranslatorNN() (i.e., the parameter translator network), may be signaled within the SEI message or indicated by a URI included in the SEI message. A flag is signaled to indicate whether the translator network is signaled or indicated in the SEI message. Furthermore, information regarding the fixed parameter format output by the parameter translator (e.g., the number of keypoints, the number of matrices, and the size of each matrix) may also be signaled within the SEI message or output by the translator itself.

[0299] Furthermore, predictive signaling is supported for signaling keypoint coordinates and matrix elements. Specifically, the difference between the keypoint coordinates or matrix elements of the base image and the current frame is signaled. A flag is also signaled indicating whether the difference value or the absolute value is being signaled.

[0300] The syntax of the generated face image SEI messages shown in Table 10 can be encoded and signaled within a bitstream and further decoded by a decoder. Figure 11 is a schematic diagram showing an exemplary SEI message 1100 according to some embodiments of the present disclosure. As shown in Figure 11, when conveying face information parameters, the SEI message 1100 may include an identification number indicator (gfv_id) 1101, a key presence indicator (gfv_key_present_flag) 1102 indicating the presence of a key indicator, a key indicator (gfv_key) 1103, a coordinate indicator (coordinate_present_flag) 1104, a matrix indicator (matrix_present_flag) 1105, and the like. The coordinate indicator 1104 and the matrix indicator 1105 are signaled cooperatively. That is, in an appropriate SEI message 1100, only one of the coordinate indicator 1104 or the matrix indicator 1105 may be activated. Furthermore, the parameter indicator 1106 can be generated and signaled according to the activated mode.

[0301] In some embodiments, the SEI message 1100 may include, when used to communicate the configuration of a neural network, an identification number indicator (gfv_id) 1101, a key presence indicator (gfv_key_present_flag) 1102 indicating the absence of a key indicator, a neural network mode indicator (gfv_nn_mode_idc) 1107, a neural network type indicator (gfv_nn_type_idc) 1108, and the like. The corresponding bit lengths of these parameters are also shown as an example in Figure 11. As is understood, the bit lengths of these parameters may be designed in other ways. Other parameters are not shown in Figure 11 to avoid ambiguity.

[0302] An example of the syntax for a common SEI message for generated facial images is shown in Table 12. As shown in Table 12, the identification number indicator gfv_id can be generated by the encoder and signaled to the decoder. In some embodiments, gfv_id may be used to indicate whether or not the SEI is used to encode the facial image.

[0303] As described above, a face image can be reconstructed based on a base image. Therefore, the encoded and signaled image to the decoder can be classified as either a base image or an image used to generate face information parameters (also referred to herein as a driving image). It is understood that face information parameters may not be signaled within the SEI message associated with the base image. Therefore, the SEI message associated with the base image can be used on the decoder side to set up (initialize or update) the neural network.

[0304] In some embodiments, the SEI message may include a base image indicator (e.g., gfv_base_pic_flag, described later) to indicate that the SEI message is currently associated with a base image. Referring back to Figure 7A, in step 706, the decoder may determine that the SEI message is applied to a neural network for generating a face image if the base image indicator indicates that the SEI message is not currently associated with a base image. Otherwise, the decoder may determine that the SEI message does not currently contain any face information parameters.

[0305] In some embodiments, if it is determined that the SEI message is currently associated with a base image, the SEI message may be used on the decoder side to set up (initialize or update) the neural network. The SEI message may also include a neural presence indicator (e.g., gfv_nn_present_flag, described below) to indicate whether information for a first subnetwork is indicated by the SEI message. In some embodiments, if the base image indicator indicates that the SEI message is associated with a base image and the neural network presence indicator indicates that information for a first subnetwork is indicated by the SEI message, the first subnetwork may be set up (initialized or updated) based on the neural network indicator (e.g., gfv_nn_base_flag, gfv_nn_mode_idc). [Table 12-1] [Table 12-2] [Table 12-3]

[0306] The generated face video (GFV) SEI message specifies a face parameter translator network, denoted as TranslatorNN(), which indicates face parameters and may be used to convert face parameters in various formats signaled within the SEI message to parameters in a fixed format, and a face image generator neural network, denoted as GenerativeNN(), which may be used to generate an output image using face parameters in a fixed format and a previously decoded output image.

[0307] It should be noted that facial parameters can be determined from the source image before encoding. Such a source image is sometimes called a driving image.

[0308] Furthermore, the decoded output image that is input to GenerativeNN() may be a base image (a decoded output image that provides a reference texture for generating a face image) and an optional image that can be fused by GenerativeNN() to improve the background texture and facial details. If the current image is not the base image, the GFV SEI message may be used to generate a face image based on the previously decoded base image, the face parameters conveyed by the GFV SEI message, and the optional current decoded image for fusion.

[0309] To use this SEI message, you need to define the following variables. - The width and height of the input image in units of luma samples (referred to as CroppedWidth and CroppedHeight, respectively, in this specification). - Luma sample array baseCroppedYPic and chroma sample arrays baseCroppedCbPic and baseCroppedCrPic of the decoded output image (denoted as BasePicture) corresponding to the source base image. - Luma sample array driveCroppedYPic and chroma sample arrays driveCroppedCbPic and driveCroppedCrPic of the decoded output image (referred to as DrivePicture) corresponding to the source-driven image. - BitDepthY of the luma sample array of the input image. - The bit depth of the chroma sample array (if any) of the input image, BitDepthC. - Chroma format indicator (described in Section 7.3 of this specification, denoted as ChromaFormatIdc). -The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc.

[0310] The gfv_id contains an identification number that identifies facial feature information and may be used to specify a neural network that may be used as GenerativeNN(). The value of gfv_id is between 0 and 2. 32 It can be within the range of -2 or less. 256 or more and 511 or less, and 2 31 The above 2 32 gfv_id values ​​less than or equal to -2 are reserved for future use by ITU-T|ISO / IEC. gfv_id values ​​in the range of 256 to 511, or 2 31 The above 2 32 Decoders that encounter a GFV SEI message within the range of -2 or less may ignore that SEI message.

[0311] For example, if an output image contains multiple faces, note that different values ​​for gfv_id in different GFV SEI messages may be used to identify different faces.

[0312] In some embodiments, if gfv_base_pic_flag is equal to 1, it indicates that the currently decoded output image corresponds to the base image. If gfv_base_pic_flag is equal to 0, it indicates that the currently decoded output image does not correspond to the base image.

[0313] The following constraints apply to the value of gfv_base_pic_flag: -If a GFV SEI message is the first GFV SEI message in CLVS that currently has a specific gfv_id value in the decoding order, the value of gfv_base_pic_flag may be equal to 1. -If a GFV SEI message with a specific gfv_id value has a gfv_base_pic_flag equal to 0, this SEI message relates to the current decoded image of the current layer and all subsequent decoded images in output order up to the end of the current CLVS or up to the decoded image (but not that decoded image). This decoded image is associated with the current decoded image in output order within the current CLVS, followed by subsequent GFV SEI messages in decoding order within the current CLVS that have a gfv_base_pic_flag equal to 0 and have that specific gfv_id value (but not the image itself).

[0314] In some embodiments, if gfv_nn_present_flag is equal to 1, it indicates that a neural network that may be used as TranslatorNN() is included in or indicated in the SEI message. If gfv_nn_present_flag is equal to 0, it indicates that a neural network that may be used as TranslatorNN() is not included in or indicated in the SEI message.

[0315] gfv_nn_base_flag, gfv_nn_mode_idc, gfv_nn_reserved_zero_bit_a, gfv_nn_tag_uri, gfv_nn_uri, and gfv_nn_payload_byte[i] specify a neural network that may be used as TranslatorNN(). gfv_nn_base_flag, gfv_nn_mode_idc, gfv_nn_reserved_zero_bit_a, gfv_nn_tag_uri, gfv_nn_uri, and gfv_nn_payload_byte[i] have the same syntax and semantics as nnpfc_base_flag, nnpfc_mode_idc, nnpfc_reserved_zero_bit_a, nnpfc_tag_uri, nnpfc_uri, and nnpfc_payload_byte[i], respectively.

[0316] In some embodiments, if gfv_drive_pic_fusion_flag (if present) is equal to 1, it indicates that the currently decoded image corresponding to the drive image that may be used for fusion may be input to GenerativeNN(). If gfv_drive_pic_fusion_flag is equal to 0, it indicates that the currently decoded image should not be input to GenerativeNN().

[0317] Note that a value of 1 for gfv_drive_pic_fusion_flag may be used, for example, to indicate that the decoded image is currently ready for processing, such as improving facial details or handling background changes.

[0318] Note that fusion takes three inputs: a base image, features from keypoints or matrices carried within GFV SEI messages, and the currently decoded image, and outputs an image.

[0319] Please note that if the decoded image corresponds to the driving image, it must be marked as not intended for output.

[0320] If gfv_coordinate_present_flag is equal to 1, it indicates that coordinate information for the keypoint exists. If gfv_coordinate_present_flag is equal to 0, it indicates that coordinate information for the keypoint does not exist.

[0321] As a bitstream conformance requirement, for any i from 0 to gfv_num_matrix_types_minus1, if gfv_matrix_type_idx[i] is equal to 0 or 1, then the value of gfv_coordinate_present_flag can be equal to 1.

[0322] gfv_coordinate_precision_factor_minus1+1 represents the bitwise length of gfv_coordinate_x_abs[i], gfv_coordinate_y_abs[i], and gfv_coordinate_z_abs[i].

[0323] gfv_num_kp_minus1+1 indicates the number of keypoints. The value of gfv_num_kp_minus1 is between 0 and 2. 10 It can be within the range of -1 or less.

[0324] If gfv_kp_pred_flag is equal to 1, it indicates that the syntax elements gfv_coordinate_dx_abs[i], gfv_coordinate_dy_abs[i], and gfv_coordinate_dz_abs[i] exist, and that the syntax elements gfv_coordinate_dx_sign_flag[i], gfv_coordinate_dy_sign_flag[i], and gfv_coordinate_dz_sign_flag[i] may exist. If gfv_kp_pred_flag is equal to 0, it indicates that gfv_coordinate_x_abs[i], gfv_coordinate_y_abs[i], and gfv_coordinate_z_abs[i] exist, and that syntax elements gfv_coordinate_x_sign_flag[i], gfv_coordinate_y_sign_flag[i], and gfv_coordinate_z_sign_flag[i] may exist.

[0325] If gfv_coordinate_z_present_flag is equal to 1, it indicates that z-axis coordinate information for the keypoint exists. If gfv_coordinate_z_present_flag is equal to 0, it indicates that z-axis coordinate information for the keypoint does not exist.

[0326] gfv_coordinate_z_max_value_minus1+1 represents the maximum absolute value of the z-axis coordinate of the keypoint.

[0327] gfv_coordinate_x_abs[i] represents the normalized absolute value of the x-axis coordinate of the i-th keypoint.

[0328] gfv_coordinate_x_sign_flag[i] specifies the sign of the x-axis coordinate of the i-th keypoint. If gfv_coordinate_x_sign_flag[i] does not exist, it is presumed to be equal to 0.

[0329] gfv_coordinate_y_abs[i] specifies the normalized absolute value of the y-axis coordinate of the i-th keypoint.

[0330] gfv_coordinate_y_sign_flag[i] specifies the sign of the y-axis coordinate of the i-th keypoint. If gfv_coordinate_y_sign_flag[i] does not exist, it is presumed to be equal to 0.

[0331] gfv_coordinate_z_abs[i] specifies the normalized absolute value of the z-axis coordinate of the i-th keypoint.

[0332] gfv_coordinate_z_sign_flag[i] specifies the sign of the z-axis coordinate of the i-th keypoint. If gfv_coordinate_z_sign_flag[i] does not exist, it is presumed to be equal to 0.

[0333] gfv_coordinate_dx_abs[i] represents the absolute difference of the normalized x-axis coordinates of the i-th keypoint.

[0334] gfv_coordinate_dx_sign_flag[i] specifies the sign of the difference value of the x-axis coordinate of the i-th keypoint. If gfv_coordinate_dx_sign_flag[i] does not exist, it is presumed to be equal to 0.

[0335] gfv_coordinate_dy_abs[i] specifies the absolute difference value of the normalized y-axis coordinates of the i-th keypoint.

[0336] gfv_coordinate_dy_sign_flag[i] specifies the sign of the difference value of the y-axis coordinate of the i-th keypoint. If gfv_coordinate_dy_sign_flag[i] does not exist, it is presumed to be equal to 0.

[0337] gfv_coordinate_dz_abs[i] specifies the absolute difference value of the normalized z-axis coordinates of the i-th keypoint.

[0338] gfv_coordinate_dz_sign_flag[i] specifies the sign of the difference value of the z-axis coordinate of the i-th keypoint. If gfv_coordinate_dz_sign_flag[i] does not exist, it is presumed to be equal to 0.

[0339] The variables coordinateDeltaX[i], coordinateDeltaY[i], and coordinateDeltaZ[i], which represent the x-axis, y-axis, and z-axis coordinates of the i-th keypoint, are derived as follows.

number

[0340] The variables coordinateX[i], coordinateY[i], and coordinateZ[i], which represent the x-axis, y-axis, and z-axis coordinates of the i-th keypoint, are derived as follows: The cases where gfv_kp_pred_flag is equal to 0 and gfv_kp_pred_flag is equal to 1 are derived as follows. Also, BaseKpCoordinateX[i], BaseKpCoordinateY[i], and BaseKpCoordinateZ[i], which represent the x, y, and z axis coordinates of the i-th keypoint of the base image, are derived as follows.

number

number

number

[0341] If gfv_matrix_present_flag is equal to 1, it indicates that the matrix parameter exists. If gfv_matrix_present_flag is equal to 0, it indicates that the matrix parameter does not exist.

[0342] gfv_matrix_element_precision_factor_minus1+1 indicates the bit length of gfv_matrix_element_dec[i][j][k][m].

[0343] gfv_num_matrix_types_minus1+1 indicates the number of matrix types signaled in the SEI message. The value of fv_matrix_type_num_minus1 is between 0 and 2. 6 It can be within the range of -1 or less.

[0344] If gfv_matrix_pred_flag is equal to 1, it indicates that the syntax elements gfv_matrix_element_int[i][j][k][m], gfv_matrix_element_dec[i][j][k][m], and gfv_matrix_element_sign_flag[i][j][k][m] may exist. If gfv_matrix_pred_flag is equal to 0, it indicates that gfv_matrix_delta_element_int[i][j][k][m] and gfv_matrix_delta_element_dec[i][j][k][m] exist, and that the syntax element gfv_matrix_delta_element_sign_flag[i][j][k][m] may exist. If gfv_matrix_pred_flag does not exist, it is inferred that the value of gfv_matrix_pred_flag is 0.

[0345] gfv_matrix_type_idx[i] indicates the index of the i-th type matrix specified in Table 13 below. [Table 13]

[0346] Note that the undefined matrix type is used to represent matrix types other than affine translation matrices, covariance matrices, rotation matrices, translation matrices, and compact feature matrices. Users may extend matrix types using this matrix type.

[0347] If gfv_num_matrices_equal_to_num_kps_flag[i] is equal to 1, it indicates that the number of matrices of the i-th type of matrix is ​​equal to gfv_num_kps_minus1+1. If gfv_num_matrices_equal_to_num_kps_flag[i] is equal to 0, it indicates that the number of matrices of the i-th type of matrix is ​​not equal to gfv_num_coordinates_minus1+1.

[0348] gfv_num_matrices_info[i] provides information for deriving the number of matrices of the i-th type of matrix.

[0349] gfv_matrix_width_minus1[i]+1 represents the width of the i-th type matrix.

[0350] gfv_matrix_height_minus1[i]+1 represents the height of the i-th type matrix.

[0351] If gfv_matrix_for_3D_space_flag[i] is equal to 1, it indicates that the i-th type of matrix is ​​a matrix defined in 3D space. If gfv_matrix_for_3D_space_flag[i] is equal to 0, it indicates that the i-th type of matrix is ​​a matrix defined in 2D space.

[0352] If gfv_matrix_width_minus1[i] does not exist, the following is inferred: -If gfv_matrix_type_idx[i] is equal to 0, 1, or 4, and either coordinate_z_present_flag or gfv_matrix_for_3D_space_flag[i] exists and is equal to 1, then gfv_matrix_width_minus1[i] is presumed to be equal to 2. - Otherwise, if matrix_type_idx[i] is equal to 0, 1, or 4, and either coordinate_z_present_flag or gfv_matrix_for_3D_space_flag[i] exists and is equal to 0, then gfv_matrix_width_minus1[i] is presumed to be equal to 1. -Otherwise (if matrix_type_idx[i] is equal to 5 or 6), gfv_matrix_height_minus1[i] is presumed to be equal to 0. If gfv_matrix_height_minus1[i] does not exist, the following is inferred: -If matrix_type_idx is equal to 0, 1, 4, 5, or 6, and either gfv_coordinate_z_present_flag or gfv_matrix_for_3D_space_flag[i] exists and is equal to 1, then gfv_matrix_height_minus1[i] is presumed to be equal to 2. -Otherwise (if gfv_matrix_type_idx is equal to 0, 1, 4, 5, or 6, and either gfv_coordinate_z_present_flag or gfv_matrix_for_3D_space_flag[i] is 0), gfv_matrix_height_minus1[i] is presumed to be equal to 1.

[0353] The variables matrixWidth[i] and matrixHeight[i], which represent the width and height of the matrix of type i, are derived as follows:

number

[0354] gfv_num_matrices_minus1[i]+1 indicates the number of matrices of the i-th type.

[0355] The variable numMatrices[i], which indicates the number of matrices of type i, is derived as follows:

number

[0356] In some of the disclosed embodiments, the position (k,m) represents the position of the k-th row and m-th column of the matrix.

[0357] gfv_matrix_element_int[i][j][k][m] represents the integer part of the value of the matrix element at position (k,m) of the j-th matrix of the i-th type of matrix.

[0358] gfv_matrix_element_dec[i][j][k][m] represents the fractional part of the value of the matrix element at position (k,m) of the j-th matrix of the i-th type of matrix.

[0359] gfv_matrix_element_sign_flag[i][j][k][m] indicates the sign of the matrix element at position (k,m) of the i-th matrix of the i-th type of matrix. If gfv_matrix_element_sign_flag[i][j][k][m] does not exist, it is presumed to be equal to 0.

[0360] gfv_matrix_element_int[i][j][k][m] represents the integer part of the difference value of the matrix element at position (k,m) of the j-th matrix of the i-th type of matrix.

[0361] gfv_matrix_element_dec[i][j][k][m] represents the fractional part of the difference value of the matrix element at position (k,m) of the j-th matrix of the i-th type of matrix.

[0362] gfv_matrix_element_sign_flag[i][j][k][m] indicates the sign of the difference value of the matrix element at position (k,m) of the j-th matrix of the i-th type of matrix. If gfv_matrix_element_sign_flag[i][j][k][m] does not exist, it is presumed to be equal to 0.

[0363] The `gfv_matrix_pred_flag` flag indicates whether the matrix element values ​​are predictively signaled. If this flag is equal to 1, the difference in the matrix element values ​​is signaled. The signaled difference may be the difference between the matrix element value of the current image and the matrix element value of the base image. Alternatively, the signaled difference may be the difference between the matrix element value of the image and the value of a previous element in the matrix of that image. Furthermore, a first difference between the matrix element value of the current image and the matrix element value of the base image can be derived, and then a second difference between the first difference value of the matrix element and the first difference value of a previous matrix element can be derived and signaled.

[0364] As an example, the variable matrixElementDeltaVal[i][j][k][m], which represents the difference value of the matrix element at position (k,m) of the j-th matrix of type i, is derived as follows:

number

[0365] The variable matrixElementVal[i][j][k][m], which represents the value of the matrix element at position (k,m) of the j-th matrix of type i, is derived as follows: The cases where gfv_matrix_pred_flag is equal to 0 and gfv_matrix_pred_flag is equal to 1 are derived as follows:

number

number

[0366] As another example, the variable matrixElementDeltaVal[i][j][k][m], which represents the difference value of the matrix element at position (k,m) of the j-th matrix of type i, can be derived as follows:

number

[0367] The variable matrixElementVal[i][j][k][m], which represents the value of the matrix element at position (k,m) of the j-th matrix of type i, is derived as follows: The cases where gfv_matrix_pred_flag is equal to 0 and gfv_matrix_pred_flag is equal to 1 are derived as follows:

number

number

[0368] In some embodiments, if gfv_nn_output_info_present_flag is equal to 1, it indicates that the syntax elements gfv_nn_output_num_kps, gfv_nn_output_num_matrices, gfv_nn_output_matrix_width_minus1[i], and gfv_nn_output_matrix_height_minus1[i] may exist. If gfv_nn_output_info_present_flag is equal to 0, it indicates that the syntax elements gfv_nn_output_num_kps, gfv_nn_output_num_matrices, gfv_nn_output_matrix_width_minus1[i], and gfv_nn_output_matrix_height_minus1[i] do not exist. These may be used to configure a first subnetwork.

[0369] gfv_nn_output_num_kps indicates the number of keypoints output by TranslatorNN(). The value of gfv_nn_output_num_kps is between 0 and 2. 10 It may be within the following range.

[0370] gfv_nn_output_num_matrices indicates the number of matrices output by TranslatorNN(). The value of gfv_nn_output_num_matrices is between 0 and 2. 10 It may be within the following range.

[0371] gfv_nn_output_matrix_width_minus1[i]+1 indicates the width of the i-th matrix output by TranslatorNN(). The value of gfv_nn_output_matrix_width_minus1[i] is between 0 and 2. 10 It can be within the range of -1 or less.

[0372] gfv_nn_output_matrix_height_minus1[i]+1 indicates the height of the i-th matrix output by TranslatorNN(). The value of gfv_nn_output_matrix_height_minus1[i] is between 0 and 2. 10 It can be within the range of -1 or less.

[0373] The following processes are used to generate video images.

number

[0374] The process DeriveSigParam(), which derives the input for TranslatorNN(), is specified as follows: The keypoint coordinate array sigKeyPoint and the matrix sigMatrix are derived as follows.

number

ζ•°

[0375] It should be noted that there may be a small error in the original text where "5L" in line ID=51 should probably be "51". This has been maintained in the translation for the purpose of following the exact original content.For example, TranslatorNN() is a process that converts face parameters in various formats carried within SEI messages into face parameters in a fixed format that are input to the generative network to generate the output image.

[0376] The input to TranslatorNN() includes sigKeyPoint and sigMatrix. The output of TranslatorNN() also includes convKeyPoint and convMatrix.

[0377] The DeriveInputTensors() process, which derives the input for GenerativeNN(), is specified as follows: If gfv_base_pic_flag is equal to 1, the BasePicture input tensors inputBaseY, inputBaseCb, and inputBaseCr are derived as follows. If gfv_drive_pic_fusion_flag is equal to 1, the DrivePicture luma sample array inputDriveY, inputDriveCb, and inputDriveCr are derived as follows. If gfv_base_pic_flag is equal to 0, the current image keypoint coordinate array inputDriveKeyPoint and matrix inputDriveMatrix are derived as follows. If gfv_base_pic_flag is equal to 1, the keypoint coordinate array inputBaseKeyPoint and the base image matrix inputBaseMatrix are derived as follows.

number

number

number

number

[0378] The functions InpY() and InpC() are specified as follows:

number

[0379] As another example, when the translator outputs fixed-format face parameters that are input to the generating network, it also outputs fixed-format face parameter information such as the number of keypoints, the number of matrices, and the size of each matrix. Therefore, it is not necessary to signal this information within the SEI message. For this reason, the syntax elements gfv_nn_output_num_kps, gfv_nn_output_num_matrices, gfv_nn_output_matrix_width_minus1[i], and gfv_nn_output_matrix_height_minus1[i] may be skipped. In this example, The input to TranslatorNN() is: -sigKeyPoint and sigMatrix The output of TranslatorNN() is: -convKeyPoint, numConvKeyPoint -convMatrix, numConvMatrix, and the arrays convMatrixWidth[i] and convMatrixHeight[i] (if numConvMatrix>0, i=0 to numConvMatrix-1)

[0380] The DeriveInputTensors() process, which derives the input for GenerativeNN(), is specified as follows: If gfv_base_pic_flag is equal to 0, the keypoint coordinate array inputDriveKeyPoint and matrix inputDriveMatrix of the current image are derived as follows. If gfv_base_pic_flag is equal to 1, the keypoint coordinate array inputBaseKeyPoint and the matrix inputBaseMatrix of the base image are derived as follows.

number

number

[0381] GenerativeNN() is a process for generating sample values ​​for the output image corresponding to the driving image. It is called only if gfc_base_pic_flag is equal to 0. The input values ​​to GenerativeNN() and the output values ​​from GenerativeNN() are real numbers.

[0382] The input to GenerativeNN() is: -If gfv_base_pic_flag is equal to 0, gfv_drive_pic_fusion_flag is equal to 0, and ChromaFormatIdc is equal to 0: inputBaseY, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix. -If gfv_base_pic_flag is equal to 0, gfv_drive_pic_fusion_flag is equal to 0, and ChromaFormatIdc is not equal to 0: inputBaseY, inputBaseCb, inputBaseCr, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix. -If gfv_base_pic_flag is equal to 0, gfv_drive_pic_fusion_flag is equal to 1, and ChromaFormatIdc is equal to 0: inputBaseY, inputDriveY, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix. -If gfv_base_pic_flag is equal to 0, gfv_drive_pic_fusion_flag is equal to 1, and ChromaFormatIdc is not equal to 0: inputBaseY, inputBaseCb, inputBaseCr, inputDriveY, inputDriveCb, inputDriveCr, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix.

[0383] The output of GenerativeNN() is, - Luma Sample Array genY -If ChromaFormatIdc is not equal to 0, two chroma sample arrays genCb and genCr are used.

[0384] The process StoreOutputTensors() for deriving the output is specified as follows: If gfv_base_pic_flag is equal to 0, the output sample arrays outYPic[x][y], outCbPic[x][y], and outCrPic[x][y] are derived as follows. If gfv_base_pic_flag is equal to 1, the output sample arrays outYPic[x][y], outCbPic[x][y], and outCrPic[x][y] are derived as follows.

number

number

[0385] The functions OutY() and OutC() are specified as follows:

number

[0386] The syntax of the generated face image SEI messages shown in Table 12 may be encoded and signaled within a bitstream and further decoded by a decoder. Figure 12 is a schematic diagram showing an exemplary SEI message 1200 according to some embodiments of the present disclosure. As shown in Figure 12, the SEI message 1200 may include, when used to communicate the configuration of a neural network, an identification number indicator (gfv_id) 1201, a base image indicator (e.g., gfv_base_pic_flag) 1202 indicating that the SEI message 1200 is associated with a base image, a neural presence indicator (e.g., gfv_nn_present_flag) 1203 indicating that information of a first subnetwork is presented by the SEI message 1200, a neural network indicator (e.g., gfv_nn_mode_idc 1204 and gfv_nn_mode_idc 1205), and the like.

[0387] In some embodiments, when the SEI message 1100 is used to transmit facial information parameters, it may include an identification number indicator (gfv_id) 1201, a base image indicator (e.g., gfv_base_pic_flag) 1202 indicating that the SEI message 1200 is not associated with a base image, a neural presence indicator (e.g., gfv_nn_prent_flag) 1203 indicating that information for a first subnetwork is not indicated by the SEI message 1200, gfv_drive_pic_flag 1206, a coordinate indicator (coordent_prent_flag) 1207, a matrix indicator (matrix_prent_flag) 1208, and so on. That is, in an appropriate SEI message 1200, only one of the coordinate indicator 1207 and the matrix indicator 1208 may be activated. Furthermore, a parameter indicator 1209 may be generated and signaled depending on the activated mode. The bit lengths corresponding to these parameters are also shown as an example in Figure 12. As is understood, the bit lengths of these parameters can be designed in other ways. Other parameters are not shown in Figure 12 to avoid ambiguity.

[0388] In some embodiments, methods for encoding a video sequence into a bitstream are also provided. Figure 13 is a flowchart of an exemplary method 1300 for encoding a video sequence into a bitstream, according to some embodiments of the present disclosure. As shown in Figure 13, method 1300 may include steps 1302 and 1304, which can be implemented by an encoder (e.g., the image / video encoder 124 in Figure 1, or the device 400 in Figure 4).

[0389] In step 1302, the encoder receives the video sequence.

[0390] In step 1304, the encoder may encode one or more images of the video sequence to generate a bitstream. The encoder may encode a base image and additional enhancement information (SEI) messages from one or more images. The SEI messages are used to indicate the mode used to encode the face image and the corresponding face information parameters. The bitstream is used in a decoder (e.g., the image / video decoder 144 in Figure 1 or the device 400 in Figure 4) to generate a face image by a neural network based on the base image and face information parameters.

[0391] As is understood, SEI messages can be encoded according to the specifications described in Figures 8-12 and shown in the table above.

[0392] In some embodiments, a non-temporary, computer-readable storage medium for storing the bitstream is also provided. The bitstream can be encoded and decoded according to the generated facial image augmentation information (SEI) messages described above (e.g., Figures 8-12).

[0393] In some embodiments, non-temporary computer-readable storage media containing instructions are also provided, and these instructions may be executed by devices for carrying out the methods described above (such as disclosed encoders and decoders). Examples of common forms of non-temporary media include floppy disks, flexible disks, hard disks, solid-state drives, magnetic tapes, or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media having a pattern of holes, RAM, PROMs, EPROMs, FLASH-EPROMs, or any other flash memory, NVRAMs, caches, registers, any other memory chips or cartridges, and networked versions thereof. The device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.

[0394] Embodiments may be further described using the following clauses. 1. A method for generating a face image, Receiving a bitstream and The encoded information of the bitstream is decoded to obtain the base image and the SEI (Special Extended Information) message, The determination of whether the aforementioned SEI message is applied to the neural network for generating a facial image, In response to the application of the SEI message to the neural network for generating the face image, the mode used for encoding the face image and the corresponding face information parameters are determined based on the SEI message. A method comprising generating a face image using the neural network based on the base image and the face information parameters. 2. The SEI message includes an identification number indicator to indicate whether the SEI message is used to encode the face image, The method according to Clause 1, wherein determining whether the SEI message is applied to the neural network for generating the face image includes determining whether the SEI message is used to encode the face image based on the identification number indicator. 3. The method according to Clause 2, wherein the identification number indicator indicates a generated face image filter. 4. The SEI message includes a mode indicator. The method according to any one of the claims 1 to 3, wherein determining the mode used for coding the facial image includes determining the mode based on the mode indicator. 5. The method according to Clause 4, wherein the SEI message further includes a parameter indicator corresponding to the mode indicator, and the facial information parameter is determined based on the parameter indicator. 6. The method according to Clause 4, wherein the mode indicator indicates at least one of 2D facial landmarks, 2D keypoints, consistency regions, 3D keypoints, compact features, or facial semantics as the mode. 7. The method according to Clause 4, wherein the mode indicator includes at least one of the following: a coordinate indicator for indicating whether the SEI message includes coordinate parameters for encoding the face image; a matrix indicator for indicating whether the SEI message includes matrix parameters for encoding the face image; or a semantic indicator for indicating whether the SEI message includes semantic parameters for encoding the face image. 8. In response to the coordinate indicator indicating that the SEI message contains coordinate parameters, the SEI message further includes a parameter indicator for transmitting the coordinate parameters. The method according to Clause 7, wherein the coordinate parameters are determined as the face information parameters based on the parameter indicator. 9. In response to the matrix indicator indicating that the SEI message includes matrix parameters, the SEI message further includes a parameter indicator for communicating the matrix parameters. The method according to Clause 7, wherein the matrix parameters are determined as the face information parameters based on the parameter indicator. 10. In response to the semantic indicator indicating that the SEI message contains the semantic parameter, the SEI message further includes a parameter indicator for communicating the semantic parameter. The method according to Clause 7, wherein the semantic parameter is determined as the facial information parameter based on the parameter indicator. 11. The method according to Clause 7, wherein the coordinate parameters include at least one of a 2D facial landmark, a 2D keypoint, a consistency region, or a 3D keypoint. 12. The method according to Clause 7, wherein the matrix parameter includes at least one of an affine translation matrix, a covariance matrix, a translation matrix, a rotation matrix, or a compact feature matrix. 13. The method according to Clause 7, wherein the semantic parameter includes at least one of the semantic descriptions of a mouth, an eye, or a head. 14. The method according to Clause 7, wherein the matrix parameter includes at least one of an affine translation matrix, a covariance matrix, a compact feature matrix, or a semantic matrix. 15. The method according to Clause 14, wherein the semantic matrix includes at least one of a mouth matrix, an eye matrix, a head rotation matrix, a head translation matrix, or a head position matrix. 16. The SEI message includes a key indicator for identifying the target neural network used to generate the facial information parameters. The step of determining whether the SEI message is applied to the neural network for generating the face image is, The method according to any one of the clauses 1 to 15, comprising determining whether the neural network matches the target neural network based on the key indicator. 17. The SEI message includes an order indicator for indicating the display order of the face images, and the method is The method according to any one of the clauses 1 to 16, further comprising the step of displaying the facial images in the aforementioned display order. 18. The neural network includes a first subnetwork and a second subnetwork, The step of generating the face image based on the base image and the face information parameters using the neural network, The first subnetwork converts the facial information parameters into parameters in a predetermined format, The method according to any one of the claims 1 to 17, comprising generating the face image based on the parameters of the predetermined format and the base image using the second subnetwork. 19. The method according to clause 18, wherein the predetermined format is a flow map. 20. The SEI message includes a key presence indicator to indicate the presence of a key indicator. The step of determining whether the SEI message is applied to the neural network for generating the face image is, In response to the key presence indicator indicating the presence of the key indicator, the key indicator is obtained from the SEI message to identify the target neural network used to generate the face information parameters. The method according to clause 18, comprising determining whether the neural network matches the target neural network based on the key indicator. 21. The SEI message further includes a neural network indicator for indicating the target neural network, The aforementioned method, The method according to clause 20, further comprising configuring the neural network based on the neural network indicator in response to the key presence indicator indicating the absence of the key indicator. 22. The method according to Clause 21, wherein the neural network indicator includes at least one neural network mode indicator for indicating the format of the target neural network, or a neural network type indicator for indicating the type of the target neural network. 23. The parameters of the predetermined format are A predetermined number of key points used to describe facial features, A predetermined number of matrices used to describe the aforementioned facial features, or The method according to clause 18, comprising at least one matrix of a predetermined size used to describe the facial features. 24. The first subnetwork converts the facial information parameters into parameters in the predetermined format. Decoding syntax elements that indicate the difference between the parameters of the predetermined format and the parameters of the format associated with the base image, The method according to Clause 18, comprising determining the parameters of the predetermined format based on the difference and the parameters of the format associated with the base image. 25. The SEI message includes a base image indicator for indicating whether the SEI message is associated with the base image, and a neural network presence indicator for indicating whether the SEI message indicates information of the first subnetwork, The aforementioned method, The method according to clause 18, further comprising setting up the first subnetwork based on the neural network indicator in response to the base image indicator indicating that the SEI message is associated with a base image, and the neural network presence indicator indicating that the SEI message indicates the information of the first subnetwork. 26. A method for encoding a video sequence into a bitstream, Receiving a video sequence, The process involves encoding one or more images from the aforementioned video sequence to generate a bitstream, Encoding a base image and an additional extended information (SEI) message from one or more images, including generating an SEI message that indicates a mode used for encoding a face image and corresponding face information parameters, A method wherein the bitstream is used to generate the face image by a neural network based on the base image and the face information parameters. 27. The method according to clause 26, wherein the SEI message includes an identification number indicator for indicating whether the SEI message is used to encode the facial image. 28. The method according to clause 27, wherein the identification number indicator indicates a generated face image filter. 29. The method according to any one of the clauses 26 to 28, wherein the SEI message includes a mode indicator for indicating the mode used to encode the facial image. 30. The method according to clause 29, wherein the SEI message further includes a parameter indicator corresponding to the mode indicator for transmitting the facial information parameters. 31. The method according to clause 30, wherein the mode indicator indicates at least one of 2D facial landmarks, 2D keypoints, consistency regions, 3D keypoints, compact features, or facial semantics as the mode. 32. The method according to Clause 29, wherein the mode indicator includes at least one of the following: a coordinate indicator for indicating whether the SEI message includes coordinate parameters for encoding the face image; a matrix indicator for indicating whether the SEI message includes matrix parameters for encoding the face image; or a semantic indicator for indicating whether the SEI message includes semantic parameters for encoding the face image. 33. The method according to clause 32, wherein, in response to the coordinate indicator indicating that the SEI message contains coordinate parameters, the SEI message further includes a parameter indicator for transmitting the coordinate parameters. 34. The method according to clause 32, wherein the matrix indicator indicates that the SEI message contains a matrix parameter, and the SEI message further includes a parameter indicator for communicating the matrix parameter. 35. The method according to clause 32, wherein, in response to the semantic indicator indicating that the SEI message contains the semantic parameter, the SEI message further includes a parameter indicator for communicating the semantic parameter. 36. The method according to Clause 32, wherein the coordinate parameters include at least one of a 2D facial landmark, a 2D keypoint, a consistency region, or a 3D keypoint. 37. The method according to clause 32, wherein the matrix parameter includes at least one of an affine translation matrix, a covariance matrix, a translation matrix, a rotation matrix, or a compact feature matrix. 38. The method according to Clause 32, wherein the semantic parameter includes at least one of the semantic descriptions of a mouth, an eye, or a head. 39. The method according to clause 32, wherein the matrix parameter includes at least one of an affine translation matrix, a covariance matrix, a compact feature matrix, or a semantic matrix. 40. The method according to Clause 39, wherein the semantic matrix includes at least one of a mouth matrix, an eye matrix, a head rotation matrix, a head translation matrix, or a head position matrix. 41. The method according to any one of the clauses 26 to 40, wherein the SEI message includes a key indicator for identifying the neural network. 42. The method according to any one of the clauses 26 to 41, wherein the SEI message includes an order indicator for indicating the display order of the face images. 43. The aforementioned neural network, A first subnetwork configured to convert the facial information parameters into parameters in a predetermined format, The method according to any one of the claims 26 to 42, comprising a second subnetwork configured to generate the face image based on the parameters of the predetermined format and the base image. 44. The method according to clause 43, wherein the predetermined format is a flow map. 45. The SEI message includes a key presence indicator to indicate the presence of a key indicator. The method according to clause 43, wherein, in response to the key presence indicator indicating the presence of the key indicator, the SEI message further includes the key indicator for identifying the neural network. 46. ​​In response to the key presence indicator indicating the absence of the key indicator, the SEI message further includes a neural network indicator for indicating the neural network. The method according to clause 45, wherein the neural network indicator is used to configure a neural network for decoding the bitstream. 47. The method according to Clause 46, wherein the neural network indicator includes at least one neural network mode indicator for indicating the format of the neural network, or a neural network type indicator for indicating the type of the neural network. 48. The parameters of the predetermined format are A predetermined number of key points used to describe facial features, A predetermined number of matrices used to describe the aforementioned facial features, or The method according to clause 43, comprising at least one matrix of a predetermined size used to describe the facial features. 49. The method according to clause 43, wherein the SEI message includes syntax elements indicating the difference between the parameters of the predetermined format and the parameters of the format associated with the base image. 50. The SEI message includes a base image indicator to indicate whether the SEI message is associated with a base image, The method according to Clause 43, further comprising a neural network presence indicator for indicating whether the SEI message indicates information of the first subnetwork by the SEI message, in response to the base image indicator indicating that the SEI message is associated with a base image. 51. The aforementioned SEI message is In response to the neural network presence indicator indicating that the SEI message indicates the information of the first subnetwork, the system further includes a neural network indicator for indicating the neural network, The method according to clause 43, wherein the neural network indicator is used to configure a neural network for decoding the bitstream. 52. A non-temporary computer-readable storage medium for storing a video bitstream, wherein the bitstream is It includes a base image and an Additional Extended Information (SEI) message indicating the mode used for encoding the face image and the corresponding face information parameters. A non-temporary computer-readable storage medium in which the bitstream is used to generate the face image by a neural network based on the base image and the face information parameters. 53. A non-temporary computer-readable storage medium as described in Clause 52, which includes an identification number indicator for indicating whether the SEI message is used to encode the facial image. 54. A non-temporary computer-readable storage medium as described in Clause 53, wherein the identification number indicator indicates the generated face image filter. 55. A non-temporary computer-readable storage medium as described in Clause 54, which includes a mode indicator for indicating the mode used to encode the facial image for the SEI message. 56. A non-temporary computer-readable storage medium according to Clause 55, wherein the SEI message further includes a parameter indicator corresponding to the mode indicator for transmitting the facial information parameters. 57. A non-temporary computer-readable storage medium according to Clause 56, wherein the mode indicator indicates at least one of 2D facial landmarks, 2D keypoints, consistency regions, 3D keypoints, compact features, or facial semantics as the mode. 58. A non-temporary computer-readable storage medium according to Clause 55, wherein the mode indicator includes at least one of the following: a coordinate indicator for indicating whether the SEI message includes coordinate parameters for encoding the face image; a matrix indicator for indicating whether the SEI message includes matrix parameters for encoding the face image; or a semantic indicator for indicating whether the SEI message includes semantic parameters for encoding the face image. 59. A non-temporary computer-readable storage medium according to Clause 58, further comprising a parameter indicator for transmitting the coordinate parameters in response to the coordinate indicator indicating that the SEI message contains coordinate parameters. 60. A non-temporary computer-readable storage medium according to Clause 58, further comprising a parameter indicator for transmitting the matrix parameter to the SEI message in response to the matrix indicator indicating that the SEI message contains the matrix parameter. 61. A non-temporary computer-readable storage medium according to Clause 58, further comprising a parameter indicator for transmitting the semantic parameter in response to the semantic indicator indicating that the SEI message contains the semantic parameter. 62. A non-temporary computer-readable storage medium as described in Clause 58, wherein the coordinate parameters include at least one of a 2D facial landmark, a 2D keypoint, a consistency region, or a 3D keypoint. 63. A non-temporary computer-readable storage medium according to Clause 58, wherein the matrix parameters include at least one of an affine translation matrix, a covariance matrix, a translation matrix, a rotation matrix, or a compact feature matrix. 64. A non-temporary computer-readable storage medium as described in Clause 58, wherein the semantic parameter includes at least one of the semantic descriptions of a mouth, an eye, or a head. 65. A non-temporary computer-readable storage medium according to Clause 58, wherein the matrix parameters include at least one of an affine translation matrix, a covariance matrix, a compact feature matrix, or a semantic matrix. 66. A non-temporary computer-readable storage medium according to Clause 65, wherein the semantic matrix includes at least one of a mouth matrix, an eye matrix, a head rotation matrix, a head translation matrix, or a head position matrix. 67. A non-temporary computer-readable storage medium according to any one of clauses 52 to 66, wherein the SEI message includes a key indicator for identifying the neural network. 68. A non-temporary computer-readable storage medium according to any one of clauses 52 to 67, wherein the SEI message includes an order indicator for indicating the display order of the facial images. 69. The aforementioned neural network, A first subnetwork configured to convert the facial information parameters into parameters in a predetermined format, A non-temporary computer-readable storage medium according to any one of the clauses 52 to 68, comprising a second subnetwork configured to generate the face image based on the parameters of the predetermined format and the base image. 70. A non-temporary computer-readable storage medium as described in Clause 69, wherein the predetermined format is a flow map. 71. The SEI message includes a key presence indicator to indicate the presence of a key indicator. A non-temporary computer-readable storage medium according to Clause 69, wherein, in response to the key presence indicator indicating the presence of the key indicator, the SEI message further includes the key indicator for identifying the neural network. 72. In response to the key presence indicator indicating the absence of the key indicator, the SEI message further includes a neural network indicator for indicating the neural network, A non-temporary computer-readable storage medium as described in Clause 71, wherein the neural network indicator is used to configure a neural network for decoding the bitstream. 73. A non-temporary computer-readable storage medium according to Clause 72, wherein the neural network indicator includes at least one of a neural network mode indicator for indicating the format of the neural network, or a neural network type indicator for indicating the type of the neural network. 74. The parameters of the predetermined format are A predetermined number of key points used to describe facial features, A predetermined number of matrices used to describe the aforementioned facial features, or A non-temporary computer-readable storage medium according to Clause 69, comprising at least one matrix of a predetermined size used to describe the aforementioned facial features. 75. A non-temporary computer-readable storage medium according to Clause 69, wherein the SEI message includes syntax elements indicating the difference between the parameters of the predetermined format and the parameters of the format associated with the base image. 76. The SEI message includes a base image indicator to indicate whether the SEI message is associated with a base image, A non-temporary computer-readable storage medium according to Clause 69, further comprising a neural network presence indicator for indicating whether the SEI message indicates information of the first subnetwork by the SEI message, in response to the base image indicator indicating that the SEI message is associated with a base image. 77. The aforementioned SEI message, In response to the neural network presence indicator indicating that the SEI message indicates the information of the first subnetwork, the system further includes a neural network indicator for indicating the neural network, A non-temporary computer-readable storage medium as described in Clause 69, wherein the neural network indicator is used to configure a neural network for decoding the bitstream.

[0395] It should be noted that, in this specification, relational terms such as β€œfirst” and β€œsecond” are used solely to distinguish one entity or action from another, and do not require or imply any actual relationship or order between these entities or actions. Furthermore, the words β€œinclude,” β€œhave,” β€œcontain,” β€œinclude,” and other similar forms are semantically equivalent, and the items following any of these words are not exhaustive, nor are they limited to, the items listed.

[0396] As used herein, the term β€œor” shall, unless otherwise specified, encompass all possible combinations unless impractical. For example, if it is stated that a database may contain A or B, then unless otherwise specified or impractical, the database may contain A, or B, or A and B. As a second example, if it is stated that a database may contain A, B, or C, then unless otherwise specified or impractical, the database may contain A, or B, or C, or A and B, or A and C, or B and C, or A, B and C.

[0397] It will be understood that the embodiments described above can be implemented by hardware, software (program code), or a combination of hardware and software. When implemented by software, it may be stored in the computer-readable medium described above. When the software is executed by a processor, it can perform the disclosed methods. The computing units and other functional units described in this disclosure can be implemented by hardware, software, or a combination of hardware and software. Those skilled in the art will understand that multiple modules / units described above may be combined into a single module / unit, or each of the modules / units described above may be further divided into multiple submodules / subunits.

[0398] In the aforementioned specification, embodiments have been described with reference to numerous specific details, which may vary from implementation to implementation. Certain modifications and changes can be made to the described embodiments. Other embodiments may become apparent to those skilled in the art by examining the specification and practices of the invention disclosed herein. The specification and examples are for illustrative purposes only, and the true scope and spirit of the invention are intended to be shown by the following claims. Furthermore, the order of steps shown in the drawings is for illustrative purposes only and is not intended to limit the sequence of any particular steps. Thus, those skilled in the art will understand that these steps can be performed in different orders while implementing the same method.

[0399] The drawings and specification disclose exemplary embodiments; however, many variations and modifications can be made to these embodiments. Therefore, whereever specific terms are used, they are used in a general and descriptive sense only and are not intended to be limiting.

Claims

1. A method for generating facial images, Receiving a bitstream and The encoded information of the bitstream is decoded to obtain the base image and the additional extended information (SEI) message, The determination of whether the aforementioned SEI message is applied to a neural network for generating a facial image, In response to the application of the SEI message to the neural network for generating the face image, the mode used for encoding the face image and the corresponding face information parameters are determined based on the SEI message. A method comprising generating a face image using the neural network based on the base image and the face information parameters.

2. The SEI message includes an identification number indicator to indicate whether the SEI message is used to encode the facial image. The method according to claim 1, wherein determining whether the SEI message is applied to the neural network for generating the face image includes determining whether the SEI message is used to encode the face image based on the identification number indicator.

3. The method according to claim 2, wherein the identification number indicator indicates a generated face image filter.

4. The aforementioned SEI message includes a mode indicator. The method according to claim 1, wherein determining the mode used for encoding the facial image includes determining the mode based on the mode indicator.

5. The method according to claim 4, wherein the SEI message further includes a parameter indicator corresponding to the mode indicator, and the face information parameter is determined based on the parameter indicator.

6. The method according to claim 4, wherein the mode indicator indicates at least one of 2D facial landmarks, 2D keypoints, consistency regions, 3D keypoints, compact features, or facial semantics as the mode.

7. The method according to claim 4, wherein the mode indicator includes at least one of the following: a coordinate indicator for indicating whether the SEI message includes coordinate parameters for encoding the face image; a matrix indicator for indicating whether the SEI message includes matrix parameters for encoding the face image; or a semantic indicator for indicating whether the SEI message includes semantic parameters for encoding the face image.

8. In response to the coordinate indicator indicating that the SEI message contains coordinate parameters, the SEI message further includes a parameter indicator for transmitting the coordinate parameters. The method according to claim 7, wherein the coordinate parameters are determined as the face information parameters based on the parameter indicator.

9. In response to the matrix indicator indicating that the SEI message includes matrix parameters, the SEI message further includes a parameter indicator for communicating the matrix parameters. The method according to claim 7, wherein the matrix parameters are determined as the face information parameters based on the parameter indicator.

10. A method for encoding a video sequence into a bitstream, Receiving a video sequence, The process involves encoding one or more images from the aforementioned video sequence to generate a bitstream, Encoding a base image and an additional extended information (SEI) message from one or more images, including generating an SEI message that indicates a mode used for encoding a face image and corresponding face information parameters, A method wherein the bitstream is used to generate the face image by a neural network based on the base image and the face information parameters.

11. The method according to claim 10, wherein the SEI message includes an identification number indicator for indicating whether the SEI message is used to encode the facial image.

12. The method according to claim 11, wherein the identification number indicator indicates a generated face image filter.

13. The method according to claim 10, wherein the SEI message includes a mode indicator for indicating the mode used to encode the facial image.

14. The method according to claim 13, wherein the SEI message further includes a parameter indicator corresponding to the mode indicator for transmitting the facial information parameters.

15. The method according to claim 14, wherein the mode indicator indicates at least one of 2D facial landmarks, 2D keypoints, consistency regions, 3D keypoints, compact features, or facial semantics as the mode.

16. The method according to claim 13, wherein the mode indicator includes at least one of the following: a coordinate indicator for indicating whether the SEI message includes coordinate parameters for encoding the face image; a matrix indicator for indicating whether the SEI message includes matrix parameters for encoding the face image; or a semantic indicator for indicating whether the SEI message includes semantic parameters for encoding the face image.

17. The method according to claim 16, wherein, in response to the coordinate indicator indicating that the SEI message contains coordinate parameters, the SEI message further includes a parameter indicator for transmitting the coordinate parameters.

18. The method according to claim 16, wherein, in response to the matrix indicator indicating that the SEI message includes matrix parameters, the SEI message further includes a parameter indicator for transmitting the matrix parameters.

19. A non-temporary computer-readable storage medium for storing a video bitstream, wherein the bitstream is It includes a base image and an Additional Extended Information (SEI) message indicating the mode used for encoding the face image and the corresponding face information parameters, A non-temporary computer-readable storage medium in which the bitstream is used to generate the face image by a neural network based on the base image and the face information parameters.

20. The non-temporary computer-readable storage medium according to claim 19, wherein the SEI message includes an identification number indicator for indicating whether the SEI message is used to encode the facial image.