Method for video generation and compression and non-temporary computer-readable storage medium
By integrating SEI messages and deep learning techniques into video coding systems, the method addresses the challenge of high compression efficiency in advanced video coding standards, achieving improved compression and quality in video encoding and decoding.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ALIBABA (CHINA) CO LTD
- Filing Date
- 2024-04-12
- Publication Date
- 2026-06-22
AI Technical Summary
Existing video coding standards face challenges in achieving high compression efficiency, particularly with the development of advanced standards like VVC/H.266, which require improved methods for encoding and decoding video sequences to reduce storage and transmission bandwidth while maintaining quality.
The method involves encoding and decoding video sequences by incorporating Supplemental Enhancement Information (SEI) messages, allowing for enhanced video generation and compression through block-based coding processes, utilizing deep learning techniques, and integrating these processes into video coding systems to improve compression efficiency.
This approach enhances video coding efficiency, enabling higher compression rates with maintained or improved quality, supporting advanced standards like VVC/H.266 by optimizing encoding and decoding processes.
Smart Images

Figure 2026520081000001_ABST
Abstract
Description
Technical Field
[0001] (Cross - reference to Related Applications) This disclosure claims the benefit of priority to U.S. Provisional Application No. 63 / 496,049, filed on April 14, 2023, U.S. Provisional Application No. 63 / 511,897, filed on July 15, 2023, and U.S. Application No. 18 / 628,002, filed on April 5, 2024, all of which are hereby incorporated by reference in their entirety.
[0002] This disclosure generally relates to video processing, and more specifically, to methods and non - transient computer - readable storage media for video generation compression.
Background Art
[0003] Video is a series of still pictures (or "frames") that capture visual information. To reduce memory storage and transmission bandwidth, video can be compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, which are most commonly based on prediction, transformation, quantization, entropy coding, and in - loop filtering. Video coding standards that specify particular video coding formats, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, have been developed by standardization organizations. As more advanced video coding techniques are adopted in video standards, the coding efficiency of new video coding standards becomes higher and higher.
Summary of the Invention
[0004] Embodiments of the present disclosure provide a method for decoding a bitstream to obtain one or more pictures of a video stream. The method includes the steps of receiving a bitstream and decoding the bitstream to obtain one or more pictures. The decoding step includes decoding a picture unit containing one or more supplemental extension information (SEI) messages and generating the one or more pictures based on a key picture and the one or more SEI messages, respectively.
[0005] Embodiments of the present disclosure provide a method for encoding a video sequence into a bitstream. The method includes the steps of receiving a video sequence, encoding one or more pictures of the video sequence, and generating a bitstream. The encoding step includes encoding one or more supplemental extension information (SEI) messages corresponding to one or more pictures, and encoding the one or more SEI messages into picture units. Embodiments of the present disclosure provide a non-temporary, computer-readable storage medium for storing a bitstream of video. The bitstream includes picture units, each containing one or more Supplemental Enhancement Information (SEI) messages, which are used to generate one or more frames.
[0006] Embodiments and various aspects of this disclosure are shown in the following detailed description and accompanying drawings. Various features shown in the drawings are not depicted to scale. [Brief explanation of the drawing]
[0007] [Figure 1] This is a schematic diagram illustrating an exemplary system for coding image data according to some embodiments of the present disclosure.
[0008] [Figure 2A] This is a schematic diagram illustrating an exemplary block-based coding process according to some embodiments of the present disclosure.
[0009] [Figure 2B] This is a schematic diagram illustrating another exemplary block-based coding process according to some embodiments of the present disclosure.
[0010] [Figure 3A] This is a schematic diagram illustrating an exemplary block-based decoding process according to some embodiments of the present disclosure.
[0011] [Figure 3B] This is a schematic diagram illustrating another exemplary block-based decoding process according to some embodiments of the present disclosure.
[0012] [Figure 4] This is a block diagram of an exemplary device for encoding or decoding video, consistent with embodiments of the present invention.
[0013] [Figure 5] The following are exemplary structures of coded video sequences (CVS) according to some embodiments of the present disclosure.
[0014] [Figure 6] The following shows an exemplary picture unit (PU) structure according to some embodiments of the present disclosure.
[0015] [Figure 7] Table 1 shows an exemplary syntax structure of a network abstraction layer (NAL) unit according to some embodiments of this disclosure.
[0016] [Figure 8] Table 2 shows an exemplary syntax structure of rbsp_trailing_bits() according to some embodiments of this disclosure.
[0017] [Figure 9-1]Table 3 showing an exemplary association between a raw byte sequence payload (RBSP) syntax structure and a NAL unit according to some embodiments of the present disclosure is shown.
[0018] [Figure 9-2] Table 3 showing an exemplary association between a raw byte sequence payload (RBSP) syntax structure and a NAL unit according to some embodiments of the present disclosure is shown.
[0019] [Figure 10] Table 4 showing an exemplary syntax structure of a NAL unit header according to some embodiments of the present disclosure is shown.
[0020] [Figure 11] FIG. is a schematic diagram showing an exemplary deep learning-based video generation compression framework according to some embodiments of the present disclosure.
[0021] [Figure 12] An exemplary deep learning-based video generation compression framework according to some embodiments of the present disclosure is shown.
[0022] [Figure 13] FIG. is a schematic diagram showing a general encoder-decoder generation compression framework for 3DMM-assisted conversational face video according to some embodiments of the present disclosure.
[0023] [Figure 14] An exemplary SEI message having a plurality of feature parameter sets according to some embodiments of the present disclosure is shown.
[0024] [Figure 15] A syntax signaling process for one SEI message having a plurality of feature parameter sets according to some embodiments of the present disclosure is shown.
[0025] [Figure 16]A flowchart illustrating an exemplary method for generating one or more frames according to some embodiments of this disclosure is shown.
[0026] [Figure 17] A first set of exemplary picture unit structures according to some embodiments of the present disclosure is shown.
[0027] [Figure 18] Table 5 shows exemplary feature SEI message syntax according to some embodiments of this disclosure.
[0028] [Figure 19] A second set of exemplary picture unit structures according to some embodiments of the present disclosure is shown.
[0029] [Figure 20] A third set of exemplary picture unit structures according to some embodiments of the present disclosure is shown.
[0030] [Figure 21] A fourth set of exemplary picture unit structures according to some embodiments of the present disclosure is shown.
[0031] [Figure 22] Table 6 shows exemplary syntax of PUD RBSP according to some embodiments of this disclosure.
[0032] [Figure 23] Table 7 shows exemplary interpretations of pud_pic_type according to some embodiments of this disclosure.
[0033] [Figure 24] Table 8 shows exemplary feature SEI message syntax according to some embodiments of this disclosure. [Modes for carrying out the invention]
[0034] Herein, exemplary embodiments shown in the accompanying drawings will be described in detail. The following description refers to the accompanying drawings, in which, unless otherwise noted, the same number in different drawings represents the same or similar elements. The embodiments described below in the description of exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, these embodiments are merely examples of apparatus and methods consistent with aspects related to the present invention enumerated in the accompanying claims. Specific aspects of this disclosure will be described in more detail below. In the event of any conflict between terms and / or definitions incorporated by reference and those provided herein, the terms and definitions provided herein shall prevail.
[0035] The Joint Video Expert Team (JVET), comprised of the ITU-T Video Coding Expert Group (ITU-TVCEG) and the ISO / IEC Video Expert Group (ISO / IECMPEG), is currently developing the Multipurpose Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 with half the bandwidth.
[0036] To achieve the same subjective quality as HEVC / H.265 with half the bandwidth, JVET has developed a technology that surpasses HEVC using Joint Search Model (JEM) reference software. Because coding techniques are incorporated into JEM, JEM achieves significantly higher coding performance than HEVC.
[0037] The VVC standard is a relatively recent development and continues to evolve to include even more coding techniques that offer superior compression performance. VVC is based on the same hybrid video coding system used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.
[0038] Video is a series of still pictures (or "frames") arranged in chronological order to store visual information. A video capture device (e.g., a camera) can capture and store these pictures in chronological order, and a video playback device (e.g., a television, computer, smartphone, tablet computer, video player, or any end-user device with display capabilities) can display these pictures in chronological order. Depending on the application, such as for surveillance, conferences, or live broadcasts, a video capture device can also transmit the captured video in real time to a video playback device (e.g., a computer with a monitor).
[0039] To reduce the storage space and transmission bandwidth required for such applications, video can be compressed before storage and transmission and decompressed before display. Compression and decompression can be performed by software executed by a processor (e.g., a general-purpose computer processor) or by dedicated hardware. The module for compression is generally called an "encoder," and the module for decompression is generally called a "decoder." Encoders and decoders can be collectively called a "codec." Encoders and decoders can be implemented as various suitable hardware, software, or a combination thereof. For example, a hardware implementation of an encoder and decoder may include circuits such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. A software implementation of an encoder and decoder may include program code, computer-executable instructions, firmware, or a suitable computer implementation algorithm or process fixed in a computer-readable medium. Video compression and decompression can be implemented using various algorithms and standards such as MPEG-1, MPEG-2, MPEG-4, and the H.26x series. Depending on the application, a codec can decompress video from a first coding standard and recompress the decompressed video using a second coding standard; in this case, the codec may be called a "transcoder."
[0040] The video coding process can identify and retain useful information that can be used to reconstruct the picture, while ignoring information that is not important for reconstruction. If the ignored, non-essential information cannot be fully reconstructed, such an encoding process may be called "lossy." Otherwise, it may be called "lossy." Most encoding processes are lossy, which is a trade-off to reduce the required storage space and transmission bandwidth.
[0041] Useful information about the picture being encoded (referred to as the "current picture") includes changes to the reference picture (e.g., a previously encoded and reconstructed picture). Such changes may include changes in pixel position, brightness, or color, but changes in position are of greatest interest. Changes in the position of a group of pixels representing an object can reflect the movement of the object between the reference picture and the current picture.
[0042] A picture coded without referencing another picture (i.e., it is its own reference picture) is called an "I-picture". A picture is called a "P-picture" if some or all of its blocks (e.g., blocks that generally refer to a portion of a video picture) are predicted using one reference picture in intra-prediction or inter-prediction (e.g., uni-prediction). A picture is called a "B-picture" if at least one block in it is predicted using two reference pictures (e.g., bi-prediction).
[0043] Figure 1 is a block diagram of a system 100 for coding image data according to some disclosed embodiments. Image data may include images (also called “pictures” or “frames”), a set of images, or video. An image is a still picture. A set of images may or may not be spatially or temporally related. A video is a set of images arranged in chronological order.
[0044] As shown in Figure 1, the system 100 includes a source device 120 that provides encoded video data to be decoded at a later time by a destination device 140. In accordance with the disclosed embodiments, the source device 120 and the destination device 140 may each include any of a wide range of devices, including desktop computers, notebook (e.g., laptop) computers, servers, tablet computers, set-top boxes, mobile phones, vehicles, cameras, image sensors, robots, televisions, wearable devices (e.g., smartwatches or wearable cameras), display devices, digital media players, video game consoles, video streaming devices, and the like. The source device 120 and the destination device 140 may be equipped for wireless or wired communication.
[0045] Referring to Figure 1, the source device 120 may include an image / video preprocessor 122, an image / video encoder 124, and an output interface 126. The destination device 140 may include an input interface 142, an image / video decoder 144, and a machine vision application 146. The image / video encoder 124 encodes the input bitstream and outputs the encoded bitstream 162 via the output interface 126. The encoded bitstream 162 is transmitted through the communication medium 160 and received by the input interface 142. The image / video decoder 144 then decodes the encoded bitstream 162 to produce the decoded data.
[0046] More specifically, the source device 120 may further include various devices (not shown) for providing source image data to be processed by the image / video encoder 124. These devices may include image / video capture devices such as cameras, image / video archive or storage devices containing previously captured images / videos, or image / video feed interfaces for receiving images / videos from image / video content providers.
[0047] The image / video encoder 124 and the image / video decoder 144 can each be implemented as one or more suitable encoder or decoder circuits, including one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. If encoding or decoding is partially implemented in software, the image / video encoder 124 or image / video decoder 144 can perform the technology of this disclosure by storing instructions for the software in a suitable non-temporary computer-readable medium and executing these instructions in hardware using one or more processors. The image / video encoder 124 or image / video decoder 144 can each be included in one or more encoders or decoders, any of which can be integrated as part of a combined encoder / decoder (CODEC) in their respective devices.
[0048] The image / video encoder 124 and image / video decoder 144 may operate according to any video coding standard, such as Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), Multipurpose Video Coding (VVC), AOMedia Video 1 (AV1), Joint Photographic Professional Group (JPEG), and Video Professional Group (MPEG). Alternatively, the image / video encoder 124 and image / video decoder 144 may be customized devices that do not conform to existing standards. Although not shown in Figure 1, in some embodiments, the image / video encoder 124 and image / video decoder 144 may be integrated with an audio encoder and decoder, respectively, and may include a suitable MUX-DEMUX unit or other hardware and software to handle the encoding of both audio and video on a common data stream or separate data streams.
[0049] The output interface 126 may include any type of medium or device capable of transmitting the encoded bitstream 162 from the source device 120 to the destination device 140. For example, the output interface 126 may include a transmitter or transceiver configured to transmit the encoded bitstream 162 directly from the source device 120 to the destination device 140 in real time. The encoded bitstream 162 may be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device 140.
[0050] The communication medium 160 may include a temporary medium such as a wireless broadcast or a wired network transmission. For example, the communication medium 160 may include a radio frequency (RF) spectrum or one or more physical transmission lines (e.g., a cable). The communication medium 160 may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. In some embodiments, the communication medium 160 may include a router, a switch, a base station, or any other equipment that may be useful in facilitating communication from the source device 120 to the destination device 140. For example, a network server (not shown) may receive an encoded bitstream 162 from the source device 120, for example, via network transmission, and provide the encoded bitstream 162 to the destination device 140.
[0051] The communication medium 160 may also be in the form of a storage medium (e.g., a non-temporary storage medium), such as a hard disk, flash drive, compact disc, digital video disc, Blu-ray disc, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded image data. In some embodiments, a computing device of a media manufacturing facility, such as a disc stamping machine, may receive encoded image data from the source device 120 and generate a disc containing the encoded video data.
[0052] The input interface 142 may include any type of medium or device capable of receiving information from the communication medium 160. The received information includes an encoded bitstream 162. For example, the input interface 142 may include a receiver or transceiver configured to receive the encoded bitstream 162 in real time.
[0053] System 100 can be configured to perform video encoding and decoding based on block-based video compression technology, deep learning-based video compression technology, conversational face video compression technology, and the like.
[0054] Video coding has multiple operational stages, examples of which are shown in Figures 2A-2B and 3A-3B. At each stage, the size of the basic processing unit may still be too large to process, and therefore it may be further divided into segments referred to in this disclosure as “basic processing subunits.” In some embodiments, a basic processing subunit may be referred to as a “block” in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC) or as a “coding unit” (“CU”) in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). A basic processing subunit may be the same size as or smaller than a basic processing unit. Like a basic processing unit, a basic processing subunit is a logical unit that can contain groups of different types of video data (e.g., Y, Cb, Cr, and associated syntax elements) stored in computer memory (e.g., a video frame buffer). Any operation performed on a basic processing subunit can be repeated on its luma and chroma components, respectively. Note that such divisions may be performed at further levels depending on processing needs. Also note that different schemes may be used to divide the basic processing unit at different stages.
[0055] For example, in the mode determination stage (an example of which is shown in Figure 2B), the encoder can determine which prediction mode (e.g., intra-picture prediction or inter-picture prediction) should be used for the base processing unit, which may be too large to make such a decision. The encoder can divide the base processing unit into multiple base processing subunits (such as CUs in H.265 / HEVC or H.266 / VVC) and determine the prediction type for each individual base processing subunit.
[0056] As another example, in the prediction stage (an example of which is shown in Figures 2A and 2B), the encoder can perform prediction operations at the level of the basic processing subunit (e.g., CU). However, in some cases, the basic processing subunit may still be too large to process. The encoder can further divide the basic processing subunit into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), at which level prediction operations can be performed.
[0057] As another example, in the transformation stage (examples of which are shown in Figures 2A and 2B), the encoder can perform transformation operations on residual basic processing subunits (e.g., CUs). However, in some cases, the basic processing subunit may still be too large to process. The encoder can further divide the basic processing subunit into smaller segments (e.g., called "transformation blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), at which level transformation operations can be performed. Note that the division scheme for the same basic processing subunit may differ between the prediction stage and the transformation stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transformation blocks of the same CU may have different sizes and numbers.
[0058] In some implementations, a picture can be divided into several regions for processing to provide parallel processing and error tolerance for video encoding and decoding. This allows the encoding or decoding process for a given region of the picture to be independent of information from any other region of the picture. In other words, each region of the picture can be processed independently. This allows the codec to process different regions of the picture in parallel, thus improving coding efficiency. Furthermore, if data in a region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thus providing error tolerance. In some video coding standards, a picture can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles." It should also be noted that different pictures in a video sequence can have different partition schemes for dividing the picture into several regions.
[0059] Figure 2A is a schematic diagram of an exemplary encoding process 200A consistent with embodiments of the present disclosure. For example, the encoding process 200A can be performed by an encoder. As shown in Figure 2A, the encoder can encode a video sequence 202 into a video bitstream 228 according to process 200A. Each original picture of the video sequence 202 can be divided by the encoder into a basic processing unit, a basic processing subunit, or a region for processing. In some embodiments, the encoder can perform process 200A at the level of a basic processing unit for each original picture of the video sequence 202. For example, the encoder can perform process 200A iteratively, in which case the encoder can encode one basic processing unit in one iteration of process 200A. In some embodiments, the encoder can perform process 200A in parallel for each region of the original picture of the video sequence 202.
[0060] In Figure 2A, the encoder can generate predicted data 206 and predicted BPU 208 by sending the basic processing unit (called the "original BPU") of the original picture of the video sequence 202 to the prediction stage 204. The encoder can generate residual BPU 210 by subtracting predicted BPU 208 from original BPU. The encoder can generate quantization conversion coefficients 216 by sending residual BPU 210 to the conversion stage 212 and quantization stage 214. The encoder can generate video bitstream 228 by sending predicted data 206 and quantization conversion coefficients 216 to the binary encoding stage 226. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be called the "forward path". During process 200A, after the quantization stage 214, the encoder can generate a reconstructed residual BPU 222 by sending the quantization conversion coefficients 216 to the inverse quantization stage 218 and the inverse conversion stage 220. The encoder can then generate a prediction criterion 224, which will be used for the next iteration of process 200A in the prediction stage 204, by adding the reconstructed residual BPU 222 to the prediction BPU 208. Components 218, 220, 222, and 224 of process 200A may be referred to as the “reconstruction path”. The reconstruction path can be used to ensure that both the encoder and decoder use the same reference data for prediction.
[0061] The encoder can iteratively perform process 200A to encode each original BPU of the original picture (in the forward path) and generate a predictive criterion 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all original BPUs of the original picture, the encoder can proceed to encoding the next picture in the video sequence 202.
[0062] Referring to process 200A, the encoder can receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term “receive” can mean any action by any means of receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or inputting data.
[0063] In prediction stage 204, in the current iteration, the encoder receives the original BPU and prediction criterion 224 and can perform prediction operations to generate prediction data 206 and prediction BPU 208. The prediction criterion 224 can be generated from the reconfiguration path in the previous iteration of process 200A. The objective of prediction stage 204 is to reduce information redundancy by extracting prediction data 206 from the prediction data 206 and prediction criterion 224 that can be used to reconfigure the original BPU as prediction BPU 208.
[0064] Ideally, the predicted BPU 208 can be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 is usually slightly different from the original BPU. To record such differences, the encoder can generate the residual BPU 210 by subtracting the predicted BPU 208 from the original BPU after generating it. For example, the encoder can subtract the pixel values (e.g., grayscale values or RGB values) of the predicted BPU 208 from the corresponding pixel values of the original BPU. Each pixel of the residual BPU 210 may have a residual value resulting from such a subtraction between the corresponding pixels of the original BPU and the predicted BPU 208. Compared to the original BPU, the predicted data 206 and residual BPU 210 may have fewer bits, but they can be used to reconstruct the original BPU without a significant loss of quality. Thus, the original BPU is compressed.
[0065] To further compress the residual BPU210, in the conversion stage 212, the encoder can reduce the spatial redundancy of the residual BPU210 by decomposing it into a set of secondary original "base patterns," each associated with a "conversion coefficient." The base patterns can have the same size (e.g., the size of the residual BPU210). Each base pattern can represent the fluctuating frequency component (e.g., brightness fluctuation frequency) of the residual BPU210. No base pattern can be reproduced from any combination of any other base patterns (e.g., a linear combination). In other words, the decomposition allows the fluctuations of the residual BPU210 to be decomposed into the frequency domain. Such decomposition is analogous to the discrete Fourier transform of a function, where the base patterns are analogous to the basis functions of the discrete Fourier transform (e.g., trigonometric functions), and the conversion coefficients are analogous to the coefficients associated with the basis functions.
[0066] Different transformation algorithms can use different base patterns. Various transformation algorithms can be used in the transformation stage 212, such as discrete cosine transform and discrete sine transform. The transformation in the transformation stage 212 is reversible; that is, the encoder can reconstruct the residual BPU 210 by performing the inverse operation of the transformation (called the "inverse transformation"). For example, to reconstruct the pixels of the residual BPU 210, the inverse transformation might involve multiplying the values of the corresponding pixels in the base pattern by their respective associated coefficients and adding their products to produce a weighted sum. In video coding standards, both the encoder and decoder can use the same transformation algorithm (and therefore the same base pattern). Therefore, the encoder can record only the transformation coefficients, and the decoder can reconstruct the residual BPU 210 from these transformation coefficients without receiving the base pattern from the encoder. While the transformation coefficients may have fewer bits compared to the residual BPU 210, they can be used to reconstruct the residual BPU 210 without significant quality degradation. Thus, the residual BPU 210 is further compressed.
[0067] The encoder can further compress the conversion coefficients in the quantization stage 214. In the conversion process, different base patterns can represent different fluctuation frequencies (e.g., brightness fluctuation frequencies). Generally, the human eye is better at recognizing low-frequency fluctuations, so the encoder can ignore information about high-frequency fluctuations without causing significant quality degradation during decoding. For example, in the quantization stage 214, the encoder can generate quantization conversion coefficients 216 by dividing each conversion coefficient by an integer value (called the "quantization scale coefficient") and rounding the quotient to its nearest integer. After such an operation, some conversion coefficients for high-frequency base patterns may be converted to zero, and conversion coefficients for low-frequency base patterns may be converted to smaller integers. The encoder can ignore quantization conversion coefficients 216 with zero values, thereby further compressing the conversion coefficients. The quantization process is also reversible, in which the quantization conversion coefficients 216 can be reconstructed into conversion coefficients in the inverse operation of quantization (called "inverse quantization").
[0068] Because the encoder ignores the remainder of such division in rounding operations, the quantization stage 214 can be irreversible. Typically, the quantization stage 214 can cause the greatest information loss in process 200A. The greater the information loss, the fewer bits the quantization conversion coefficient 216 may require. To obtain different levels of information loss, the encoder can use different values for the quantization syntax elements or any other syntax elements of the quantization process.
[0069] In the binary coding stage 226, the encoder can encode the predicted data 206 and quantization conversion coefficients 216 using binary coding techniques such as entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other reversible or irreversible compression algorithm. In some embodiments, in addition to the predicted data 206 and quantization conversion coefficients 216, the encoder can encode other information in the binary coding stage 226, such as the prediction mode used in the prediction stage 204, syntax elements of the prediction operation, the conversion type in the conversion stage 212, syntax elements of the quantization process (e.g., quantization syntax elements), and encoder control syntax elements (e.g., bitrate control syntax elements). The encoder can use the output data from the binary coding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 may be further packetized for network transmission.
[0070] Referring to the reconstruction path of process 200A, in the inverse quantization stage 218, the encoder can generate reconstruction transformation coefficients by performing inverse quantization on the quantization transformation coefficients 216. In the inverse transformation stage 220, the encoder can generate reconstruction residual BPU 222 based on the reconstruction transformation coefficients. The encoder can add the reconstruction residual BPU 222 to the prediction BPU 208 to generate a prediction criterion 224 to be used in the next iteration of process 200A.
[0071] It should be noted that other variations of process 200A can be used to encode the video sequence 202. In some embodiments, the stages of process 200A may be executed in a different order by the encoder. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be divided into multiple stages. For example, the conversion stage 212 and the quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may omit one or more stages in Figure 2A.
[0072] Figure 2B shows a schematic diagram of another exemplary coding process 200B consistent with embodiments of the present disclosure. Process 200B may be a modification of process 200A. For example, process 200B may be used by an encoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared with process 200A, the forward path of process 200B further includes a mode determination stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconfiguration path of process 200B further includes a loop filter stage 232 and a buffer 234.
[0073] Generally, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-prediction") can predict the current BPU by using pixels from one or more already coded neighboring BPUs in the same picture. That is, the prediction criterion 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter-prediction") can predict the current BPU by using regions from one or more already coded pictures. That is, the prediction criterion 224 in temporal prediction can include coded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.
[0074] Referring to process 200B, in the forward path, the encoder performs prediction operations in the spatial prediction stage 2042 and the temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder can perform intra-prediction. For the original BPU of the encoded picture, the prediction criterion 224 can include one or more adjacent BPUs that are encoded (in the forward path) and reconstructed (in the reconstruction path) in the same picture. The encoder can generate a prediction BPU 208 by extrapolating adjacent BPUs. Extrapolation techniques may include, for example, linear extrapolation or interpolation, or polynomial extrapolation or interpolation. In some embodiments, the encoder can perform extrapolation at the pixel level, such as by extrapolating the values of the corresponding pixels for each pixel of the prediction BPU 208. The adjacent BPU used for extrapolation can be positioned from various directions relative to the original BPU, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., lower left, lower right, upper left, or upper right of the original BPU), or in any direction defined by the video coding standard used. In intra-prediction, the prediction data 206 may include, for example, the position of the adjacent BPU used (e.g., coordinates), the size of the adjacent BPU used, the extrapolation syntax elements, and the orientation of the adjacent BPU used relative to the original BPU.
[0075] As another example, in the temporal prediction stage 2044, the encoder can perform interpretation. For the original BPU of the current picture, the prediction criterion 224 may include one or more pictures (referred to as "reference pictures") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures may be encoded and reconstructed for each BPU. For example, the encoder may generate a reconstructed BPU by adding the reconstructed residual BPU 222 to the prediction BPU 208. Once all the reconstructed BPUs for the same picture have been generated, the encoder can generate a reconstructed picture as a reference picture. The encoder may perform a "motion estimation" operation to search for a matching region within the range of the reference picture (referred to as the "search window"). The position of the search window in the reference picture may be determined based on the position of the original BPU in the current picture. For example, the search window may be centered at a position in the reference picture that has the same coordinates as the original BPU in the current picture and may extend outward by a predetermined distance. When the encoder identifies a region in the search window that is similar to the original BPU (for example, by using a pixel recursion algorithm, a block matching algorithm, etc.), the encoder can determine such a region as a matching region. The matching region may have different dimensions from the original BPU (for example, smaller than, equal to, larger than, or a different shape from the original BPU). Because the reference picture and the current picture are separated in time on the timeline, the matching region can be considered to "move" to the position of the original BPU over time. The encoder can record the direction and distance of such movement as a "motion vector". If multiple reference pictures are used, the encoder can search for a matching region for each reference picture and determine the associated motion vector. In some embodiments, the encoder can assign weights to the pixel values of the matching region of each matching reference picture.
[0076] Motion estimation can be used to identify various types of motion, such as translation, rotation, and zooming. In interpretation, the prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference pictures, and the weights associated with the reference pictures.
[0077] To generate a predicted BPU 208, the encoder can perform a “motion compensation” operation. Motion compensation can be used to reconstruct the predicted BPU 208 based on prediction data 206 (e.g., motion vectors) and prediction criteria 224. For example, the encoder can move the matching region of a reference picture according to the motion vector, thereby allowing the encoder to predict the original BPU of the current picture. If multiple reference pictures are used, the encoder can move the matching region of the reference pictures according to the respective motion vector and average pixel value of the matching region. In some embodiments, if the encoder has assigned weights to the pixel values of the matching region of each matching reference picture, the encoder can add the weighted sum of the pixel values of the moved matching region.
[0078] Continuing to refer to the forward path of process 200B, after the spatial prediction 2042 and the temporal prediction stage 2044, in the mode determination stage 230, the encoder can select a prediction mode (e.g., either intra-prediction or inter-prediction) for the current iteration of process 200B. For example, the encoder can perform a rate-distortion optimization technique, in which the encoder can select a prediction mode based on the bit rate of a candidate prediction mode and the distortion of the reconstructed reference picture in the candidate prediction mode to minimize the value of the cost function. Depending on the selected prediction mode, the encoder can generate the corresponding prediction BPU 208 and predicted data 206.
[0079] In the reconstruction path of process 200B, if the intra-prediction mode is selected in the forward path, after generating the prediction criterion 224 (e.g., the current BPU encoded and reconstructed with the current picture), the encoder can send the prediction criterion 224 directly to the spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). The encoder can also send the prediction criterion 224 to the loop filter stage 232, where the encoder can apply loop filters to the prediction criterion 224 to reduce or remove distortions (e.g., blocking artifacts) introduced during the coding of the prediction criterion 224. In the loop filter stage 232, the encoder can apply various loop filtering techniques, such as deblocking, sample-adaptive offset, or adaptive loop filtering. The loop-filtered reference picture can be stored in the buffer 234 (or "Decoded Picture Buffer (DPB)") for later use (e.g., to use as an inter-prediction criterion picture for future pictures in the video sequence 202). The encoder can store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder can encode the syntax elements of the loop filter (e.g., the strength of the loop filter) along with the quantization transformation coefficients 216, the prediction data 206, and other information in the binary coding stage 226.
[0080] Figure 3A shows a schematic diagram of an exemplary decoding process 300A consistent with embodiments of the present disclosure. Process 300A may be a decompression process corresponding to the compression process 200A in Figure 2A. In some embodiments, process 300A may be analogous to the reconstruction path of process 200A. The decoder can decode the video bitstream 228 into a video stream 304 according to process 300A. The video stream 304 may be very similar to the video sequence 202. However, due to the loss of information in the compression and decompression processes (e.g., the quantization stage 214 in Figures 2A and 2B), the video stream 304 is usually not identical to the video sequence 202. Similar to processes 200A and 200B in Figures 2A and 2B, the decoder may execute process 300A at the level of the basic processing unit (BPU) for each picture encoded in the video bitstream 228. For example, the decoder can iteratively execute process 300A, in which case the decoder can decode one basic processing unit in one iteration of process 300A. In some embodiments, the decoder can execute process 300A in parallel for each region of the encoded picture in the video bitstream 228.
[0081] In Figure 3A, the decoder can send a portion of the video bitstream 228 associated with the basic processing unit of the encoded picture (called the "encoded BPU") to the binary decoding stage 302. In the binary decoding stage 302, the decoder can decode this portion into prediction data 206 and quantization conversion coefficients 216. The decoder can generate a reconstructed residual BPU 222 by sending the quantization conversion coefficients 216 to the inverse quantization stage 218 and the inverse transformation stage 220. The decoder can generate a prediction BPU 208 by sending the prediction data 206 to the prediction stage 204. The decoder can generate a prediction criterion 224 by adding the reconstructed residual BPU 222 to the prediction BPU 208. In some embodiments, the prediction criterion 224 can be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder can send the prediction criterion 224 to the prediction stage 204 to perform the prediction operation in the next iteration of process 300A.
[0082] The decoder can decode each encoding BPU of the encoded picture by iteratively performing process 300A and generate a prediction criterion 224 to encode the next encoding BPU of the encoded picture. After decoding all encoding BPUs of the encoded picture, the decoder can output the picture to the video stream 304 for display and proceed to decode the next encoded picture in the video bitstream 228.
[0083] In the binary decoding stage 302, the decoder can perform the inverse operation of the binary coding technique used by the encoder (e.g., entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in the binary decoding stage 302, the decoder can decode other information in addition to the predicted data 206 and quantization conversion coefficients 216, such as the prediction mode, syntax elements of the prediction operation, conversion type, syntax elements of the quantization process (e.g., quantization syntax elements), and encoder control syntax elements (e.g., bitrate control syntax elements). In some embodiments, if the video bitstream 228 is transmitted in packet form over the network, the decoder can depacketize the video bitstream 228 before sending it to the binary decoding stage 302.
[0084] Figure 3B shows a schematic diagram of another exemplary decoding process 300B consistent with embodiments of the present disclosure. Process 300B may be a modification of process 300A. For example, process 300B may be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared with process 300A, process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and further includes a loop filter stage 232 and a buffer 234.
[0085] In process 300B, for the encoding base processing unit ("current BPU") of the encoded picture being decoded ("current picture"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 may contain various types of data depending on which prediction mode was used by the encoder to encode the current BPU. For example, if intra-prediction was used by the encoder to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra-prediction, syntax elements of the intra-prediction operation, etc. Syntax elements of the intra-prediction operation may include, for example, the location (e.g., coordinates) of one or more adjacent BPUs used as reference, the size of the adjacent BPUs, syntax elements of extrapolation, the orientation of the adjacent BPUs relative to the original BPU, etc. As another example, if inter-prediction was used by the encoder to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter-prediction, syntax elements of the inter-prediction operation, etc. Syntax elements of an interpretation operation may include, for example, the number of reference pictures associated with the current BPU, the weights associated with each reference picture, the location (e.g., coordinates) of one or more matching regions in each reference picture, and one or more motion vectors associated with each matching region.
[0086] Based on the prediction mode indicator, the decoder can decide whether to perform a spatial prediction (e.g., intra-prediction) in the spatial prediction stage 2042 or a temporal prediction (e.g., inter-prediction) in the temporal prediction stage 2044. Details of performing such spatial or temporal predictions are shown in Figure 2B and will not be repeated below. After performing such spatial or temporal predictions, the decoder can generate a prediction BPU 208. The decoder can generate a prediction criterion 224 by adding the prediction BPU 208 and the reconstructed residual BPU 222, as shown in Figure 3A.
[0087] In process 300B, the decoder can send the prediction criterion 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 to perform the prediction operation in the next iteration of process 300B. For example, if the current BPU is decoded using intra-prediction in the spatial prediction stage 2042, after generating the prediction criterion 224 (e.g., the decoded current BPU), the decoder can send the prediction criterion 224 directly to the spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). If the current BPU is decoded using inter-prediction in the temporal prediction stage 2044, after generating the prediction criterion 224 (e.g., the reference picture with all BPUs decoded), the decoder can send the prediction criterion 224 to the loop filter stage 232 to reduce or remove distortion (e.g., blocking artifacts). The decoder can apply a loop filter to the prediction criterion 224 in the manner shown in Figure 2B. The loop-filtered reference picture can be stored in buffer 234 (e.g., a decoded picture buffer (DPB) in computer memory) for later use (e.g., for use as an inter-prediction reference picture for future encoded pictures of the video bitstream 228). The decoder may store one or more reference pictures in buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the prediction data may further include the syntax elements of the loop filter (e.g., the strength of the loop filter). In some embodiments, the prediction data includes the syntax elements of the loop filter if the prediction mode indicator of the prediction data 206 indicates that inter-prediction was used to encode the current BPU.
[0088] Figure 4 is a block diagram of an exemplary apparatus 400 for encoding or decoding video, consistent with embodiments of the present disclosure. As shown in Figure 4, the apparatus 400 may include a processor 402. When the processor 402 executes the instructions described herein, the apparatus 400 can become a machine dedicated to encoding or decoding video. The processor 402 may be any type of circuit capable of manipulating or processing information. For example, the processor 402 may include any number and any combination of such as a central processing unit (or "CPU"), graphics processing unit (or "GPU"), neural processing unit ("NPU"), microcontroller unit ("MCU"), optical processor, programmable logic controller, microcontroller, microprocessor, digital signal processor, intellectual property (IP) core, programmable logic array (PLA), programmable array logic (PAL), generic array logic (GAL), complex programmable logic device (CPLD), field programmable gate array (FPGA), system-on-a-chip (SoC), application-specific integrated circuit (ASIC), etc. In some embodiments, the processor 402 may be a set of processors grouped as a single logical component. For example, as shown in Figure 4, the processor 402 may include a plurality of processors, including processor 402a, processor 402b, and processor 402n.
[0089] The device 400 may also include a memory 404 configured to store data (e.g., instruction sets, computer code, intermediate data, etc.). For example, as shown in Figure 4, the stored data may include program instructions (e.g., program instructions for implementing stages of process 200A, 200B, 300A, or 300B) and processing data (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 can access the program instructions and processing data (e.g., via bus 410), execute the program instructions, and perform operations or manipulations on the processing data. The memory 404 may include a high-speed random-access memory or a non-volatile memory. In some embodiments, the memory 404 may include any number and any combination of random-access memory (RAM), read-only memory (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, compact flash (CF) cards, etc. Memory 404 may be a group of memories grouped as a single logical component (not shown in Figure 4).
[0090] Bus 410 may be a communication device that transfers data between components in device 400, such as an internal bus (e.g., CPU memory bus) or an external bus (e.g., a universal serial bus port, a peripheral component interconnection express port).
[0091] To facilitate explanation without creating ambiguity, the processor 402 and other data processing circuits are collectively referred to as “data processing circuits” in this disclosure. The data processing circuits may be implemented entirely in hardware, or as a combination of software, hardware, or firmware. Furthermore, the data processing circuits may be a single, independent module, or may be combined whole or partially with any other component of the device 400.
[0092] The device 400 may further include a network interface 406 for providing wired or wireless communication to a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.). In some embodiments, the network interface 406 may include any number and any combination of a network interface controller (NIC), radio frequency (RF) module, transponder, transceiver, modem, router, gateway, wired network adapter, wireless network adapter, Bluetooth adapter, infrared adapter, near-field communication ("NFC") adapter, cellular network chip, etc.
[0093] In some embodiments, the apparatus 400 may further include a peripheral interface 408 for selectively providing connectivity to one or more peripheral devices. As shown in Figure 4, peripheral devices may include, but are not limited to, cursor control devices (e.g., mouse, touchpad, or touchscreen), keyboards, displays (e.g., cathode ray tube displays, liquid crystal displays, or light-emitting diode displays), video input devices (e.g., cameras, or input interfaces connected to video archives), and the like.
[0094] It should be noted that a video codec (for example, a codec that runs processes 200A, 200B, 300A, or 300B) may be implemented as any combination of any software or hardware modules in device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more software modules of device 400, such as program instructions that can be loaded into memory 404. As another example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of device 400, such as dedicated data processing circuits (e.g., FPGA, ASIC, NPU, etc.).
[0095] The present invention provides a method for signaling Supplemental Extension Information (SEI) messages used in video generation and compression.
[0096] First, exemplary bitstream structures used in the art will be described. Similar to AVC and HEVC, a bitstream in VVC includes one or more coded video sequences (CVS). CVS are coded independently of other CVS. Figure 5 shows exemplary structures of CVS 500 according to some embodiments of the present disclosure. Each CVS 500 includes one or more layers (e.g., 520a, 520b, ... 520m), each layer being a representation of video having a specific quality or spatial resolution, or a representation of a component interpretation property, such as a depth or transparency map, or a perspective view. In another dimension, each CVS consists of one or more access units (AUs) (e.g., 530a, 530b, ... 530n), each AU consisting of one or more picture units (PUs) of different layers. For example, AU 530a includes PUs 531a, 531b, ... 531m, each corresponding to layers 520a, 520b, ... 520m. A coded layer video sequence (CLVS) (e.g., 510a, 510b…510n) is a layer-by-layer CVS consisting of a series of PUs in the same layer. If the bitstream has multiple layers, the CVS in the bitstream includes multiple CLVSs (e.g., 510a, 510b…510n). Otherwise, the CVS is identical to the CLVS. As shown in Figure 5, an AU (e.g., 530a, 530b, or 530n) is represented by a group of vertical PUs belonging to the same time, and a CLVS510 is represented by a group of horizontal PUs belonging to the same layer spanning several AUs (e.g., 530a, 530b,…530n). Figure 6 shows the structure of an exemplary picture unit (PU) 600 according to some embodiments of the present disclosure. A simplified structure of PU600 is shown in Figure 6. PU600 includes one coded picture 610, each coded picture 610 including one or more coded slice network abstraction layer (NAL) units 611 (also known as VCL NAL units). In addition to coded slice NAL units 611, PU600 may include non-VCL NAL units 621 such as parameter sets and SEI NAL units.NAL units (including VCL NAL units 611 and non-VCL NAL units 621) are typically ordered within PU600, such as Decoding Capability Information (DCI), Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Prefix Adaptive Parameter Set (APS), Picture Header (PU), Prefix Supplemental Extension Information (SEI), VCL, Suffix SEI, Suffix APS, and Sequence End Symbol (EOS) (some NAL units are selectable). Additionally, Access Unit Delimiter (AUD) NAL units and Operation Point Information (OPI) NAL units are permitted only at the beginning of the AU, and Bitstream End (EOB) NAL units are permitted only at the end of the AU.
[0097] VVC NAL units, like HEVC, include a 2-byte NAL unit header and NAL unit payload. However, the syntax of the NAL unit header differs slightly from that of HEVC. The NAL unit header in both HEVC and VVC includes a layer identifier (ID), NAL unit type, and temporal ID, which are important for both the decoding process and system usage. For example, some systems may use the information carried in the NAL unit header to perform actions such as accessing the bitstream from a specific point in time or stream adaptation through (temporal) layer pruning during network transmission.
[0098] A NAL unit includes a NAL unit header and a payload. Figure 7 shows Table 1 illustrating an exemplary syntax structure 700 of a NAL unit according to some embodiments of this disclosure. Figure 8 shows Table 2 illustrating an exemplary syntax structure 800 of rbsp_trailing_bits() according to some embodiments of this disclosure.
[0099] The semantics related to syntax structures 700 and 800 are explained below.
[0100] NumBytesInNalUnit specifies the size of the NAL unit in bytes. This value is used to decode the NAL unit. Some form of demarcation of the NAL unit boundary enables the inference of NumBytesInNalUnit.
[0101] The Video Coding Layer (VCL) is specified to efficiently represent the content of video data. The NAL is specified to format the data and provide header information in a manner suitable for transmission over various communication channels or storage media. All data is contained within NAL units, each containing an integer byte. NAL units specify a general format for use in both packet-oriented and bitstream systems. In bytestream formats, the format of NAL units is identical for both packet-oriented transport and bytestreams, except that each NAL unit is preceded by a start code prefix and additional padding bytes.
[0102] rbsp_byte[i] is the i-th byte of the raw byte sequence payload (RBSP). An RBSP is specified as a sequence of bytes ordered as follows:
[0103] RBSP contains a string of data bits (SODB) as follows: - If SODB is empty (i.e., its length is 0 bits), then RBSP is also empty. - Otherwise, RBS includes SODB as follows: 1) The first byte of the RBSP contains the first (most significant, leftmost) 8 bits of the SODB, the next byte of the RBSP contains the next 8 bits of the SODB, and so on, until the remaining bits of the SODB are less than 8. 2) The rbsp_trailing_bits() syntax structure is located after SODB, as follows: i) The first (most significant, leftmost) bit of the last RBSP byte contains the remaining bits of the SODB (if any). ii) The next bit consists of a single bit equal to 1 (i.e., rbsp_stop_one_bit). iii) If rbsp_stop_one_bit is not the last bit of the byte-aligned byte, then one or more zero-value bits (i.e., instances of rbsp_alignment_zero_bit) exist to achieve byte alignment. 3) In some RBSPs, there may be one or more 16-bit syntax elements of rbsp_cabac_zero_word equal to 0x0000 after rbsp_trailing_bits() at the end of the RBSP.
[0104] Syntax structures having these RBSP properties are indicated in the syntax table by the "_rbsp" suffix. These structures are carried within the NAL unit as the content of the rbsp_byte[i] data byte. Figure 9 shows Table 3 illustrating exemplary associations between RBSP syntax structures and NAL units in some embodiments of this disclosure.
[0105] If the boundaries of the RBSP are known, the decoder can extract the SODB from the RBSP by concatenating the bits of the RBSP bytes, discarding the last (least significant, rightmost) bit, rbsp_stop_one_bit, which is equal to 1, and discarding the subsequent (lower significant, rightmost) bits that are equal to 0. The data required for the decoding process is contained in the SODB portion of the RBSP.
[0106] emulation_prevention_three_byte is a byte equal to 0x03. If emulation_prevention_three_byte exists in the NAL unit, it will be discarded by the decryption process.
[0107] The last byte of the NAL unit is not equal to 0x00.
[0108] Within a NAL unit, the following 3-byte sequence does not appear at any byte alignment position. - 0x000000; - 0x000001; - 0x000002.
[0109] Within the NAL unit, no 4-byte sequence starting with 0x000003 other than the following sequence will appear at any byte alignment position. - 0x00000300; - 0x00000301; - 0x00000302; - 0x00000303.
[0110] Figure 10 shows Table 4 illustrating an exemplary syntax structure of a NAL unit header 1000 according to some embodiments of this disclosure. The relevant semantics are described below.
[0111] forbidden_zero_bit is equal to 0.
[0112] nuh_reserved_zero_bit is equal to 0. A value of 1 for nuh_reserved_zero_bit may be specified in the future by ITU-T|ISO / IEC. The VVC specification requires that the value of nuh_reserved_zero_bit be equal to 0, but a VVC-compliant decoder will allow a value of 1 for nuh_reserved_zero_bit to appear in the syntax and will ignore (i.e., remove and discard) NAL units with a nuh_reserved_zero_bit equal to 1.
[0113] nuh_layer_id specifies the identifier of the layer to which a VCL NAL unit belongs, or the identifier of the layer to which a non-VCL NAL unit applies. The value of nuh_layer_id is in the range of 0 to 55. Other values for nuh_layer_id are reserved for future use by ITU-T|ISO / IEC. The VVC specification requires that the value of nuh_layer_id be in the range of 0 to 55, but VVC-compliant decoders allow nuh_layer_id values greater than 55 to appear in the syntax and ignore (i.e., remove and discard) NAL units with nuh_layer_ids greater than 55.
[0114] The value of nuh_layer_id is the same for all VCL NAL units in an encoded picture. The value of nuh_layer_id for an encoded picture or PU is the value of nuh_layer_id for the VCL NAL unit of that encoded picture or PU.
[0115] If nal_unit_type is equal to PH_NUT or FD_NUT, then nuh_layer_id is equal to the nuh_layer_id of the associated VCL NAL unit.
[0116] If nal_unit_type is equal to EOS_NUT, then nuh_layer_id is equal to one of the nuh_layer_id values of the layer that exists in CVS.
[0117] The value of nuh_layer_id for DCI, OPI, VPS, AUD, and EOB NAL units is not restricted.
[0118] nal_unit_type specifies the NAL unit type, that is, the type of RBSP data structure included in the NAL unit specified in Table 3 shown in Figure 9.
[0119] NAL units with nal_unit_type in the range of UNSPEC_28…UNSPEC_31 for which no semantics are specified do not affect the decoding process specified in the VVC specification.
[0120] NAL unit types in the range UNSPEC_28…UNSPEC_31 can be used as determined by the application. This specification does not specify a decoding process for these values of nal_unit_type. Because different applications may use these NAL unit types for different purposes, particular care must be taken in designing encoders that produce NAL units with these nal_unit_type values, and decoders that interpret the contents of NAL units with these nal_unit_type values. The VVC specification also does not define any control over these values. These nal_unit_type values may only be suitable for use in contexts where usage "conflicts" (i.e., different definitions of the meaning of NAL unit contents for the same nal_unit_type value) are not important, impossible, or are controlled (e.g., defined or controlled by a control application or transport specification, or defined or controlled by controlling the environment in which the bitstream is distributed).
[0121] Aside from determining the amount of data in the bitstream on the PU, the decoder ignores (removes and discards) the contents of all NAL units that use the pending nal_unit_type value.
[0122] nuh_temporal_id_plus1 minus 1 specifies the time identifier for the NAL unit.
[0123] The value of nuh_temporal_id_plus1 is not 0.
[0124] Next, we will discuss Supplemental Extension Information (SEI) messages. SEI messages are intended to be transmitted within the coded video bitstream in the manner specified by the video coding specification, or by other means determined by the specifications of the system using such coded video bitstream. SEI messages can include various types of data, such as indicating the timing of video pictures, or describing various properties of the coded video and how to use or enhance them. SEI messages are also defined to include arbitrary user-defined data. SEI messages do not affect the core decoding process, but they can indicate recommendations on how to post-process or display the video.
[0125] SEI assists processes related to decoding, display, or other purposes. Similar to Video Usability Information (VUI), SEI does not affect signal processing operations in the decoding process. SEI syntax for various purposes is transmitted in a syntax structure called an SEI message, with one or more SEI messages carried within a NAL unit called a SEINAL unit. SEI messages can have a very high level of operational range, similar to VUI, or a narrower range, such as being applied to individual pictures or slices. As the name suggests, SEI messages are intended to supplement video content. Decoder support for SEI messages is usually selective, and even when a decoder uses an SEI message, in most cases the decoder is not required to use the SEI message exactly as described in the standard. However, SEI messages do affect bitstream conformance (for example, if the syntax of an SEI message in a bitstream does not conform to the specification, the bitstream is not conforming to the specification). Also, some SEI messages are used to specify HRD operations and HRD-based bitstream conformance requirements.
[0126] To specify SEI messages, the JVET Working Group also developed the H.274 standard, which specifies the syntax and semantics of Video Usability Information (VUI) parameters and Supplemental Extension Information (SEI) messages, intended for use with coded video bitstreams specifically specified in the VVC standard. However, since VUI parameters and SEI messages do not affect the decoding process, SEI messages in H.274 can also be used with other types of coded video bitstreams such as H.265 / HEVC and H.264 / AVC.
[0127] Next, we will discuss face video generation and compression. With the advent of deep generative models, including variational autoencoding (VAE) and generative adversarial networks (GAN), face video compression has achieved promising performance improvements. In 2018, Wiles designed X2Face, which controls face generation via image, audio, and pose codes. Also, Zakharov et al. published a realistic neural talking head model through adversarial learning with a small number of shots. For video-to-video synthesis tasks, an NVIDIA research team first proposed Face-vidtovid in 2019. Subsequently, in 2020, they proposed a new scheme that leverages compact 3D keypoint representations to drive generative models and render target frames. In addition, a Facebook research team designed a mobile-compatible video chat system based on FOMM. Feng et al. proposed VSBNet, which reconstructs the original frame from landmarks using adversarial learning. Furthermore, Chen et al. propose an end-to-end talking head video compression framework based on compact feature learning (CFTE), which is cleverly designed for highly efficient talking face video compression in ultra-low bandwidth scenarios. The CFTE scheme leverages compact feature representations to compensate for temporal evolution and reconstruct the target face video frame end-to-end. This can be incorporated into a video coding framework that has rate distortion target monitoring.
[0128] Figure 11 is a schematic diagram showing an exemplary deep learning-based video generation and compression framework 1100 according to some embodiments of the present disclosure. For example, the framework 1100 may be based on a first-order motion model (FOMM). The FOMM deforms a reference source frame to follow the motion of the driven video. This method can be applied to various types of video (e.g., video and cartoons), but can also be used for face animation applications. The FOMM follows an encoder-decoder architecture having a motion transfer component, which includes the following steps:
[0129] First, a keypoint extractor (also called a motion module) is trained using constant loss without explicit labels. This keypoint extractor computes two sets of 10 trained keypoints for the source frame and the driving frame. The trained keypoints are transformed from a feature map of size channel × 64 × 64 using a Gaussian map function, so that each corresponding keypoint can represent feature information from a different channel. Note that each keypoint is a (x, y) point that can represent the most important information in the feature map.
[0130] Next, the high-density motion network uses landmarks and source frames to generate a high-density motion field and an occlusion map.
[0131] Next, the encoder 1110 encodes the source frame using conventional image / video compression methods such as HEVC / VVC or JPEG / BPG. Here, VVC is used to compress the source frame.
[0132] In a later stage, the resulting feature map is warped using a high-density motion field (using differentiable grid sampling operations) and then multiplied with the occlusion map.
[0133] Finally, decoder 1120 generates an image from the warp map.
[0134] Figure 12 shows an exemplary deep learning-based video generation compression framework 1200 according to some embodiments of the present disclosure. Figure 12 provides another basic framework of a deep-based video generation compression scheme based on compact feature representation, i.e., CFTE. This follows an encoder / decoder architecture that applies a context-based coding scheme.
[0135] On the encoder 1210 side, the compression framework consists of three modules: an encoder (also called a VVC coding module) for compressing keyframes, a feature extractor for extracting compact human features from other interframes, and a feature coding module for compressing interpredictive residuals of compact human features. First, keyframes representing human textures are compressed by the VVC encoder. Through the compact feature extractor, each subsequent interframe is represented as a compact feature matrix with a size of 1 × 4 × 4. Note that the size of the compact feature matrix is not fixed, and the number of feature parameters can be increased or decreased depending on the specific requirements of bit consumption. Next, these extracted features are interpredicted and quantized, and the residuals are finally entropy coded to form the final bitstream.
[0136] On the decoder 1220 side, this compression framework also includes three main modules, including decoding to reconstruct keyframes, reconstruction of compact features by entropy decoding and correction, and generation of the final image using the reconstructed features and decoded keyframes. More specifically, during the generation of the final image, the decoded keyframes from the VVC bitstream are further represented in the form of features by compact feature extraction. Next, given features from the keyframes and interframes, the associated sparse motion field is computed, facilitating the generation of high-density motion and occlusion maps per pixel. Finally, based on a deep generative model, the decoded keyframes, high-density motion maps per pixel, and occlusion maps with implicit motion field properties are used to generate a final image with accurate appearance, pose, and facial expressions.
[0137] Numerous studies have focused on 3D faces to further pursue coding performance. 3D head models have been employed, and only the pose parameters of a face-specific video compression task have been encoded. Subsequently, both eigenspace and principal component analysis (PCA) models have been used for this task. However, based on these conventional 3D techniques, the visual quality of the reconstructed images is unacceptable. The development of deep generative models may yield promising results for this 3D MMM-assisted face video generation task.
[0138] Figure 13 is a schematic diagram showing a general encoder-decoder generation and compression framework 1300 for 3DMM auxiliary conversational face video according to some embodiments of the present disclosure. Generally, the generation of 3DMM auxiliary face video can provide accurate 3D face reconstruction based on a combination of shape and S texture T given by the following formula.
number
number
number
[0139] There are several problems with signaling SEI messages. For example, the first problem stems from transmitting all parameters for all frames in a single SEI message. Specifically, to support the generation compression in current HEVC or VVC standards, it has been proposed to encode a first frame, called a keyframe, as an intrapicture using the HEVC or VVC standard. All existing coding tools and encoder optimization methods in the standard can be applied. Furthermore, for subsequent frames, instead of directly encoding them using conventional encoders, the features used for frame generation are abstracted and signaled in the bitstream. The generated video may include, but is not limited to, images of faces or human bodies. Therefore, SEI messages were proposed. Features are signaled in the proposed SEI message. After decoding the keyframe, the decoder also decodes the SEI message to obtain the features and generates subsequent frames based on the first keyframe and features. Since there are multiple subsequent frames generated, they may all have different features. In this disclosure, the features required to generate a single frame are referred to as the feature parameter set. Therefore, the number of feature parameter sets being signaled is the same as the number of frames in the generated video sequence.
[0140] One way to signal multiple feature parameter sets is to send all the necessary parameters in a single SEI message. Figure 14 shows an exemplary SEI message 1400 having multiple feature parameter sets according to some embodiments of the Disclosure. As shown in Figure 14, a single SEI message 1400 contains multiple feature parameter sets 1401. The number of feature parameter sets in an SEI message is the same as the number of frames to be generated. Figure 15 shows a syntax signaling process 1500 for a single SEI message having multiple feature parameter sets according to some embodiments of the Disclosure. Referring to Figure 15, step 1510 signals an SEI ID indicating that the current SEI message is a feature SEI message for frame generation. Step 1520 signals multiple feature parameter sets. Step 1530 signals feature parameters for each feature parameter set that is being signaled. Step 1540 determines whether the most recently signaled feature parameter set is the last feature parameter set. If it is the last feature parameter set, signaling ends. If it is not the last feature parameter set, the signaling process returns to step 1530 and continues signaling feature parameters.
[0141] In current existing signaling processes, the encoder determines the number of frames to generate before sending the entire SEI message, and the decoder decodes all feature parameters before generating the first frame, which results in considerable delay, especially when generating a large number of frames. In some real-time and low-latency applications, the encoder sends feature parameters for each frame in real time, and the decoder also generates the frames in real time. Therefore, sending all parameters for all frames in a single generated SEI message is not suitable for real-time and low-latency applications.
[0142] The second problem concerns ensuring that the encoder and decoder use matching feature analysis models. Specifically, in generative compression, non-key picture features are extracted by the feature analysis model and encoded into a bitstream. The decoder then decodes the features from the bitstream, and the generative model uses those features as input to generate the picture. Therefore, the generative model on the decoder side must match the analysis model on the encoder side; otherwise, the picture cannot be generated correctly. In the current design, only features are transmitted in the SEI message, and no information about the model is included. Therefore, the decoder does not know which model the encoder used or whether it matches the generative model.
[0143] The present invention provides an SEI signaling method for solving one or more of the above problems.
[0144] In some embodiments, a method is provided for signaling a set of feature parameters for a single frame in a single SEI message in order to solve the first problem described above.
[0145] Figure 16 shows a flowchart illustrating an exemplary method 1600 for generating one or more frames according to some embodiments of the present disclosure. Method 1600 may be performed by a decoder (e.g., process 300A in Figure 3A or process 300B in Figure 3B) or by one or more software or hardware components of an apparatus (e.g., apparatus 400 in Figure 4). For example, a processor (e.g., processor 402 in Figure 4) can perform method 1600. In some embodiments, method 1600 may be implemented by a computer program product embodied in a computer-readable medium, which includes computer-executable instructions such as program code executed by a computer (e.g., apparatus 400 in Figure 4). Referring to Figure 16, method 1600 may include the following steps 1602 and 1604.
[0146] In step 1602, the picture units are decoded. A picture unit contains one or more SEI (Supplementary Extension Information) messages. For example, a single picture unit (PU) contains multiple feature SEI messages, each feature SEI message used to generate a single frame. Figure 17 shows a first set 1700 of exemplary picture unit structures according to some embodiments of the present disclosure. As shown in Figure 17, for each picture unit (e.g., picture unit 0, picture unit 1, ... picture unit N), first the picture header NAL unit 1701 is decoded. Next, one or more coded slice NAL units 1702 for keyframes (e.g., keyframe 0, keyframe 1, ... keyframe N) are decoded. In some embodiments, if there is only one coded slice NAL unit 1702, the picture header NAL unit 1701 is selective; that is, in this case, there may be no picture header NAL unit 1701 to decode. According to the HEVC and VVC specifications, one picture unit contains one coded picture, so only one keyframe (e.g., keyframe 0, keyframe 1, ..., or keyframe N) is coded and included in the picture unit. Following the coded slice NAL unit 1702, there is a feature SEI message NAL unit 1703. A picture unit may have one or more feature SEI message NAL units 1703. Each feature SEI message NAL unit 1703 contains one feature SEI message used to generate one frame based on the keyframes in the same picture unit. Therefore, the number of frames generated is the same as the number of feature SEI message NAL units 1703 in the same picture unit. If a second keyframe (e.g., keyframe 1) is coded and signaled in the bitstream, then a second picture unit, e.g., picture unit 1, will exist.The second picture unit (e.g., picture unit 1) also includes a picture header NAL unit 1701 (if present) and one or more coded slice NAL units 1702 for the second keyframe (e.g., keyframe 1). After these coded slice NAL units 1702 are decoded, one or more feature SEI messages 1703 are decoded. These feature SEI message NAL units 1703 contain feature parameters used to generate a frame based on the second keyframe (e.g., keyframe 1). Thus, in this example, the number of picture units in the bitstream is the same as the number of coded keyframes.
[0147] In step 1604, one or more frames are generated based on a keyframe and one or more SEI messages. For example, referring to Figure 17, taking picture unit 0 as an example, the first frame is generated based on keyframe 0 and the first SEI message in feature SEI message NAL unit 1703a, the second frame is generated based on keyframe 0 and the second SEI message in feature SEI message NAL unit 1703b, and so on. Thus, in this example, the feature SEI messages used to generate the frames are located within the same picture unit as the keyframes on which the frames were generated. In some embodiments, the keyframes and SEI messages may be contained in different picture units, which will be discussed later in this disclosure.
[0148] From the perspective of video coding standards, only keyframes are coded pictures and should be output. The generation of frames containing feature SEI messages is a post-processing step, and the generated frames are not the decoder's output. Therefore, the post-processor should properly manage the frames decoded by the decoder and the frames generated by the post-processor. If only the first frame is coded as a keyframe, there is only one picture unit in the entire CLVS.
[0149] Because multiple frames may be generated using multiple feature SEI messages contained within a single PU, a picture order count pic_order_cnt is signaled in the feature SEI message to distinguish the display order of these generated frames. Figure 18 shows Table 5 illustrating exemplary feature SEI message syntax 1800 in some embodiments of the present disclosure. Referring to Figure 18, the picture order count pic_order_cnt 1801 is signaled in the feature SEI message. The picture order count pic_order_cnt specifies the display order count of the modulo 1 << 31 of the pictures generated in the current SEI message. Thus, even if the SEI messages are transmitted in a random order, the generated frames will be displayed in the correct order specified by pic_order_cnt. This provides the transmission system with flexibility in transmitting feature SEI messages. In some embodiments, the picture order count may be coded in a fixed-length code or a variable-length code. In some embodiments, the picture sequence count is divided into the most significant bit (MSB) and the least significant bit (LSB), and to save signaling cost, only the LSB is signaled in the bitstream. The MSB may be derived in response to the change in the LSB of the picture sequence count signaled in two consecutive SEI messages. In this example, on the decoder side, the decoder is configured to first decode the LSB and then derive the MSB in response to the change between the first LSB of the picture sequence count decoded in the previous SEI message and the second LSB of the current SEI message.
[0150] Since one picture unit contains only one feature SEI message that can be used to generate one frame, the number of frames is the same as the number of picture units. Figure 19 shows a second set 1900 of exemplary picture unit structures according to some embodiments of the present disclosure. As shown in Figure 19, in this example, the first frame is actually coded as a keyframe, so the first picture unit, i.e., picture unit 0, contains a coded slice NAL unit 1902 for the keyframe 1910. Also, the first picture unit, i.e., picture unit 0, does not have a feature SEI message NAL unit. In subsequent picture units (e.g., picture unit 1, ... picture unit N), each picture unit contains a feature SEI message NAL unit 1903, and the feature SEI message in the feature SEI message NAL unit 1903 is used to generate a frame based on the keyframe 1910 in the first picture unit, i.e., picture unit 0.
[0151] In HEVC and VVC, each picture unit is specified to contain one coded picture, so the resulting frame is not a coded picture, but a dummy picture 1920 coded for the next picture unit (e.g., picture unit 1…picture unit N). To save bit overhead transmitted in the bitstream, in some embodiments, all samples of the dummy picture can be set to the same value, e.g., 1<<(bitdpeth-1), where bitdepth is the bit depth of the picture sample. In some embodiments, the encoder can skip as many slice data level syntax elements as possible and maintain syntax elements signaled by the same picture across different blocks by directly signaling the slice data with no CU splitting, first most likely mode (MPM), and zero residuals. Thus, the decoder can decode the dummy picture with first most likely mode (MPM), zero residuals, and slice data without coded unit splitting, and there are no slice data level syntax elements to decode.
[0152] In some embodiments, when context-adaptive binary arithmetic coding (CABAC) is used, bit overhead can be very small if the encoder always signals the same value for the same syntax element. Since the picture header NAL unit (or picture header structure in the slice header) and slice header are required for each coded picture, the encoder can reduce the bit overhead of the slice header by using a single slice for the dummy picture. The encoder can then disable all intercoding tools in the sequence parameter set (SPS) and code the dummy picture as an intra-slice, skipping all inter-slice or inter-prediction-related control flags, configurations, or parameters in the picture header and slice header. It is also possible to disable these intercoding tools where parameters are signaled at the picture or slice level, thereby skipping the signaling of the relevant parameters. For example, Adaptive Loop Filter (ALF), Cross-Component Adaptive Loop Filter (CCALF), Luminance Mapping with Chroma Scaling (LMCS), Scaling List, Virtual Boundary, CU Delta QP, the presence of Deblocking Parameters, Bidirectional Optical Flow (BDOF), Improved Prediction with Optical Flow (PROF), and Improved Decoder-Side Motion Vector (DMVR) can all be disabled. With the syntax elements that must be signaled being signaled, the slice header and picture header occupy approximately 2-3 bytes according to the VVC syntax structure. The bit cost of the dummy picture is relatively small compared to the bit cost of the feature SEI message itself.
[0153] In this example, referring to Figure 19, the decoder decodes the dummy picture 1920 and generates a frame with the keyframe 1910 in picture unit 0 based on the SEI message of the feature SEI message NAL unit 1903 in picture unit 1. The dummy picture is then replaced with the generated frame and discarded. In some examples, the dummy picture is not used, so the decoder may skip decoding the dummy picture and generate only the frame to be used.
[0154] In the method described above, the dummy picture is coded and transmitted in a bitstream, but discarded on the decoder side. Therefore, a certain number of bits are required to signal the dummy picture. In some embodiments, the encoder transmits an NAL unit with a special NAL unit type to represent the coded picture, but does not transmit any payload to store the number of bits.
[0155] Figure 20 shows a third set 2000 of exemplary picture unit structures according to some embodiments of the present disclosure. As shown in Figure 20, one picture unit contains one coded picture. For the first picture unit, i.e., picture unit 0, it contains the first frame of a video sequence coded as a keyframe 2010. The first picture unit, i.e., picture unit 0, does not contain a feature SEI message NAL unit. In subsequent picture units (e.g., picture unit 1…picture unit N), each picture unit contains a coded picture and a feature SEI message NAL unit 2003. The coded picture is represented by an empty picture NAL unit 2004, and the feature SEI message NAL unit 2003 contains a feature SEI message used for frame generation. The empty picture NAL unit 2004 represents an empty picture 2020.
[0156] On the decoder side, upon receiving an empty picture NAL unit 2004, the decoder recognizes it as a new picture unit and can directly discard the empty picture NAL unit 2004. No decoding process is required for the current picture. After decoding the feature SEI message, the decoder generates a frame containing the signaled feature parameters based on keyframe 2010 and outputs the generated frame.
[0157] An unspecified NAL unit type in HEVC or VVC can be used for an empty picture NAL unit 2004. For example, in VVC, NAL units with nal_unit_type values from 28 to 31 are not specified. Therefore, an application can use nal_unit_type 28 to 31 as the NAL unit type for an empty picture NAL unit 2004. The size of an empty picture NAL unit (NumBytesInNalUnit) is set to 2 bytes, the same size as the NAL unit header. Therefore, an empty picture NAL unit does not have an RBSP.
[0158] In this example, the encoder sends an empty NAL unit 2004 with a nal_unit_type of 28 to 31 and containing only a NAL unit header as the coded picture. The empty NAL unit 2004 is followed by a feature SEI message NAL unit 2003, which contains the feature SEI message used to generate the frame. The decoder decodes the NAL units with nal_unit_type 28 to 31 and discards the empty NAL unit 2004. Next, the decoder decodes the feature SEI message NAL unit 2003 and determines the feature parameters. Based on the already decoded keyframe 2010, the decoder generates a frame with the feature parameters.
[0159] Since a picture unit must contain a coded picture that includes one or more VCL NAL units, in the above embodiment, an empty picture NAL unit is transmitted to save bits that would otherwise be allocated to the coded picture. In other embodiments, a picture unit may have zero or one coded picture; that is, a coded picture is not required for a picture unit. In that case, to determine the first NAL unit of a new picture unit, this example proposes a picture unit delimiter (PUD), which is a new NAL unit. The PUD is used to delimit picture units, and this is optional. A picture can contain at most one PUD NAL unit. If a picture unit does not have any VCL NAL units, then a PUD must be present in the picture unit. If a PUD unit is present in the picture unit, then the PUD unit should be the first NAL unit of the picture unit.
[0160] By introducing PUD, it becomes possible to skip empty coded picture NAL units. The encoder sends a PUD to start a picture unit, and then sends a feature SEI message NAL unit in the picture unit for frame generation. When the decoder decodes the PUD, it determines the new picture unit to be decoded. A feature SEI message NAL unit before the PUD in the bitstream is determined to be in the previous picture unit, and a feature SEI message NAL unit after the PUD is determined to be in the next picture unit. Using PUD, the decoder can clearly determine the picture unit and know which picture unit the decoded feature SEI message belongs to. As a result, the frame generated by the feature SEI message is output for the current picture unit.
[0161] Figure 21 shows a fourth set 2100 of exemplary picture unit structures according to some embodiments of the present disclosure. For a first picture unit, i.e., picture unit 0, it contains the first frame of a video sequence coded as a keyframe 2110. The first picture unit, i.e., picture unit 0, does not contain a feature SEI message NAL unit. For subsequent picture units (e.g., picture unit 1, ... picture unit N), each picture unit contains a PUD 2104 and a feature SEI message NAL unit 2103. These picture units (e.g., picture unit 1, ... picture unit N) do not contain coded pictures, and the picture unit can be determined by the PUD 2104. The feature SEI message NAL unit 2103 contains a feature SEI message that can be used to generate a frame, and the generated frame is treated as a decoded picture in the picture unit (e.g., picture unit 1, ... picture unit N) containing the feature SEI message NAL unit 2103.
[0162] Figure 22 shows Table 6 illustrating exemplary syntax of PUD RBSP2200 according to some embodiments of this disclosure. The semantics will be discussed later. Figure 23 shows Table 7 illustrating exemplary interpretation of pud_pic_type2300 according to some embodiments of this disclosure.
[0163] A value of 1 for pud_irap_or_gdr_flag indicates that the picture unit containing the PUD is an intra-random access point (IRAP) or a progressively decoded refresh (GDR) picture unit. A value of 0 for pud_irap_or_gdr_flag indicates that the picture unit containing the PUD is neither an IRAP nor a GDR picture unit.
[0164] The pud_pic_type value indicates that, for a given pud_pic_type value, the sh_slice_type values of all slices of the coded picture in a picture unit containing a PUD NAL unit are members of the set listed in Table 7 shown in Figure 23. The value of pud_pic_type is equal to 0, 1, or 2. Other values for pud_pic_type are reserved for future use. The decoder ignores the reserved values of pud_pic_type.
[0165] In some embodiments, various mechanisms can be used to inform the decoder about the feature analysis model used by the encoder in order to solve the second problem described above (i.e., ensuring that the encoder and decoder extract features of non-key pictures using a matching feature analysis model). In some embodiments, it is proposed to send a key in the SEI message to indicate the analysis model used to extract the features signaled in the SEI message. Figure 24 shows Table 8, which illustrates exemplary feature SEI message syntax 2400 according to some embodiments of the present disclosure. As shown in Figure 24, key 2401 is signaled in the SEI message. On the decoder side, the key is used to identify the analysis network used to generate the syntax elements of the current feature SEI message. The key may also be used to determine whether the analysis model in the encoder matches the generation model stored in the decoder. The generation model in the decoder can generate an image picture using the SEI message only if it recognizes the key. The value of the key may be specified by the application that requires the encoder and decoder to use a matching model. In such applications, if parameters in an SEI message are extracted by the encoder using an analytical model that does not match the generative model in the decoder, the generative model may ignore the received SEI message.
[0166] In some embodiments, for some applications other than generational compression, the generational model (decoder side) and the analysis model (encoder side) do not need to match each other. For example, in use cases such as user-specified animation or filtering, the decoder only needs to perform face animation or face filtering. In these cases, since there is no need to generate a picture, the decoder does not need to have a generational model that matches the encoder's analysis model. Instead, the decoder can use any model to modify or filter faces in the picture or video. Because the encoder has an unencoded face image, it can extract more accurate facial landmarks or keypoints, resulting in higher quality face animation than on the decoder side. This allows the transmitter to have better control over the quality of the animated or filtered face picture. The facial landmarks or keypoints extracted on the encoder side can be signaled with feature SEI messages that can be signaled by the method proposed in this disclosure. That is, instead of using picture generation on the decoder side, feature SEI messages are used to provide landmark or keypoint information to assist in face animation / filtering.
[0167] In some embodiments, a non-temporary, computer-readable storage medium for storing the bitstream is also provided. The bitstream may include coded syntax elements for implementing the disclosed signaling method for Supplementary Enhancement Information (SEI) messages.
[0168] Embodiments may be further described using the following clauses. 1. A method for decoding a bitstream to obtain one or more pictures of a video stream, The steps include receiving a bitstream, The steps include decoding the bitstream to obtain one or more pictures, The decryption step is, A step of decrypting a picture unit containing one or more Supplemental Extended Information (SEI) messages, A method comprising the steps of generating a key picture and one or more pictures based on one or more SEI messages. 2. The method according to Clause 1, wherein the key picture is decoded by a decoder based on the received bitstream. 3. The method according to Clause 1, wherein the key picture and the one or more SEI messages are in the same picture unit. 4. The method according to Clause 3, wherein the SEI message includes a count that identifies the SEI message within the picture unit. 5. The method according to Clause 4, wherein the count indicates the sequence of the one or more pictures corresponding to the SEI message. 6. The method according to Clause 2, wherein the first picture unit includes one SEI message and the second picture unit includes the key picture. 7. The method according to Clause 6, wherein the first picture unit further includes a coded slice network abstraction layer (NAL) unit for coded picture. 8. The steps include: decrypting the encoded picture to obtain a decoded picture; The steps include generating a picture based on the key picture and the SEI message included in the first picture unit, The steps include outputting a generated picture in place of the decoded picture, The method described in Clause 7, further including the method described in Clause 7. 9. The sample of the coded picture is set to a fixed value, as described in Clause 7. 10. The decryption step is: The method according to Clause 7, further comprising the step of decoding the coded picture using the first most likely mode (MPM), zero residuals, and slice data without code unit splitting. 11. The method according to Clause 7, wherein the signaled syntax elements are the same for different blocks in the coded picture. 12. The method according to Clause 7, wherein one slice is used for the coded picture. 13. The method according to Clause 7, wherein all encoding tools in the Sequence Parameter Set (SPS) are disabled for the coded picture. 14. The method of Clause 13, wherein an intercoding tool having parameters signaled at the picture level or slice level is disabled. 15. The method according to Clause 6, wherein the first picture unit includes an empty picture NAL unit indicating an empty picture. 16. The decryption step is: The method according to clause 15, further comprising the step of skipping the decoding process for the empty picture in the empty picture NAL unit. 17. The method according to Clause 15, wherein the empty picture NAL unit does not contain a raw byte sequence payload (RBSP). 18. The empty picture NAL unit is a NAL unit of type as described in Clause 15. 19. The method according to Clause 1, wherein the bitstream includes a picture unit delimiter (PUD) indicating that a subsequent NAL unit belongs to the next picture unit. 20. The decryption step is: A step of decoding a flag in the PUD to indicate whether the subsequent picture unit is an intra-random access point (IRAP) picture unit or a progressively decoded refresh (GDR) picture unit, The method according to clause 19, further comprising the step of decoding an index in the PUD indicating the picture type of the coded picture in the subsequent picture unit. 21. A method for decoding a bitstream to obtain one or more pictures of a video stream, The steps include receiving a bitstream, The steps include decoding the bitstream to obtain one or more pictures, The decryption step is, The steps include: decrypting the Supplementary Extended Information (SEI) message, The steps include determining whether the generative model in the decoder matches the analytical model in the encoder based on the decoded SEI message, A method comprising the steps of generating a picture based on the SEI message using the generation model in response to the generation model matching the analysis model. 22. The step of determining whether the generative model matches the analytical model is: The steps include: decoding the key signaled by the bitstream; The method according to Clause 21, further comprising the step of determining whether the generative model matches the analytical model based on the value of the key. 23. A method for decoding a bitstream to obtain one or more pictures of a video stream, The steps include receiving a bitstream, The steps include decoding the bitstream to obtain one or more pictures, The decryption step is, The steps include: decoding a Supplementary Extended Information (SEI) message containing a landmark or keypoint; A method comprising the step of animating or filtering one or more pictures using the landmark or keypoint in the decoded SEI message. 24. A method for encoding a video sequence into a bitstream, The steps include receiving a video sequence and The steps include encoding one or more pictures of the aforementioned video sequence, The step includes generating a bitstream, The above step of encoding is The steps include encoding one or more supplemental extension information (SEI) messages corresponding to one or more pictures, A method comprising the step of encoding one or more SEI messages into picture units. 25. The above step of encoding is: The method according to clause 24, further comprising the step of encoding a key picture into the picture unit. 26. The above step of encoding is: The method according to clause 24, further comprising the step of encoding a count that identifies the SEI message within the picture unit into the SEI message. 27. The method according to clause 26, wherein the count indicates the sequence of the one or more pictures corresponding to the SEI message. 28. The above step of encoding is: The steps include encoding one SEI message into the first picture unit, The method according to Clause 24, further comprising the step of encoding a key picture into a second picture unit. 29. The above step of encoding is: The method according to clause 28, further comprising the step of encoding a slice network abstraction layer (NAL) unit for a picture into the first picture unit. 30. The sample of the picture is set to a fixed value, as described in Clause 29. 31. The above step of encoding is: The method according to Clause 29, further comprising the step of encoding the picture with first most likely mode (MPM), zero residuals, and slice data without code unit splitting. 32. The method according to Clause 29, wherein the signaled syntax elements are the same for different blocks in the picture. 33. The method according to Clause 29, wherein one slice is used for the picture. 34. The above step of encoding is: The method according to Clause 29, further comprising the step of disabling all encoding tools in the sequence parameter set (SPS) for the aforementioned picture. 35. The above step of encoding is The method according to clause 34, further comprising the step of disabling an intercoding tool that has picture or slice level parameters. 36. The method according to Clause 28, wherein the first picture unit includes an empty picture NAL unit indicating an empty picture. 37. The above step of encoding is: The method according to clause 36, further comprising the step of skipping the encoding process for the empty picture in the empty picture NAL unit. 38. The method according to Clause 37, wherein the empty picture NAL unit does not contain a raw byte sequence payload (RBSP). 39. The empty picture NAL unit is the method described in Clause 37, where the NAL unit type is specified. 40. The above step of encoding is The method according to clause 24, further comprising the step of encoding a picture unit delimiter (PUD) in the bitstream indicating that a subsequent NAL unit belongs to the next picture unit. 41. The above step of encoding is: The steps include encoding a flag in the PUD to indicate whether the subsequent picture unit is an intra-random access point (IRAP) picture unit or a progressively decoded refresh (GDR) picture unit, The method according to clause 40, further comprising the step of encoding an index in the PUD to indicate the picture type of the coded picture in the subsequent picture unit. 42. The above step of encoding is: The method according to clause 24, further comprising the step of encoding a key indicating an analysis model in the encoder into the bitstream. 43. The method described in Clause 24, wherein the SEI message includes a landmark or key point. 44. A non-temporary computer-readable storage medium for storing a video bitstream, wherein the bitstream is: Includes a picture unit containing one or more Supplemental Extended Information (SEI) messages, The aforementioned one or more SEI messages are used to generate one or more pictures on a non-temporary, computer-readable storage medium. 45. The picture unit further includes a key picture, a non-temporary computer-readable storage medium as described in Clause 44. 46. The picture unit further includes a count in the SEI message that identifies the SEI message in the picture unit, as described in Clause 45. 47. A non-temporary computer-readable storage medium as described in Clause 46, in which the count indicates the sequence of the one or more pictures corresponding to the SEI message. 48. The bitstream is Each contains one or more first picture units, each containing one SEI message, A non-temporary computer-readable storage medium as described in Clause 44, including a second picture unit containing a key picture. 49. The non-temporary computer-readable storage medium described in Clause 48, wherein the first picture unit further comprises a coded slice network abstraction layer (NAL) unit for coded pictures. 50. A non-temporary computer-readable storage medium as described in Clause 49, in which the sample of the coded picture is set to a fixed value. 51. A non-temporary computer-readable storage medium as described in Clause 49, having the coded picture with first most probable mode (MPM), zero residual, and slice data without code unit division. 52. The non-temporary computer-readable storage medium described in Clause 49, wherein the bitstream further includes the same syntax elements for different blocks in the coded picture. 53. A non-temporary computer-readable storage medium as described in Clause 49, wherein one slice is used for the coded picture. 54. A non-temporary computer-readable storage medium as described in Clause 49, wherein all decoding tools in the Sequence Parameter Set (SPS) are disabled for the coded picture. 55. A non-temporary computer-readable storage medium as described in Clause 54, in which an intercoding tool having parameters signaled at the picture level or slice level is disabled. 56. The first picture unit is a non-temporary computer-readable storage medium as described in Clause 48, which includes an empty picture NAL unit indicating an empty picture. 57. The empty picture NAL unit is a non-temporary computer-readable storage medium as described in Clause 56, which does not contain a raw byte sequence payload (RBSP). 58. The empty picture NAL unit is a non-temporary computer-readable storage medium as specified in Clause 56, which is of NAL unit type. 59. The non-temporary computer-readable storage medium described in Clause 44, wherein the bitstream further includes a picture unit delimiter (PUD) in the bitstream indicating that a subsequent NAL unit belongs to the next picture unit. 60. The PUD is, A flag indicating whether the subsequent picture unit is an intra-random access point (IRAP) picture unit or a progressively decoded refresh (GDR) picture unit, A non-temporary computer-readable storage medium according to Clause 59, further comprising an index indicating the picture type of the coded picture in the subsequent picture unit including the PUD. 61. The non-temporary computer-readable storage medium described in Clause 44, wherein the bitstream further includes keys indicating an analysis model in the encoder. 62. The SEI message is contained in a non-temporary computer-readable storage medium as described in Clause 44, including landmarks or key points.
[0169] It should be noted that relational terms such as “first” and “second” in this specification are used solely to distinguish one entity or operation from another, and do not require or imply any actual relationship or order between these entities or operations. Furthermore, the words “comprising,” “having,” “containing,” and “including,” as well as other similar forms, are intended to be open-ended in that they are equivalent in meaning, and that the one or more items following any one of these words do not exhaustively list such one or more items, or are not limited to only the listed one or more items.
[0170] As used herein, unless otherwise specified, the term “or” encompasses all possible combinations, unless impractical. For example, if it is stated that a database may contain A or B, then unless otherwise specified or impractical, the database may contain A, or B, or A and B. As a second example, if it is stated that a database may contain A, B, or C, then unless otherwise specified or impractical, the database may contain A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.
[0171] It will be understood that the embodiments described above may be implemented by hardware, software (program code), or a combination of hardware and software. If implemented by software, it may be stored in the computer-readable medium described above. When executed by a processor, the software may perform the methods disclosed herein. The computing units and other functional units described herein may be implemented by hardware, software, or a combination of hardware and software. It will also be understood by those skilled in the art that several of the above modules / units may be integrated into a single module / unit, and each of the above modules / units may be further divided into several submodules / subunits.
[0172] In the above specification, embodiments have been described with reference to numerous specific details that may vary depending on the embodiment. Specific adaptations and modifications may be made to the above embodiments. Other embodiments may become apparent to those skilled in the art in consideration of the specification and practice of the present invention disclosed herein. This specification and examples are to be considered merely illustrative, and the true scope and spirit of the invention are intended to be shown by the appended claims. Furthermore, the order of steps shown in the figures is for illustrative purposes only and is not intended to limit the steps to any particular order. Therefore, those skilled in the art will understand that these steps may be performed in different orders while carrying out the same method.
[0173] Exemplary embodiments are disclosed in the drawings and specification. However, many variations and modifications can be made to these embodiments. Accordingly, certain terms are used, but only in a general and descriptive sense, and not as limiting.
Claims
1. A method for decrypting a bitstream to obtain one or more pictures of a video stream, The steps include receiving a bitstream, The steps include decoding the bitstream to obtain one or more pictures, The decryption step is, The steps include: decoding a picture unit containing one or more Supplemental Extended Information (SEI) messages; A method comprising the steps of generating a key picture and one or more pictures based on one or more SEI messages.
2. The method according to claim 1, wherein the key picture is decoded by a decoder based on the received bitstream.
3. The method according to claim 1, wherein the key picture and the one or more SEI messages are in the same picture unit.
4. The method according to claim 2, wherein the first picture unit includes one SEI message and the second picture unit includes the key picture.
5. The method according to claim 4, wherein the first picture unit further comprises a coded slice network abstraction layer (NAL) unit for coded pictures.
6. The steps include: decrypting the encoded picture to obtain a decoded picture; The steps include generating a picture based on the key picture and the SEI message contained in the first picture unit, The steps include outputting a generated picture in place of the decoded picture, The method according to claim 5, further comprising:
7. The method according to claim 1, wherein the bitstream includes a picture unit delimiter (PUD) indicating that a subsequent NAL unit belongs to the next picture unit.
8. A step of decoding a flag in the PUD indicating whether the subsequent picture unit is an intra-random access point (IRAP) picture unit or a progressively decoded refresh (GDR) picture unit, The steps include decoding the index in the PUD that indicates the picture type of the coded picture in the subsequent picture unit, The method according to claim 7, further comprising:
9. A method for encoding a video sequence into a bitstream, The steps include receiving a video sequence and The steps include encoding one or more pictures of the aforementioned video sequence, The step includes generating a bitstream, The above step of encoding is The steps include encoding one or more Supplemental Extension Information (SEI) messages corresponding to one or more pictures, A method comprising the step of encoding one or more SEI messages into picture units.
10. The above step of encoding is The method according to claim 9, further comprising the step of encoding a key picture into the picture unit.
11. The above step of encoding is The steps include encoding one SEI message into the first picture unit, The method according to claim 9, further comprising the step of encoding a key picture into a second picture unit.
12. The above step of encoding is The method according to claim 11, further comprising the step of encoding a slice network abstraction layer (NAL) unit for a picture into the first picture unit.
13. The above step of encoding is The method according to claim 9, further comprising the step of encoding a picture unit delimiter (PUD) in the bitstream that indicates that a subsequent NAL unit belongs to the next picture unit.
14. The above step of encoding is The steps include encoding a flag in the PUD to indicate whether the subsequent picture unit is an intra-random access point (IRAP) picture unit or a progressively decoded refresh (GDR) picture unit, The method according to claim 13, further comprising the step of encoding an index in the PUD to indicate the picture type of the coded picture in the subsequent picture unit.
15. A non-temporary computer-readable storage medium for storing a video bitstream, wherein the bitstream is Includes a picture unit containing one or more Supplemental Extended Information (SEI) messages, The aforementioned one or more SEI messages are used to generate one or more pictures on a non-temporary, computer-readable storage medium.
16. The non-temporary computer-readable storage medium according to claim 15, wherein the picture unit further includes a key picture.
17. The aforementioned bitstream is Each contains one or more first picture units, each containing one SEI message, A non-temporary computer-readable storage medium according to claim 15, comprising a second picture unit containing a key picture.
18. The non-temporary computer-readable storage medium according to claim 17, wherein the first picture unit further comprises a coded slice network abstraction layer (NAL) unit for coded pictures.
19. The non-temporary computer-readable storage medium according to claim 15, wherein the bitstream further includes a picture unit delimiter (PUD) in the bitstream indicating that a subsequent NAL unit belongs to the next picture unit.
20. The PUD is, A flag indicating whether the subsequent picture unit is an intra-random access point (IRAP) picture unit or a progressively decoded refresh (GDR) picture unit, A non-temporary computer-readable storage medium according to claim 19, further comprising an index indicating the picture type of the coded picture in the subsequent picture unit.