Image encoding method, image encoding apparatus, image decoding method, image decoding apparatus, method for transmitting bitstream, and recording medium storing bitstream
The image encoding method and device address the high-cost challenge of high-resolution video by employing GFV SEI messages and advanced encoding techniques, achieving efficient transmission and storage of high-quality video.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2025-10-20
- Publication Date
- 2026-04-23
AI Technical Summary
The increasing demand for high-resolution, high-quality video has led to higher transmission and storage costs due to the increase in transmitted information or bits, necessitating high-efficiency video compression technology.
An image encoding method and device that includes generating Generative Face Video (GFV) SEI messages and encoding picture units, along with improved encoding and decoding efficiency through techniques like multi-type tree splitting, Low-Frequency Non-Separable Transform (LFNST), and Context Adaptive Binary Arithmetic Coding (CABAC), and a method for transmitting and storing bitstreams efficiently.
The solution provides enhanced encoding and decoding efficiency, reducing transmission and storage costs while maintaining high-quality video performance.
Smart Images

Figure KR2025016556_23042026_PF_FP_ABST
Abstract
Description
Image encoding method, image encoding device, image decoding method, image decoding device, method for transmitting a bitstream and a recording medium storing a bitstream
[0001] The embodiments relate to an image encoding method, an image encoding device, an image decoding method, an image decoding device, a method for transmitting a bitstream, and a recording medium storing a bitstream.
[0002] Recently, the demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition), has been increasing across various fields. As video data becomes higher in resolution and quality, the relative amount of information or bits transmitted increases compared to conventional video data. This increase in transmitted information or bits leads to higher transmission and storage costs.
[0003] Accordingly, high-efficiency video compression technology is required to effectively transmit, store, and play back high-resolution, high-quality video information.
[0004] The embodiments provide an image encoding method, an image encoding device, an image decoding method, an image decoding device, a method for transmitting a bitstream, and a recording medium storing a bitstream.
[0005] The embodiments provide an image encoding method with improved encoding and decoding efficiency, an image encoding device, an image decoding method, an image decoding device, a method for transmitting a bitstream, and a recording medium storing a bitstream.
[0006] However, the scope of rights of the embodiments is not limited to the technical problems described above, and may be extended to other technical problems that a person skilled in the art can infer based on the entire content described.
[0007] A method according to the embodiments may include the step of acquiring GFV (Generative face video) SEI messages; and the step of generating an output picture based on the GFV (Generative face video) SEI messages. A method according to the embodiments may include the step of encoding a picture unit for the picture; and the step of generating GFV (Generative face video) SEI messages for the picture unit.
[0008] The embodiments provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0009] The embodiments provide a non-transient computer-readable recording medium that stores a bitstream generated by an image encoding method.
[0010] The embodiments provide a non-transient computer-readable recording medium that stores a bitstream received and decoded by an image decoding device and used for image restoration.
[0011] The embodiments provide a method for transmitting a bitstream generated by an image encoding method.
[0012] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs from the description below.
[0013] Drawings are included to further understand the embodiments, and the drawings illustrate the embodiments along with descriptions related to the embodiments. For a better understanding of the various embodiments described below, one must refer to the description of the embodiments below in relation to the following drawings, which include parts corresponding to similar reference numerals throughout the drawings.
[0014] FIG. 1 shows a video and / or image coding system according to embodiments.
[0015] FIG. 2 shows an encoding device according to embodiments.
[0016] FIG. 3 shows a decoding device according to embodiments.
[0017] Figure 4 shows the structure of a content streaming system according to embodiments.
[0018] FIG. 5 shows an example of a picture divided into Coding Tree Units (CTUs) according to embodiments.
[0019] FIG. 6 shows an example of a picture partitioned into tiles and raster-scan slices according to embodiments.
[0020] FIG. 7 shows an example of a picture partitioned into tiles and raster-scan slices according to embodiments.
[0021] FIG. 8 shows an example of a picture partitioned into tiles, bricks, and rectangular slices according to embodiments.
[0022] FIG. 9 shows an example of a picture including subpictures according to embodiments.
[0023] FIG. 10 shows an example of a picture including tiles and CTUs according to embodiments.
[0024] FIG. 11 shows a multi-type tree splitting mode according to embodiments.
[0025] FIG. 12 shows splitting flags within a quad tree of a multi-type tree coding structure according to embodiments.
[0026] FIG. 13 shows an example of a quad tree of a multi-type tree coding block structure according to embodiments.
[0027] FIG. 14 shows the prohibition of TT (Ternary Tree) division for coding blocks according to embodiments.
[0028] FIG. 15 shows the transform and inverse transform according to the embodiments.
[0029] FIG. 16 shows a Low-Frequency Non-Separable Transform (LFNST) according to embodiments.
[0030] FIG. 17 shows CABAC (Context Adaptive Binary Arithmetic Coding) encoding according to embodiments.
[0031] FIG. 18 illustrates an entropy encoding method according to embodiments.
[0032] FIG. 19 illustrates an entropy decoding method according to embodiments.
[0033] FIG. 20 illustrates a picture decoding method according to embodiments.
[0034] FIG. 21 illustrates a picture encoding method according to embodiments.
[0035] FIG. 22 shows a hierarchical structure for a coded image according to embodiments.
[0036] FIGS. 23a, FIGS. 23b, FIGS. 23c, FIGS. 23d, and FIGS. 23e show picture header structures according to embodiments.
[0037] FIGS. 24a, FIGS. 24b, and FIGS. 24c show neural-network post-filter characteristics SEI message syntax according to embodiments.
[0038] FIG. 25 illustrates the process of inducing a luma channel in a luma component according to the embodiments.
[0039] FIG. 26 shows the syntax of a neural network post-filter activation SEI (Supplemental enhancement information) message according to embodiments.
[0040] FIG. 27 shows the syntax of a neural network post-pillar group characteristic SEI message according to embodiments.
[0041] FIG. 28 shows the syntax of a neural network post-filter group activation SEI message according to embodiments.
[0042] FIG. 29 shows source picture timing information according to embodiments.
[0043] FIGS. 30a and FIGS. 30b show object mask information SEI messages according to embodiments.
[0044] FIG. 31 shows an SEI processing order SEI message according to embodiments.
[0045] FIG. 32 shows a processing order nesting SEI message according to embodiments.
[0046] FIG. 33 shows the syntax of an encoder optimization information SEI message according to embodiments.
[0047] FIG. 34 shows the syntax of a text description information SEI message according to embodiments.
[0048] FIGS. 35a, 35b, 35c, and 35d show the syntax of a generative face video SEI message according to the embodiments.
[0049] FIG. 36 shows the syntax of a generative face video SEI message according to the embodiments.
[0050] FIG. 37 shows a encoding method according to embodiments.
[0051] FIG. 38 illustrates a decoding method according to embodiments.
[0052] Preferred embodiments of the embodiments are described in detail, and examples thereof are shown in the accompanying drawings. The following detailed description, with reference to the accompanying drawings, is intended to describe preferred embodiments of the embodiments rather than merely embodiments that may be implemented according to the embodiments. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it is obvious to those skilled in the art that the embodiments may be practiced without these details.
[0053] Most terms used in the embodiments are selected from those commonly used in the field, but some terms are chosen at the applicant's discretion, and their meanings are described in detail in the following description as necessary. Accordingly, the embodiments should be understood based on the intended meaning of the terms, rather than their mere names or meanings.
[0054] Related technical fields: Versatile Video Coding (VVC), Versatile supplemental enhancement information messages for coded video bitstreams (VSEI), Additional SEI messages for VSEI (Draft 3), SEI processing order and processing order nesting SEI messages in VVC (draft 7), Technologies under consideration for future extensions of VSEI (draft 4), SEI messages for VSEI version 4 (Draft 2).
[0055] FIG. 1 shows a video and / or image coding system according to embodiments.
[0056] As shown in FIG. 1, a video / image coding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image information or data to the receiving device in the form of a file or streaming via a digital storage medium or a network.
[0057] The source device may include a video source, an encoding device, and a transmission unit. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.
[0058] A video source may acquire video / images through processes such as video / image capture, synthesis, or generation. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include, for example, a computer, a tablet, and a smartphone, etc., and may generate video / images (electronically). For example, virtual video / images may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process in which related data is generated.
[0059] The encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0060] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device in the form of a file or streaming via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmission unit may include elements for creating a media file through a predetermined file format and elements for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0061] The decoding device can decode video / images by performing a series of procedures such as inverse quantization, inverse transform, and prediction corresponding to the operation of the encoding device.
[0062] The renderer can render the decoded video / image. The rendered video / image can be displayed through the display unit.
[0063] This document relates to video / video coding. For example, the methods / executions disclosed in this document may be applied to methods disclosed in the VVC (versatile video coding) standard, EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard) or next-generation video / video coding standards (e.g., H.267 or H.268).
[0064] This document presents various embodiments regarding video / image coding, and unless otherwise noted, the embodiments may be performed in combination with one another.
[0065] In this document, "video" may refer to a set of images over time. "Picture" generally refers to a unit representing a single image at a specific time, and "slice" or "tile" are units that constitute a part of a picture in coding. A slice or tile may contain one or more CTUs (coding tree units). A single picture may consist of one or more slices or tiles. A single picture may consist of one or more tile groups. A tile group may contain one or more tiles. A "brick" may represent a rectangular area of rows of CTUs within a tile in a picture.
[0066] A brick can represent a rectangular area of a row of CTUs within a tile in a picture. A tile can be divided into multiple bricks, and each brick consists of one or more rows of CTUs within the tile. A tile that is not divided into multiple bricks is also referred to as a brick. A brick scan is a specific sequential order of CTUs that divides a picture. In a brick, CTUs are arranged sequentially as CTU raster scans; in a brick within a tile, bricks are arranged sequentially as brick raster scans of the tile; and in a tile within a picture, tiles are arranged sequentially as tile raster scans of the picture. A tile is a rectangular area of CTUs within a specific tile column and a specific tile row in a picture. A tile column is a rectangular area of CTUs that is equal to the height of the picture and has a width specified by the syntax element of the picture parameter set. A tile row is a rectangular area of CTUs that has a height specified by the syntax element of the picture parameter set and a width equal to the width of the picture. A tile scan is a specific sequential order of CTUs that divides a picture. In tiles, CTUs are continuously aligned by CTU raster scans, and in pictures, tiles are continuously aligned by tile raster scans. A slice contains an integer number of bricks of a picture, which can be contained exclusively in a single NAL unit. A slice can consist of multiple complete tiles or a sequence in which the complete bricks of a single tile are arranged continuously.
[0067] In this document, tile group and slice may be used interchangeably. For example, in this document, tile group / tile group header may be referred to as slice / slice header.
[0068] A pixel or pel can refer to the smallest unit that constitutes a picture (or image). Additionally, the term 'sample' may be used as a counterpart to pixel. Generally, a sample can represent a pixel or its value, and it may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chroma component.
[0069] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to that area. A unit may include one luminance block and two chroma (e.g., cb, cr) blocks. Depending on the case, the term unit may be used interchangeably with terms such as block or area. In general, an MxN block may include samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows.
[0070] FIG. 2 shows an encoding device according to embodiments.
[0071] FIG. 2 shows a schematic block diagram of an encoding device to which the embodiment(s) of the present document can be applied and to which video / image signal encoding is performed.
[0072] As shown in FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a predictor (220), a residual processor (230), an entropy encoder (240), an adder (250), a filter (260), and a memory (270). The predictor (220) may include an inter-predictor (221) and an intra-predictor (222). The residual processor (230) may include a transformer (232), a quantizer (233), a dequantizer (234), and an inverse transformer (235). The residual processor (230) may further include a subtractor (231). The addition unit (250) may be referred to as a reconstructor or a reconstructed block generator. The above-described image segmentation unit (210), prediction unit (220), residual processing unit (230), entropy encoding unit (240), addition unit (250), and filtering unit (260) may be configured by one or more hardware components (e.g., an encoder chipset or processor) according to the embodiment. Additionally, the memory (270) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.
[0073] The image segmentation unit (210) can divide an input image (or picture, frame) input to an encoding device (200) into one or more processing units. For example, a processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, a single coding unit may be divided into multiple coding units of a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first and the binary-tree structure and / or ternary structure may be applied later. Or, the binary-tree structure may be applied first. A coding procedure according to this document may be performed based on the final coding unit that is no longer divided. In this case, based on coding efficiency according to image characteristics, the maximum coding unit may be used directly as the final coding unit, or, if necessary, the coding unit may be recursively divided into lower-depth coding units so that a coding unit of the optimal size is used as the final coding unit. Here, the term "coding procedure" may include procedures such as prediction, transformation, and restoration described below. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may each be divided or partitioned from the aforementioned final coding unit. The prediction unit may be a unit for sample prediction, and the transformation unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from transformation coefficients.
[0074] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the context. In general, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chroma component. A sample may be used to refer to a single picture (or image) as a term corresponding to a pixel or pel.
[0075] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (predicted block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input image signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, as illustrated, the unit that subtracts the prediction signal (predicted block, prediction sample array) from the input image signal (original block, original sample array) within the encoder (200) may be called a subtraction unit (231). The prediction unit performs a prediction for a block to be processed (hereinafter referred to as the current block) and can generate a predicted block containing prediction samples for the current block. The prediction unit can determine whether intra prediction is applied or inter prediction is applied at the current block or CU level. The prediction unit can generate various information regarding prediction, such as prediction mode information, as described below in the description of each prediction mode, and transmit it to the entropy encoding unit (240). The information regarding prediction can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.
[0076] The intra prediction unit (222) can predict the current block by referencing samples within the current picture. The referenced samples may be located near the current block or away from it, depending on the prediction mode. In intra prediction, the prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a Planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is merely an example, and depending on the settings, more or fewer directional prediction modes may be used. The intra prediction unit (222) may also determine the prediction mode applied to the current block by using the prediction mode applied to the surrounding blocks.
[0077] The inter prediction unit (221) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference picture indices. Motion information may further include information on inter prediction directions (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture containing the reference blocks and the reference picture containing the temporal neighboring blocks may be the same or different. Temporal surrounding blocks may be referred to by names such as collocated reference block, collocated CU (colCU), etc., and a reference picture containing temporal surrounding blocks may be referred to as a collocated picture (colPic). For example, the inter prediction unit (221) may construct a list of motion information candidates based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of surrounding blocks as motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0078] The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit may apply intra prediction or inter prediction for a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called combined inter and intra prediction (CIIP). Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, for example, screen content coding (SCC). IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. IBC may utilize at least one of the inter prediction techniques described in this document. The palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, sample values within the picture can be signaled based on information regarding the palette table and palette index.
[0079] The prediction signal generated through the prediction unit (including the inter prediction unit (221) and / or the intra prediction unit (222)) may be used to generate a restored signal or to generate a residual signal. The transformation unit (232) may generate transform coefficients by applying a transformation technique to the residual signal. For example, the transformation technique may include at least one of the Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loeve Transform (KLT), Graph-Based Transform (GBT), or Conditionally Non-linear Transform (CNT). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transformation obtained based on generating a prediction signal using all previously reconstructed pixels. In addition, the transformation process can be applied to pixel blocks of the same square size, or to non-square blocks of variable size.
[0080] The quantization unit (233) quantizes the transformation coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information regarding the quantized transformation coefficients) and output it as a bitstream. The information regarding the quantized transformation coefficients may be called residual information. The quantization unit (233) can rearrange the block-shaped quantized transformation coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information regarding the quantized transformation coefficients based on the one-dimensional vector-shaped quantized transformation coefficients. The entropy encoding unit (240) can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit (240) may encode information necessary for video / image restoration (e.g., values of syntax elements) together or separately, in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information regarding various parameter sets, such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Additionally, the video / image information may further include general constraint information. In this document, information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded through the encoding procedure described above and included in the bitstream.The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) that transmits the signal output from the entropy encoding unit (240) and / or a storage unit (not shown) that stores it may be configured as internal / external elements of the encoding device (200), or the transmission unit may be included in the entropy encoding unit (240).
[0081] The quantized transform coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, a residual signal (residual block or residual samples) can be restored by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (234) and the inverse transformation unit (235). The adder (155) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter-prediction unit (221) or the intra-prediction unit (222). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The adder (250) may be called a reconstructed unit or a reconstructed block generation unit. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, and can also be used for inter prediction of the next picture after filtering as described below.
[0082] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.
[0083] The filtering unit (260) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (260) can generate a modified restored picture by applying various filtering methods to the restored picture, and can store the modified restored picture in memory (270), specifically in the DPB of memory (270). Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (240), as described below in the description of each filtering method. The information regarding filtering can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.
[0084] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter-prediction unit (221). Through this, when inter-prediction is applied, the encoding device can avoid prediction mismatches between the encoding device (100) and the decoding device, and can also improve encoding efficiency.
[0085] The memory (270) DPB can store the modified restored picture to be used as a reference picture in the inter-prediction unit (221). The memory (270) can store motion information of blocks from which motion information is derived (or encoded) within the current picture and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (221) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (270) can store restoration samples of the blocks restored within the current picture and transmit them to the intra-prediction unit (222).
[0086] FIG. 3 shows a decoding device according to embodiments.
[0087] FIG. 3 shows a schematic block diagram of a decoding device to which the embodiment(s) of the present document can be applied and to which decoding of a video / image signal is performed.
[0088] As shown in FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-predictor (331) and an intra-predictor (332). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321). The aforementioned entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) may be configured by a single hardware component (e.g., a decoder chipset or a processor) according to an embodiment. Additionally, the memory (360) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.
[0089] When a bitstream containing video / image information is input, the decoding device (300) can restore the image in correspondence with the process in which the video / image information is processed by the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied by the encoding device. Thus, the processing unit for decoding may be, for example, a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a binary tree structure. One or more conversion units may be derived from the coding unit. And, the restored image signal decoded and output through the decoding device (300) can be played back through a playback device.
[0090] The decoding device (300) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through the entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may further include general constraint information. The decoding device can decode the picture based further on information regarding the parameter sets and / or general constraint information. The signaling / receiving information and / or syntax elements described below in this document can be obtained from the bitstream by decoding through a decoding procedure. For example, the entropy decoding unit (310) can decode information within a bitstream based on coding methods such as exponential chord coding, CAVLC, or CABAC, and output values of syntax elements required for image restoration and quantized values of transformation coefficients regarding residuals. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, determine a context model using information of the syntax element to be decoded and decoding information of surrounding and decoding target blocks or information of a symbol / bin decoded in the previous step, predict the probability of occurrence of the bin according to the determined context model, and perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model.Information regarding prediction among the information decoded in the entropy decoding unit (310) is provided to the prediction unit (inter prediction unit (332) and intra prediction unit (331)), and residual values for which entropy decoding has been performed in the entropy decoding unit (310), for example, quantized transformation coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive residual signals (residual blocks, residual samples, residual sample arrays). Additionally, information regarding filtering among the information decoded in the entropy decoding unit (310) can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (300), or the receiving unit may be a component of the entropy decoding unit (310). Meanwhile, the decoding device according to the present document may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoding unit (310), and the sample decoder may include at least one of an inverse quantization unit (321), an inverse transform unit (322), an adder (340), a filtering unit (350), a memory (360), an inter prediction unit (332), and an intra prediction unit (331).
[0091] In the inverse quantization unit (321), the quantized transformation coefficients can be inversely quantized to output transformation coefficients. The inverse quantization unit (321) can rearrange the quantized transformation coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain transformation coefficients.
[0092] In the inverse conversion unit (322), the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).
[0093] The prediction unit can perform a prediction for the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit (310), the prediction unit can determine whether an intra prediction or an inter prediction is applied to the current block and can determine a specific intra / inter prediction mode.
[0094] The prediction unit (320) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit may apply intra prediction or inter prediction for a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called combined inter and intra prediction (CIIP). Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, for example, screen content coding (SCC). IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. IBC may utilize at least one of the inter prediction techniques described in this document. The palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and palette index may be included in the video / image information and signaled.
[0095] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The referenced samples may be located near the current block or away from it, depending on the prediction mode. In intra prediction, the prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit (331) may determine the prediction mode applied to the current block by using the prediction mode applied to the surrounding blocks.
[0096] The inter prediction unit (332) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. Motion information may include a motion vector and a reference picture index. Motion information may further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) may construct a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction may be performed based on various prediction modes, and information regarding the prediction may include information indicating the mode of inter prediction for the current block.
[0097] The adder (340) can generate a restoration signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the prediction signal (predicted block, predicted sample array) output from the prediction unit (including the inter prediction unit (332) and / or the intra prediction unit (331)). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the restoration block.
[0098] The addition unit (340) may be called a restoration unit or a restoration block generation unit. The generated restoration signal may be used for intra-predicting the next block to be processed within the current picture, may be output after filtering as described below, or may be used for inter-predicting the next picture.
[0099] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0100] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can generate a modified restored picture by applying various filtering methods to the restored picture, and can transmit the modified restored picture to memory (360), specifically to the DPB of memory (360). Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0101] The (modified) restored picture stored in the DPB of the memory (360) can be used as a reference picture in the inter-prediction unit (332). The memory (360) can store motion information of blocks from which motion information within the current picture has been derived (or decoded) and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (260) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (360) can store restoration samples of blocks restored within the current picture and transmit them to the intra-prediction unit (331).
[0102] In this specification, the embodiments described in the filtering unit (260), inter prediction unit (221), and intra prediction unit (222) of the encoding device (100) may be applied to the filtering unit (350), inter prediction unit (332), and intra prediction unit (331) of the decoding device (300) in the same or corresponding manner.
[0103] Implementation and Application Examples:
[0104] The embodiments described in this document may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each figure may be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for implementation may be stored on a digital storage medium.
[0105] In addition, the decoding device and encoding device to which the embodiment(s) of this document apply may be included in multimedia broadcasting transmission and reception devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video conversation devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, Video on Demand (VoD) service providers, Over-the-top video (OTT) devices, internet streaming service providers, 3D video devices, virtual reality (VR) devices, augmented reality (AR) devices, video phone video devices, transportation terminals (e.g., vehicle terminals (including autonomous vehicles), airplane terminals, ship terminals, etc.), and medical video devices, and may be used to process video signals or data signals. For example, Over-the-top video (OTT) devices may include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, Digital Video Recorders (DVRs), etc.
[0106] Additionally, the processing method to which the embodiment(s) of this document are applied may be produced in the form of a program that is executed by a computer and may be stored on a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of this document may also be stored on a computer-readable recording medium. A computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. A computer-readable recording medium may include, for example, a Blu-ray disc (BD), a Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, a floppy disk, and an optical data storage device. Additionally, a computer-readable recording medium includes media implemented in the form of a carrier wave (e.g., transmission over the Internet). Additionally, a bitstream generated by an encoding method may be stored on a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0107] Additionally, the embodiment(s) of this document may be implemented as a computer program product by program code, and the program code may be executed on a computer by the embodiment(s) of this document. The program code may be stored on a computer-readable carrier.
[0108] Figure 4 shows the structure of a content streaming system according to embodiments.
[0109] A content streaming system to which the embodiment(s) of this document apply may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0110] The encoding server compresses content input from multimedia input devices, such as smartphones, cameras, and camcorders, into digital data to generate a bitstream, and transmits it to the streaming server. As another example, if multimedia input devices, such as smartphones, cameras, and camcorders, generate the bitstream directly, the encoding server can be omitted.
[0111] A bitstream may be generated by an encoding method or a bitstream generation method to which the embodiment(s) of this document are applied, and a streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0112] The streaming server transmits multimedia data to the user's device based on user requests made through the web server, while the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server forwards the request to the streaming server, which then transmits the multimedia data to the user. In this process, the content streaming system may include a separate control server, which plays the role of managing commands and responses between devices within the content streaming system.
[0113] A streaming server can receive content from a media storage and / or an encoding server. For example, if content is received from an encoding server, it can be received in real time. In this case, to provide a seamless streaming service, the streaming server may store the bitstream for a certain period of time.
[0114] Examples of user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs (head-mounted displays)), digital TVs, desktop computers, digital signage, etc.
[0115] Each server within the content streaming system can be operated as a distributed server, and in this case, data received from each server can be processed in a distributed manner.
[0116] Partitioning structure:
[0117] The video / image coding method according to this document can be performed based on the following partitioning structure. Specifically, the procedures described below, such as prediction, residual processing ((inverse)transform, (inverse)quantization, etc.), syntax element coding, and filtering, can be performed based on CTU and CU (and / or TU, PU) derived based on the partitioning structure. The block partitioning procedure is performed in the image splitting unit (210) of the encoding device described above, and the partitioning-related information can be processed (encoded) in the entropy encoding unit (240) and transmitted to the decoding device in the form of a bitstream. The entropy decoding unit (310) of the decoding device can derive the block partitioning structure of the current picture based on the partitioning-related information obtained from the bitstream, and perform a series of procedures for image decoding (e.g., prediction, residual processing, block / picture restoration, in-loop filtering, etc.) based thereon. The CU size and the TU size may be the same, or multiple TUs may exist within the CU area. Meanwhile, the term CU size generally refers to the luminance component (sample) CB size. The term TU size generally refers to the luminance component (sample) TB size. The chroma component (sample) CB or TB size can be derived based on the luminance component (sample) CB or TB size according to the component ratio based on the color format (chroma format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.) of the picture / image. The TU size can be derived based on maxTbSize. For example, if the CU size is greater than maxTbSize, multiple TUs (TBs) of maxTbSize are derived from C, and conversion / inverse conversion can be performed in TU (TB) units. In addition, for example, when intra prediction is applied, the intra prediction mode / type is derived in units of CU (or CB), and the procedure for deriving surrounding reference samples and generating prediction samples can be performed in units of TU (or TB).In this case, one or more TUs (or TBs) may exist within a single CU (or CB) region, and in this case, the multiple TUs (or TBs) may share the same intra prediction mode / type.
[0118] Additionally, in the coding of video / images according to this document, image processing units may have a hierarchical structure. A picture may be divided into one or more tiles, bricks, slices, and / or tile groups. A slice may contain one or more bricks. A brick may contain one or more CTU rows within a tile. A slice may contain an integer number of bricks in a picture. A tile group may contain one or more tiles. A tile may contain one or more CTUs. A CTU may be divided into one or more CUs. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. A tile group may contain an integer number of tiles based on a tile raster scan within a picture. A slice header may carry information / parameters that can be applied to the corresponding slice (blocks within the slice). If the encoding / decoding device has a multi-core processor, the encoding / decoding procedures for tiles, slices, bricks, and / or tile groups may be processed in parallel. In this document, the terms slice and tile group may be used interchangeably. A tile group header may be referred to as a slice header. Here, a slice may have one of the slice types, including intra (I) slice, predictive (P) slice, and bi-predictive (B) slice. For blocks within an I slice, inter-prediction is not used for prediction, and only intra-prediction may be used. Of course, even in this case, the original sample value may be coded and signaled without prediction.For blocks within a P slice, intra prediction or inter prediction may be used, and if inter prediction is used, only uni prediction may be used. Meanwhile, for blocks within a B slice, intra prediction or inter prediction may be used, and if inter prediction is used, up to bi prediction may be used.
[0119] In the encoder, tile / tile group, brick, slice, and maximum and minimum coding unit sizes are determined based on video characteristics (e.g., resolution) or by considering coding efficiency or parallel processing, and information regarding this or information that can derive it may be included in the bitstream.
[0120] The decoder can obtain information indicating whether the tile / tile group, brick, slias, and CTU within the tile of the current picture have been divided into multiple coding units. Efficiency can be increased by obtaining (transmitting) this information only under specific conditions.
[0121] A slice header (slice header syntax) may include information / parameters that can be applied commonly to slices. An APS (APS syntax) or PPS (PPS syntax) may include information / parameters that can be applied commonly to one or more pictures. An SPS (SPS syntax) may include information / parameters that can be applied commonly to one or more sequences. A VPS (VPS syntax) may include information / parameters that can be applied commonly to multiple layers. A DPS (DPS syntax) may include information / parameters that can be applied commonly across the video. A DPS may include information / parameters related to the concatenation of a CVS (coded video sequence).
[0122] In this document, the term "higher-level syntax" may include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, and slice header syntax.
[0123] In addition, for example, information regarding the division and configuration of tiles / tile groups / bricks / slices can be configured at the encoding stage through high-level syntax and transmitted to a decoding device in the form of a bitstream.
[0124] FIG. 5 shows an example of a picture divided into Coding Tree Units (CTUs) according to embodiments.
[0125] Partitioning of picture into CTUs:
[0126] Pictures can be divided into a sequence of coding tree units (CTUs). A CTU may correspond to a coding tree block (CTB). Alternatively, a CTU may include a coding tree block of luminance samples and two coding tree blocks of corresponding chroma samples. In other words, for a picture containing three sample arrays, a CTU may include an NxN block of luminance samples and two corresponding blocks of chroma samples. FIG. 5 illustrates an example in which a picture is divided into CTUs.
[0127] The maximum allowable size of a CTU for coding and prediction, etc., may differ from the maximum allowable size of a CTU for transformation. For example, the maximum allowable size of a luminance block within a CTU may be 128x128 (even though the maximum size of luminance ring blocks is 64x64).
[0128] FIG. 6 shows an example of a picture partitioned into tiles and raster-scan slices according to embodiments.
[0129] Partitioning of pictures into subpictures, slices, and tiles:
[0130] A picture is divided into one or more tile rows and one or more tile columns. A tile is a sequence of CTUs that includes a rectangular area of the picture. The CTUs within a tile are scanned in the raster scan order within that tile.
[0131] A slice consists of an integer number of complete tiles or an integer number of consecutive complete CTU rows within a picture tile.
[0132] Two modes are supported for slicing: raster scan slice mode and rectangular slice mode. In raster scan slice mode, a slice contains a complete sequence of tiles from a tile raster scan of the picture. In rectangular slice mode, a slice contains multiple complete tiles that make up a rectangular area of the picture, or multiple consecutive rows of complete CTUs of a single tile that make up a rectangular area of the picture. The tiles within a rectangular slice are scanned in the tile raster scan order within the rectangular area corresponding to that slice.
[0133] A sub-picture consists of one or more slices that cover the entire rectangular area of the picture.
[0134] Figure 6 shows an example of splitting a picture into raster scan slices. Here, the picture is divided into 12 tiles and 3 raster scan slices.
[0135] FIG. 7 shows an example of a picture partitioned into tiles and raster-scan slices according to embodiments.
[0136] Figure 7 shows an example of dividing a picture into rectangular slices. Here, the picture is divided into 24 tiles (6 tile columns and 4 tile rows) and 9 rectangular slices.
[0137] FIG. 8 shows an example of a picture partitioned into tiles, bricks, and rectangular slices according to embodiments.
[0138] Figure 8 shows an example of a picture divided into tiles and rectangular slices. Here, the picture is divided into 4 tiles (2 tile columns and 2 tile rows) and 4 rectangular slices.
[0139] FIG. 9 shows an example of a picture including subpictures according to embodiments.
[0140] Fig. 9 shows an example of sub-picture division of a picture. Here, the picture is divided into 28 sub-pictures of various dimensions.
[0141] FIG. 10 shows an example of a picture including tiles and CTUs according to embodiments.
[0142] If the picture is coded using three separate color planes (where separate_colour_plane_flag is 1), the slice contains only one CTU of a color component identified by its color_plane_id value, and each array of color components in the picture consists of slices having the same color_plane_id value. Coded slices with different color_plane_id values within the picture can be interleaved with each other under the constraint that for each color_plane_id value, the coded slice NAL unit having that color_plane_id value is in ascending order of CTU addresses in the tile scan order for the first CTU of each coded slice NAL unit.
[0143] Note - If separate_colour_plane_flag is 0, each CTU of the picture is contained in exactly one slice. If separate_colour_plane_flag is 1, each CTU of the color component is contained in exactly one slice (information for each CTU of the picture exists in exactly three slices, and these three slices have different colour_plane_id values).
[0144] Tile changes the order of CTUs in a picture. If the picture is divided into two or more tiles, the order of CTUs is the raster scan order within each tile, as shown in FIG. 10. In FIG. 10, the picture is divided into two tiles, and each tile has eight CTUs. The order of CTUs within the tile is the raster scan order.
[0145] FIG. 11 shows a multi-type tree splitting mode according to embodiments.
[0146] Partitioning of the CTUs using a tree structure
[0147] A CTU can be partitioned into CUs based on a quad-tree (QT) structure. The quad-tree structure can be referred to as a quaternary tree structure. This is intended to reflect various local characteristics. Meanwhile, in this document, a CTU can be partitioned based on a multitype tree structure partitioning that includes not only quad-trees but also binary trees (BT) and ternary trees (TT). Hereinafter, the term QTBT structure may include quad-tree and binary tree-based partitioning structures, and QTBTTT may include quad-tree, binary tree, and ternary tree-based partitioning structures. Alternatively, the QTBT structure may include quad-tree, binary tree, and ternary tree-based partitioning structures. In a coding tree structure, CUs can have a square or rectangular shape. A CTU can first be partitioned into a quad-tree structure. Subsequently, the leaf nodes of the quad-tree structure can be further partitioned by a multitype tree structure. For example, as shown in FIG. 11, a multitype tree structure may include four partition types in a schematic manner.
[0148] The four splitting types may include vertical binary splitting (SPLIT_BT_VER), horizontal binary splitting (SPLIT_BT_HOR), vertical ternary splitting (SPLIT_TT_VER), and horizontal ternary splitting (SPLIT_TT_HOR). Leaf nodes of a multitype tree structure may be called CUs. These CUs can be used for prediction and transformation procedures. In this document, CUs, PUs, and TUs generally have the same block size. However, if the maximum supported transform length is smaller than the width or height of the color component of the CU, the CU and TU may have different block sizes.
[0149] FIG. 12 shows splitting flags within a quad tree of a multi-type tree coding structure according to embodiments.
[0150] FIG. 12 exemplarily illustrates the signaling mechanism of partition splitting information in a quadtree with nested multi-type tree structure.
[0151] Here, the CTU is treated as the root of the quadtree and is initially partitioned into a quadtree structure. Each quadtree leaf node can subsequently be further partitioned into a multitype tree structure. In the multitype tree structure, a first flag (e.g., mtt_split_cu_flag) is signaled to indicate whether the node is further partitioned. If the node is further partitioned, a second flag (e.g., mtt_split_cu_vertical_flag) may be signaled to indicate the splitting direction. Subsequently, a third flag (e.g., mtt_split_cu_binary_flag) may be signaled to indicate whether the splitting type is binary or binary. For example, based on mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the multi-type tree splitting mode (MttSplitMode) of CU can be derived as shown in Table 1 (MttSplitMode derviation based on multi-type tree syntax elements).
[0152] [Table 1]
[0153]
[0154] FIG. 13 shows an example of a quad tree of a multi-type tree coding block structure according to embodiments.
[0155] FIG. 13 exemplarily illustrates a CTU being divided into multiple CUs based on a quadtree and nested multi-type tree structure.
[0156] Here, bold block edges represent quadtree partitioning, and the remaining edges represent multitype tree partitioning. Quadtree partitioning involving a multitype tree can provide a content-adapted coding tree structure. A CU can correspond to a coding block (CB). Alternatively, a CU may include a coding block of luminance samples and two coding blocks of corresponding chroma samples. The size of a CU may be as large as a CTU, or it may be 4x4 in luminance sample units. For example, in the case of a 4:2:0 color format (or chroma format), the maximum chroma CB size may be 64x64 and the minimum chroma CB size may be 2x2.
[0157] For example, in this document, the maximum allowable luma TB size may be 64x64 and the maximum allowable chroma TB size may be 32x32. If the width or height of a CB partitioned according to the tree structure is greater than the maximum conversion width or height, the CB may be automatically (or implicitly) partitioned until the horizontal and vertical TB size limits are satisfied.
[0158] Meanwhile, for a quadtree coding tree scheme involving a multitype tree, the following parameters can be defined and identified as SPS syntax elements.
[0159] CTU size: Size of the root node of a 4th-order tree
[0160] MinQTSize: Minimum allowed 4th-order tree leaf node size
[0161] MaxBtSize: Maximum allowed binary tree root node size
[0162] MaxTtSize: Maximum allowed ternary tree root node size
[0163] MaxMttDepth: The maximum allowed hierarchy depth of a multi-type tree splitting at a 4th-order tree leaf.
[0164] MinBtSize: Minimum allowed binary tree leaf node size
[0165] MinTtSize: Minimum allowed tertiary tree leaf node size
[0166] As an example of a quadtree coding tree structure involving a multitype tree, the CTU size can be set to 64x64 blocks of 128x128 luminance samples and two corresponding chroma samples (in the 4:2:0 chroma format). In this case, MinOTSize can be set to 16x16, MaxBtSize to 128x128, MaxTtSize to 64x64, MinBtSize and MinTtSize (for both width and height) to 4x4, and MaxMttDepth to 4. Quadtree partitioning can be applied to the CTU to create quadtree leaf nodes. Quadtree leaf nodes can be called leaf QT nodes. Quadtree leaf nodes can have sizes ranging from 16x16 (i.e., the MinOTSize) to 128x128 (i.e., the CTU size). If the leaf QT node is 128x128, it may not be further split into a binary tree / binary tree. This is because even if it were split in this case, it would exceed MaxBtsize and MaxTtsize (i.e., 64x64). Otherwise, the leaf QT node may be further split into a multitype tree. Therefore, the leaf QT node is the root node of the multitype tree, and the leaf QT node can have a multitype tree depth (mttDepth) value of 0. If the multitype tree depth reaches MaxMttdepth (e.g., 4), further splitting may not be considered. If the width of the multitype tree node is equal to MinBtSize and is less than or equal to 2xMinTtSize, further horizontal splitting may not be considered. If the height of a multitype tree node is equal to MinBtSize and less than or equal to 2xMinTtSize, no further vertical splitting may be considered.
[0167] FIG. 14 shows the prohibition of TT (Ternary Tree) division for coding blocks according to embodiments.
[0168] In order to allow 64x64 luminance block and 32x32 chroma pipeline designs in a hardware decoder, TT splitting may be forbidden in certain cases. For example, if the width or height of the luminance coding block is greater than 64, TT splitting may be forbidden, as shown in FIG. 14. Also, for example, if the width or height of the chroma coding block is greater than 32, TT splitting may be forbidden.
[0169] In this document, the coding tree scheme may support Luma and Chroma (component) blocks having separate block tree structures. If Luma and Chroma blocks within a single CTU have the same block tree structure, it may be denoted as SINGLE_TREE. If Luma and Chroma blocks within a single CTU have separate block tree structures, it may be denoted as DUAL_TREE. In this case, the block tree type for the Luma component may be called DUAL_TREE_LUMA, and the block tree type for the Chroma component may be called DUAL_TREE_CHROMA. For P and B slice / tile groups, Luma and Chroma CTBs within a single CTU may be restricted to having the same coding tree structure. However, for I slice / tile groups, Luma and Chroma blocks may have separate block tree structures. If individual block tree mode is applied, the Luma CTB may be divided into CUs based on a specific coding tree structure, and the Chroma CTB may be divided into Chroma CUs based on a different coding tree structure. This may mean that CUs within an I slice / tile group may consist of coding blocks of the Luma component or coding blocks of two Chroma components, and CUs within a P or B slice / tile group may consist of blocks of three color components. In this document, a slice may be referred to as a tile / tile group, and a tile / tile group may be referred to as a slice.
[0170] In the aforementioned "Partitioning of the CTUs using a tree structure," a quadtree coding tree structure involving a multitype tree was described, but the structure in which the CU is partitioned is not limited to this. For example, the BT structure and the TT structure can be interpreted as concepts included in the Multiple Partitioning Tree (MPT) structure, and the CU can be interpreted as being partitioned through the QT structure and the MPT structure. In an example where the CU is partitioned through the QT structure and the MPT structure, the partitioning structure can be determined by signaling a syntax element (e.g., MPT_split_type) containing information regarding how many blocks the leaf node of the QT structure is partitioned into, and a syntax element (e.g., MPT_split_mode) containing information regarding whether the leaf node of the QT structure is partitioned vertically or horizontally.
[0171] In another example, the CU may be divided in a way different from the QT structure, BT structure, or TT structure. That is, unlike when a lower-depth CU is divided into 1 / 4 the size of an upper-depth CU according to the QT structure, or a lower-depth CU is divided into 1 / 2 the size of an upper-depth CU according to the BT structure, or a lower-depth CU is divided into 1 / 4 or 1 / 2 the size of an upper-depth CU according to the TT structure, the lower-depth CU may, in some cases, be divided into 1 / 5, 1 / 3, 3 / 8, 3 / 5, 2 / 3, or 5 / 8 the size of an upper-depth CU, and the method of dividing the CU is not limited thereto.
[0172] Transformation / Inverse Transformation:
[0173] As described above, the encoding device can derive residual blocks (residual samples) based on blocks (predicted samples) predicted through intra / inter / IBC prediction, etc., and can derive quantized transformation coefficients by applying transformation and quantization to the derived residual samples. Information regarding the quantized transformation coefficients (residual information) can be included in the residual coding syntax and output in the form of a bitstream after encoding. The decoding device can obtain information regarding the quantized transformation coefficients (residual information) from the bitstream and derive the quantized transformation coefficients by decoding. The decoding device can derive residual samples by undergoing inverse quantization / inverse transformation based on the quantized transformation coefficients. As described above, at least one of quantization / inverse quantization and / or transformation / inverse transformation may be omitted. When the transform / inverse transform is omitted, the transform coefficients may be called coefficients or residual coefficients, or they may still be called transform coefficients for consistency of representation. Whether the transform / inverse transform is omitted can be signaled based on transform_skip_flag.
[0174] Transformation / inverse transformation can be performed based on transformation kernel(s). For example, according to this document, a multiple transform selection (MTS) scheme may be applied. In this case, some of the sets of multiple transformation kernels may be selected and applied to the current block. Transformation kernels may be referred to by various terms, such as transformation matrix or transformation type. For example, a set of transformation kernels may represent a combination of vertical transformation kernels and horizontal transformation kernels.
[0175] For example, MTS index information (or tu_mts_idx syntax elements) may be generated / encoded in an encoding device and signaled to a decoding device to indicate one of the sets of transformation kernels. For example, the sets of transformation kernels based on the values of the MTS index information may be derived as shown in Table 2 (Specification of trTypeHor and trTypeVer depending on tu_mts_idx[ x ][ y ]), Table 3 (Specification of trTypeHor and trTypeVer depending on cu_sbt_horizontal_flag and cu_sbt_pos_flag), and / or Table 4 (Specification of trTypeHor and trTypeVer depending on predModeIntra).
[0176] [Table 2]
[0177]
[0178] The set of transformation kernels may be determined, for example, based on cu_sbt_horizontal_flag and cu__sbt_pos_flag.
[0179] If cu_sbt_horizontal_flag is 1, it indicates that the current coding unit is divided horizontally into two transformation units. If cu_sbt_horizontal_flag[ x0 ][ y0 ] is 0, it indicates that the current coding unit is divided vertically into two transformation units. If cu_sbt_pos_flag is 1, it indicates that tu_cbf_luma, tu_cbf_cb, and tu_cbf_cr of the first transformation unit of the current coding unit are not in the bitstream. If cu_sbt_pos_flag is 0, it indicates that tu_cbf_luma, tu_cbf_cb, and tu_cbf_cr of the second transformation unit of the current coding unit are not in the bitstream.
[0180] [Table 3]
[0181]
[0182] The set of transformation kernels may be determined, for example, based on the intra prediction mode for the current block.
[0183] [Table 4]
[0184]
[0185] In the tables above, trTypeHor can represent a horizontal direction conversion kernel, and trTypeVer can represent a vertical direction conversion kernel. Here, a trTypeHor / trTypeVer value of 0 can represent DCT2, a trTypeHor / trTypeVer value of 1 can represent DST7, and a trTypeHor / trTypeVer value of 2 can represent DCT8. However, this is merely an example, and by convention, other values may be mapped to different DCTs / DSTs.
[0186] The following Table 5 (Transform basis functions of DCT-II / VIII and DSTVII for N-point input) illustrates exemplary basis functions for the aforementioned DCT2, DCT8, and DST7.
[0187] [Table 5]
[0188]
[0189] FIG. 15 shows the transform and inverse transform according to the embodiments.
[0190] In this document, the MTS-based transformation is applied as a primary transform, and a secondary transform may be applied. The secondary transform may be applied only to the coefficients in the upper-left wxh region of the coefficient block to which the primary transform is applied, and may be called the Reduced Secondary Transform (RST). For example, w and / or h may be 4 or 8. In the transformation, the primary and secondary transforms may be applied sequentially to the secondary block, and in the inverse transformation, the inverse secondary transform and the inverse primary transform may be applied sequentially to the transform coefficients. The secondary transform (RST transform) may be called the low frequency coefficients transform (LFCT) or low frequency non-separable transform (LFNST). The inverse secondary transform may be called the inverse LFCT or inverse LFNST.
[0191] FIG. 16 shows a Low-Frequency Non-Separable Transform (LFNST) according to embodiments.
[0192] The LFNST (Low Frequency Inseparable Transform), also known as the Reduced Secondary Transform, is applied between the forward first-order transform and quantization (encoder side), and between the inverse quantization and the inverse first-order transform (decoder side), as shown in FIG. 16. In the LFNST, a 4x4 inseparable transform or an 8x8 inseparable transform is applied depending on the block size. For example, a 4x4 LFNST is applied to small blocks (i.e., minimum value (width, height) < 8), and an 8x8 LFNST is applied to large blocks (i.e., minimum value (width, height) > 4).
[0193] The application of the inseparable transformation used in LFNST is explained as follows, using the input as an example. To apply a 4x4 LFNST, the 4x4 input block X is represented as a vector as follows.
[0194]
[0195]
[0196] Inseparable transformations are as follows: It is calculated as. Here represents the transformation coefficient vector, and T is a 16x16 transformation matrix. 16x1 coefficients The vector is then reconstructed into a 4x4 block using the scan order (horizontal, vertical, or diagonal) of the corresponding block. Factors with smaller indices are placed at smaller scan indices in the 4x4 factor block.
[0197] Transform / inverse transformation can be performed in units of CU or TU. That is, transformation / inverse transformation can be applied to residual samples within a CU or residual samples within a TU. The CU size and the TU size may be the same, or multiple TUs may exist within the CU area. Meanwhile, the term CU size generally refers to the luminance component (sample) CB size. The term TU size generally refers to the luminance component (sample) TB size. The chroma component (sample) CB or TB size can be derived based on the luminance component (sample) CB or TB size according to the component ratio based on the color format (chroma format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.). The TU size can be derived based on maxTbSize. For example, if the CU size is greater than maxTbSize, multiple TU(TB) of maxTbSize are derived from the CU, and conversion / inverse conversion can be performed in TU(TB) units. maxTbSize can be considered for determining whether to apply various intra-prediction types, such as ISP. Information regarding maxTbSize may be determined in advance, or it may be generated and encoded by an encoding device and signaled to a decoding device.
[0198] Quantization / Dequantization:
[0199] As described above, the quantization unit of the encoding device can derive quantized conversion coefficients by applying quantization to conversion coefficients, and the inverse quantization unit of the encoding device or the inverse quantization unit of the decoding device can derive conversion coefficients by applying inverse quantization to quantized conversion coefficients.
[0200] In general, in video / image coding, the quantization rate can be varied, and compression can be adjusted using the varied quantization rate. From an implementation perspective, considering complexity, quantization parameters (QP) can be used instead of directly using the quantization rate. For example, quantization parameters can be integer values from 0 to 63, and each quantization parameter value can correspond to an actual quantization rate. The quantization parameter (QPY) for the luminance component (luma sample) and the quantization parameter (QPC) for the chroma component (chroma sample) can be set differently.
[0201] The quantization process takes a transform coefficient (C) as input and divides it by a quantization rate (Qstep) to obtain a quantized transform coefficient (C'). In this case, considering computational complexity, the quantization rate can be multiplied by a scale to form an integer, and a shift operation can be performed by an amount corresponding to the scale value. A quantization scale can be derived based on the product of the quantization rate and the scale value. In other words, the quantization scale can be derived according to QP. Alternatively, the quantization scale can be applied to the transform coefficient (C) to derive the quantized transform coefficient (C').
[0202] The inverse quantization process is the reverse of the quantization process; by multiplying the quantized transformation coefficients (C') by the quantization rate (Qstep), the reconstructed transformation coefficients (C'') can be obtained based on this. In this case, a level scale can be derived depending on the quantization parameters, and the reconstructed transformation coefficients (C'') can be derived by applying this level scale to the quantized transformation coefficients (C''). The reconstructed transformation coefficients (C'') may differ slightly from the original transformation coefficients (C) due to losses during the transformation and / or quantization processes. Therefore, the encoding device performs inverse quantization in the same manner as the decoding device.
[0203] Meanwhile, adaptive frequency-weighted quantization technology, which adjusts the quantization intensity according to frequency, may be applied. Adaptive frequency-weighted quantization is a method of applying different quantization intensities for each frequency. Adaptive frequency-weighted quantization can apply different quantization intensities for each frequency by utilizing a predefined quantization scaling matrix. That is, the aforementioned quantization / de-quantization process can be performed based further on the quantization scaling matrix. For example, different quantization scaling matrices may be used depending on whether the prediction mode applied to the current block to generate the size of the current block and / or the residual signal of the current block is inter-prediction or intra-prediction. The quantization scaling matrix may be referred to as a quantization matrix or a scaling matrix. The quantization scaling matrix may be predefined. Additionally, for frequency-adaptive scaling, frequency-specific quantization scale information regarding the quantization scaling matrix may be configured / encoded in the encoding device and signaled to the decoding device. Frequency-specific quantization scale information can be referred to as quantization scaling information. Frequency-specific quantization scale information may include scaling list data (scaling_list_data). A (modified) quantization scaling matrix can be derived based on the scaling list data. Additionally, frequency-specific quantization scale information may include present flag information indicating the existence of scaling list data. Alternatively, it may further include information indicating whether scaling list data is modified at a lower level (e.g., PPS or tile group header, etc.) when scaling list data is signaled at a higher level (e.g., SPS).
[0204] Entropy Coding:
[0205] As described above in the description of FIG. 2, part or all of the video / image information may be entropied by the entropy encoding unit (240), and part or all of the video / image information described above in the description of FIG. 3 may be entropied by the entropy decoding unit (310). In this case, the video / image information may be encoded / decoded in units of syntax elements. In this document, the term "information is encoded / decoded" may include encoding / decoding by the method described in this paragraph.
[0206] FIG. 17 shows CABAC (Context Adaptive Binary Arithmetic Coding) encoding according to embodiments.
[0207] Figure 17 shows a block diagram of a CABAC for encoding a single syntax element. The encoding process of the CABAC first converts the input signal into a binary value through binarization if the input signal is a syntax element rather than a binary value. If the input signal is already a binary value, it is bypassed without undergoing binarization. Here, each binary digit 0 or 1 constituting the binary value is called a bin. For example, if the binary string after binarization (bin string) is 110, each of 1, 1, and 0 is called a bin. The bin(s) for a single syntax element can represent the value of the corresponding syntax element.
[0208] Binary bins are input into a regular coding engine or a bypass coding engine. The regular coding engine assigns a context model reflecting probability values to the corresponding bin and encodes the bin based on the assigned context model. The regular coding engine can update the probability model for each bin after performing coding for it. Bins coded in this way are called context-coded bins. The bypass coding engine omits the procedure of estimating probabilities for input bins and the procedure of updating the probability model applied to the bin after coding. Instead of assigning context, it improves coding speed by coding input bins using a uniform probability distribution (e.g., 50:50). Bins coded in this way are called bypass bins. The context model can be assigned and updated per context-coded (regularly coded) bin, and the context model can be indicated based on ctxidx or ctxInc. ctxidx can be derived based on ctxInc. Specifically, for example, the context index (ctxidx) pointing to the context model for each normally coded bean can be derived as the sum of the context index increment (ctxInc) and the context index offset (ctxIdxOffset). Here, ctxInc can be derived differently for each bean. ctxIdxOffset can be represented as the lowest value of ctxIdx. The lowest value of ctxIdx can be called the initial value (initValue) of ctxIdx. ctxIdxOffset is a value generally used to distinguish context models for other syntax elements, and the context model for a single syntax element can be distinguished / derived based on ctxinc.
[0209] In the entropy encoding procedure, it is determined whether to perform encoding through a regular coding engine or a bypass coding engine, and the coding path can be switched. Entropy decoding performs the same process as entropy encoding in reverse order.
[0210] FIG. 18 illustrates an entropy encoding method according to embodiments.
[0211] The entropy coding described in Fig. 17 can be performed, for example, as shown in Fig. 18.
[0212] Referring to FIG. 18, an encoding device (entropy encoding unit) performs an entropy coding procedure regarding image / video information. The image / video information may include partitioning-related information, prediction-related information (e.g., inter / intra prediction distinction information, intra prediction mode information, inter prediction mode information, etc.), residual information, in-loop filtering-related information, etc., or may include various syntax elements related thereto. Entropy coding may be performed on a syntax element basis. S600 to S610 may be performed by the entropy encoding unit (240) of the encoding device of FIG. 2 described above.
[0213] The encoding device performs binarization on the target syntax element (S600). Here, the binarization may be based on various binarization methods, such as the Truncated Rice binarization process and the Fixed-length binarization process, and the binarization method for the target syntax element may be predefined. The binarization procedure may be performed by the binarization unit (242) within the entropy encoding unit (240).
[0214] The encoding device performs entropy encoding on the target syntax element (S610). The encoding device may encode the empty string of the target syntax element based on a regular coding-based (context-based) or bypass coding-based method, such as CABAC (context-adaptive arithmetic coding) or CAVLC (context-adaptive variable length coding), and the output may be included in a bitstream. The entropy encoding procedure may be performed by an entropy encoding processing unit (243) within the entropy encoding unit (240). As previously mentioned, the bitstream may be transmitted to a decoding device via a (digital) storage medium or a network.
[0215] FIG. 19 illustrates an entropy decoding method according to embodiments.
[0216] As shown in FIG. 19, a decoding device (entropy decoding unit) can decode encoded image / video information. The image / video information may include partitioning-related information, prediction-related information (e.g., inter / intra prediction distinction information, intra prediction mode information, inter prediction mode information, etc.), residual information, in-loop filtering-related information, etc., or may include various syntax elements related thereto. Entropy coding can be performed on a syntax element basis. S700 to S710 can be performed by the entropy decoding unit (310) of the decoding device of FIG. 3 described above.
[0217] The decoding device performs binarization on the target syntax element (S700). Here, the binarization may be based on various binarization methods, such as the Truncated Rice binarization process and the Fixed-length binarization process, and the binarization method for the target syntax element may be predefined. The decoding device may derive available empty strings (empty string candidates) for the available values of the target syntax element through the binarization procedure. The binarization procedure may be performed by the binarization unit (312) within the entropy decoding unit (310).
[0218] The decoding device performs entropy decoding for the target syntax element (S710). The decoding device sequentially decodes and parses each bin for the target syntax element from the input bit(s) in the bitstream, and compares the derived bin string with the available bin strings for the corresponding syntax element. If the derived bin string is equal to one of the available bin strings, the value corresponding to the bin string is derived as the value of the corresponding syntax element. If not, the next bit in the bitstream is parsed further, and the procedure described above is performed again. Through this process, information (specific syntax element) can be signaled using variable-length bits without using start bits or end bits for specific information within the bitstream. Through this, relatively fewer bits can be allocated to low values, and overall coding efficiency can be increased.
[0219] The decoding device can decode each bin within a bin string from a bitstream in a context-based or bypass-based manner based on an entropy coding technique such as CABAC or CAVLC. The entropy decoding procedure can be performed by an entropy decoding processing unit (313) within the entropy decoding unit (310). As described above, the bitstream may contain various information for image / video decoding. As previously stated, the bitstream may be transmitted to the decoding device via a (digital) storage medium or a network.
[0220] In this document, a table containing syntax elements (syntax table) may be used to represent the signaling of information from an encoding device to a decoding device. The order of the syntax elements in the table containing syntax elements used in this document may represent the parsing order of the syntax elements from the bitstream. The encoding device may configure and encode the syntax table so that the syntax elements can be parsed by the decoding device in the parsing order, and the decoding device may obtain the values of the syntax elements by parsing and decoding the syntax elements of the corresponding syntax table from the bitstream according to the parsing order.
[0221] FIG. 20 illustrates a picture decoding method according to embodiments.
[0222] General Video / Video Coding Procedures:
[0223] In video coding, the pictures constituting the video can be encoded / decoded according to a series of decoding orders. The picture order corresponding to the output order of the decoded pictures can be set differently from the decoding order, and based on this, not only forward prediction but also reverse prediction can be performed during inter-prediction.
[0224] FIG. 20 illustrates an example of a schematic picture decoding procedure to which the embodiment(s) of the present document are applicable. In FIG. 20, S900 may be performed in the entropy decoding unit (310) of the decoding device described in FIG. 3, S910 may be performed in the prediction unit (330), S920 may be performed in the residual processing unit (320), S930 may be performed in the addition unit (340), and S940 may be performed in the filtering unit (350). S900 may include the information decoding procedure described in the present document, S910 may include the inter / intra prediction procedure described in the present document, S920 may include the residual processing procedure described in the present document, S930 may include the block / picture restoration procedure described in the present document, and S940 may include the in-loop filtering procedure described in the present document.
[0225] As shown in FIG. 20, the picture decoding procedure may include, schematically as described in FIG. 3, a procedure for obtaining image / video information (through decoding) from a bitstream (S900), a picture restoration procedure (S910–S930), and an in-loop filtering procedure for the restored picture (S940). The picture restoration procedure may be performed based on prediction samples and residual samples obtained through the inter / intra prediction (S910) and residual processing (S920, inverse quantization and inverse transformation of quantized transformation coefficients) described in this document. A modified restored picture may be generated through an in-loop filtering procedure for the restored picture generated through the picture restoration procedure, and the modified restored picture may be output as a decoded picture and may also be stored in the decoded picture buffer or memory (360) of the decoding device and used as a reference picture in the inter prediction procedure when decoding the picture thereafter. In some cases, the in-loop filtering procedure may be omitted, in which case the restored picture may be output as a decoded picture and may also be stored in the decoded picture buffer or memory (360) of the decoding device and used as a reference picture in the inter-prediction procedure during subsequent decoding of the picture. The in-loop filtering procedure (S940) may include a deblocking filtering procedure, a sample adaptive offset (SAO) procedure, an adaptive loop filter (ALF) procedure, and / or a bilateral filter procedure, as described above, and some or all of these may be omitted. Additionally, one or some of the deblocking filtering procedure, the sample adaptive offset (SAO) procedure, the adaptive loop filter (ALF) procedure, and the bilateral filter procedure may be applied sequentially, or all of them may be applied sequentially. For example, the SAO procedure may be performed after the deblocking filtering procedure is applied to the restored picture.Alternatively, for example, the ALF procedure may be performed after a deblocking filtering procedure has been applied to the restored picture. This can be performed in the same way on the encoding device.
[0226] FIG. 21 illustrates a picture encoding method according to embodiments.
[0227] FIG. 21 illustrates an example of a schematic picture encoding procedure to which the embodiment(s) of the present document are applicable. In FIG. 21, S800 may be performed in the prediction unit (220) of the encoding device described above in FIG. 2, S810 may be performed in the residual processing unit (230), and S820 may be performed in the entropy encoding unit (240). S800 may include the inter / intra prediction procedure described in the present document, S810 may include the residual processing procedure described in the present document, and S820 may include the information encoding procedure described in the present document.
[0228] As shown in FIG. 21, the picture encoding procedure may include not only a procedure for encoding information for picture restoration (e.g., prediction information, residual information, partitioning information, etc.) in a general manner as described in FIG. 02 and outputting it in the form of a bitstream, but also a procedure for generating a restored picture for the current picture and a procedure for applying in-loop filtering to the restored picture (optional). The encoding device may derive (modified) residual samples from quantized transform coefficients through the inverse quantization unit (234) and the inverse transform unit (235), and may generate a restored picture based on the prediction samples and (modified) residual samples which are the outputs of S800. The restored picture thus generated may be identical to the restored picture generated by the decoding device described above. A modified restored picture can be generated through an in-loop filtering procedure for the restored picture, which can be stored in a decoded picture buffer or memory (270), and, as in the case of a decoding device, can be used as a reference picture in an inter-prediction procedure during the encoding of the picture thereafter. As described above, in some cases, part or all of the in-loop filtering procedure may be omitted. When the in-loop filtering procedure is performed, filtering-related information (parameters) can be encoded in the entropy encoding unit (240) and output in the form of a bitstream, and the decoding device can perform the in-loop filtering procedure in the same way as the encoding device based on the filtering-related information.
[0229] Through this in-loop filtering procedure, noise generated during video coding, such as blocking and ringing artifacts, can be reduced, and subjective and objective visual quality can be enhanced. Furthermore, by performing the in-loop filtering procedure in both the encoding and decoding devices, they can derive identical prediction results, increase the reliability of picture coding, and reduce the amount of data that must be transmitted for picture coding.
[0230] As described above, the picture restoration procedure can be performed not only in the decoding device but also in the encoding device. Restoration blocks can be generated based on intra-prediction / inter-prediction on a block-by-block basis, and a restored picture containing the restoration blocks can be generated. If the current picture / slice / tile group is the I picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based solely on intra-prediction. Meanwhile, if the current picture / slice / tile group is the P or B picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based on intra-prediction or inter-prediction. In this case, inter-prediction may be applied to some blocks within the current picture / slice / tile group, and intra-prediction may be applied to the remaining blocks. The color components of the picture may include luminance components and chroma components, and unless explicitly limited in this document, the methods and embodiments proposed in this document may be applied to luminance components and chroma components.
[0231] Examples of coding hierarchy and structure:
[0232] The coded video / image according to this document can be processed according to, for example, the coding layers and structures described below.
[0233] FIG. 22 shows a hierarchical structure for a coded image according to embodiments.
[0234] FIG. 22 is a diagram illustrating the hierarchical structure of a coded image.
[0235] The coded video is divided into a video coding layer (VCL) that handles the decoding processing of the video and the video itself, a subsystem that transmits and stores the encoded information, and a network abstraction layer (NAL) that exists between the VCL and the subsystem and is responsible for network adaptation functions.
[0236] In VCL, VCL data containing compressed image data (slice data) can be generated, or parameter sets containing information such as Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS), or SEI (Supplemental Enhancement Information) messages that are additionally required in the decoding process of the image can be generated.
[0237] In NAL, a NAL unit can be created by adding header information (NAL unit header) to the Raw Byte Sequence Payload (RBSP) generated in VCL. In this case, the RBSP refers to slice data, parameter sets, SEI messages, etc. generated in VCL. The NAL unit header may include NAL unit type information specified according to the RBSP data included in the NAL unit.
[0238] NAL units can be classified into VCL NAL units and Non-VCL NAL units depending on the RBSP generated in VCL. A VCL NAL unit may refer to a NAL unit containing information about an image (slice data), and a Non-VCL NAL unit may refer to a NAL unit containing information necessary to decode an image (parameter set or SEI message).
[0239] The aforementioned VCL NAL unit and Non-VCL NAL unit can be transmitted over a network by attaching header information according to the data specifications of the underlying system. For example, the NAL unit can be transformed into a data format of a specified specification, such as H.266 / VVC file format, RTP (Real-time Transport Protocol), TS (Transport Stream), etc., and transmitted over various networks.
[0240] The NAL unit type can be determined according to the RBSP data structure included in the NAL unit, and information about this NAL unit type can be stored in the NAL unit header and signaled.
[0241] For example, NAL units can be broadly classified into VCL NAL unit types and Non-VCL NAL unit types depending on whether they contain information about the image (slice data). VCL NAL unit types can be classified according to the properties and types of the picture included in the VCL NAL unit, while Non-VCL NAL unit types can be classified according to the types of parameter sets.
[0242] The following is an example of a NAL unit type specified based on the type of parameter set included by the Non-VCL NAL unit type: APS (Adaptation Parameter Set) NAL unit: A type for a NAL unit containing APS. DPS (Decoding Parameter Set) NAL unit: A type for a NAL unit containing DPS. VPS (Video Parameter Set) NAL unit: A type for a NAL unit containing VPS. SPS (Sequence Parameter Set) NAL unit: A type for a NAL unit containing SPS. PPS (Picture Parameter Set) NAL unit: A type for a NAL unit containing PPS.
[0243] The above-described NAL unit types have syntax information for the NAL unit type, and the syntax information can be stored in the NAL unit header and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit types can be specified by the nal_unit_type value.
[0244] A slice header (slice header syntax) may include information / parameters that can be applied commonly to slices. An APS (APS syntax) or a PPS (PPS syntax) may include information / parameters that can be applied commonly to one or more slices or pictures. An SPS (SPS syntax) may include information / parameters that can be applied commonly to one or more sequences. A VPS (VPS syntax) may include information / parameters that can be applied commonly to multiple layers. A DPS (DPS syntax) may include information / parameters that can be applied commonly across the video. A DPS may include information / parameters related to the concatenation of a CVS (coded video sequence). In this document, High-level syntax (HLS) may include at least one of an APS syntax, a PPS syntax, an SPS syntax, a VPS syntax, a DPS syntax, and a slice header syntax.
[0245] In this document, the image / video information that is encoded from an encoding device to a decoding device and signaled in the form of a bitstream includes not only information related to partitioning within a picture, intra / inter prediction information, residual information, in-loop filtering information, etc., but may also include information included in a slice header, information included in an APS, information included in the PPS, information included in an SPS, and / or information included in a VPS.
[0246] Coding descriptors:
[0247] The following descriptors represent the parsing process for each syntax element: ae(v): Context-adaptive arithmetic entropy-coded syntax element. b(8): A byte containing a bit string of arbitrary patterns (8 bits). The parsing process for this descriptor is specified by the return value of the read_bits(8) function. f(n): A fixed-pattern bit string of n bits written from left to right with the left bit coming first. The parsing process for this descriptor is specified by the return value of the read_bits(n) function. i(n): A signed integer using n bits. If n is "v" in the syntax table, the number of bits depends on the values of other syntax elements. The parsing process for this descriptor is specified by the return value of the read_bits(n) function and is interpreted as a two's complement integer representation with the most significant bit written first. se(v): A signed integer zero-ordered Exp-Golomb-coded syntax element, with the left bit coming first. The parsing process for this descriptor is specified by the order of k being zero. st(v): A null-terminated string encoded in Universal Coded Character Set (UCS) Transfer Format-8 (UTF-8) characters as specified in ISO / IEC 10646. The parsing process is as follows: st(v) moves the bitstream pointer (stringLength + 1) * 8 bit positions starting from the current position in the bitstream's byte alignment position to the next byte alignment byte, such as 0x00 (excluding that byte), where stringLength is equal to the number of bytes returned. The st(v) syntax descriptor is used in this specification only when the current position in the bitstream is the byte alignment position. tu(v): A truncated unary operator using up to maxVal bits. maxVal is defined in the semantics of the symtax element. u(n): An unsigned integer using n bits.In the syntax table, if n is "v", the number of bits depends on the values of other syntax elements. The parsing process of this descriptor is specified by the return value of the function read_bits(n) and is interpreted as a binary representation of an unsigned integer with the most significant bit written first. ue(v): An unsigned integer of a zero-order exponent Colomb-coded syntax element with the left bit written first. The parsing process of this descriptor is specified by setting the order of k to 0.
[0248] High-level syntax signaling and semantics are described below with reference to each figure.
[0249] FIG. 23 shows a picture header structure according to embodiments.
[0250] Picture header and slice header:
[0251] A coded picture may consist of one or more slices. Parameters describing the coded picture are passed within the picture header (PH), and parameters describing the slice are passed within the slice header. The PH is passed as its own NAL unit type. The SH is located at the beginning of the NAL unit containing the slice's payload (e.g., slice data). For details on the syntax and semantics of the PH and SH, refer to Section 7 of the VVC specification.
[0252] SEI Messages:
[0253] Neural-network post-filter SEI messages
[0254] General post-processing filtering processes using NNPFs
[0255] The input to this process is a bitstream BitstreamToFilter. The output of this process is a list of NNPF output pictures, ListNnpfOutputPics.
[0256] First, BitstreamToFilter is decoded, and the CroppedDecodedPictures list is set as a list of decoded pictures cropped in the order of the BitstreamToFilter decoding output.
[0257] Second, a filtering process for one picture is in CroppedDecodedPictures and is repeatedly called in output order for each cropped decoded picture with one or more NNPFs enabled.
[0258] The order of the pictures in ListNnpfOutputPics is the output order.
[0259] There is only one picture associated with a specific output time instance within ListNnpfOutputPics. If there are multiple NNPFs enabled for a specific picture in CroppedDecodedPictures and only one NNPF can be selected to apply (other NNPFs can also be selected), the above constraints apply regardless of which NNPF is applied to the specific picture.
[0260] Single Picture Filtering Process Using NNPF:
[0261] The filtering process is applied to each cropped decoded picture (referred to as the current picture) that belongs to CroppedDecodedPictures and has one or more NNPFs enabled.
[0262] When applying NNPF to the current picture, the filtered and / or interpolated picture is generated by NNPF by applying the NNPF process specified in the semantics of the NNPFFC SEI message to the current picture in a patch manner.
[0263] When applying NNPF to the current picture, the order of the picture generated by NNPF by applying the NNPF process is the same as the output order stored in the output tensor of NNPF.
[0264] If the applied NNPF is the last NNPF applied to the current picture, the picture generated by the NNPF and the picture output from the NNPF process are included in ListNnpfOutputPics, in the same order as when the picture is stored in the output tensor of the NNPF.
[0265] FIG. 24 shows the syntax of a neural-network post-filter characteristics SEI message according to embodiments.
[0266] The syntax of the NNPFC SEI message associated with the Neural-network post-filter characteristics SEI message (NNPFC) is as shown in FIG. 24.
[0267] The NNPFC SEI message represents a neural network that can be used as a post-processing filter. The use of a specified neural network post-processing filter (NNPF) for a particular picture is indicated by the Neural Network Post-processing Filter Activation (NNPFA) SEI message.
[0268] To use this SEI message, the following variables must be defined.
[0269] Input picture width and height in Luma sample units (labeled as CroppedWidth and CroppedHeight, respectively).
[0270] CroppedYPic[idx], an array of luminance samples, and CroppedCbPic[idx] and CroppedCrPic[idx] (if any), chroma sample arrays of input pictures whose index idx used as input to NNPF is in the range from 0 to numInputPics - 1 (inclusive).
[0271] Bit depth for the luminance sample array of the input picture BitDepthY.
[0272] BitDepthC for the chroma sample array (if any) of the input picture.
[0273] Chroma format indicator displayed as ChromaFormatIdc.
[0274] If nnpfc_auxiliary_inp_idc is 1, the filtering strength control value array StrengthControlVal[idx] contains real numbers in the range of 0 to 1 (inclusive) for input pictures where index idx is in the range of 0 to numInputPics - 1 (inclusive).
[0275] The input picture with index 0 corresponds to the picture in which the NNPF defined in this NNPFC SEI message is activated by the NNPFA SEI message. Input pictures with index i in the range from 1 to numInputPics-1 take precedence over the input picture with index i-1 in the output order.
[0276] The SubWidthC and SubHeightC variables are derived from ChromaFormatIdc.
[0277] Two or more NNPFC SEI messages may exist for the same picture. If two or more NNPFC SEI messages with different nnpfc_id values exist or are enabled for the same picture, the nnpfc_purpose and nnpfc_mode_idc values of those messages may be the same or different.
[0278] nnpfc_purpose represents the purpose of the NNPF specified in Table 6 (Definition of nnpfc_purpose). Here, if (nnpfc_purpose & bitMask) is not 0, it indicates that the NNPF has a purpose associated with the bitMask value in Table 6. If nnpfc_purpose is greater than 0 and (nnpfc_purpose & bitMask) is 0, the purpose associated with the bitMask value cannot be applied to the NNPF. If nnpfc_purpose is 0, the NNPF can be used as determined by the application.
[0279] The value of nnpfc_purpose is in the range of 0 to 63 in bitstreams conforming to this version of this document. Values for nnpfc_purpose from 64 to 65,535 (inclusive) are reserved for future use by ITU-T | ISO / IEC and do not exist in bitstreams conforming to this version of this document. Decoders conforming to this version of this document ignore NNPFC SEI messages with nnpfc_purpose in the range of 64 to 65,535.
[0280] [Table 6]
[0281]
[0282] The variables chromaUpsamplingFlag, resolutionResamplingFlag, pictureRateUpsamplingFlag, bitDepthUpsamplingFlag, and colourizationFlag, which respectively specify whether nnpfc_purpose includes chroma upsampling, resolution resampling, picture rate upsampling, bit depth upsampling, and colourization in the NNPF purpose, are derived as follows: chromaUpsamplingFlag = ( ( nnpfc_purpose & 0x02 ) > 0 ) ? 1 : 0
[0283] resolutionResamplingFlag = ( ( nnpfc_purpose & 0x04 ) > 0 ) ? 1:0
[0284] pictureRateUpsamplingFlag = ((nnpfc_purpose & 0x08) > 0)? 1:0 (76)
[0285] bitDepthUpsamplingFlag = ( ( nnpfc_purpose & 0x10 ) > 0 ) ? 1:0
[0286] colourizationFlag = ( ( nnpfc_purpose & 0x20 ) > 0 ) ? 1:0
[0287] If the reserved value of nnpfc_purpose is used in the future by ITU-T | ISO / IEC, the syntax of this SEI message may be expanded with syntax elements depending on whether nnpfc_purpose is the same as that value.
[0288] If ChromaFormatIdc is 3, chromaUpsamplingFlag becomes 0.
[0289] If ChromaFormatIdc or chromaUpsamplingFlag is not 0, colourizationFlag becomes 0.
[0290] If the input picture with pictureRateUpsamplingFlag 1 and index 0 is associated with a frame packing array SEI message with fp_arrangement_type 5, then all input pictures are associated with a frame packing array SEI message with fp_arrangement_type 5 and the same fp_current_frame_is_frame0_flag value.
[0291] nnpfc_id contains an identification number that can be used to identify the NNPF. The nnpfc_id value ranges from 0 to 2 32It is in the range up to -2 (including 0). Among nnpfc_id values, 256 to 511 (including 0), and 231 to 2 32 Values up to -2 (including 0) are reserved for future use by ITU-T | ISO / IEC. Decoders compliant with this version of this document use nnpfc_id values from 256 to 511 (including 0) or from 231 to 2 32 If an NNPFC SEI message within the range of -2 (including 0) is found, that SEI message is ignored.
[0292] If the NNPFC SEI message is the first NNPFC SEI message with a specific nnpfc_id value within the current CLVS in the decoding order, the following applies.
[0293] This SEI message represents the default NNPF.
[0294] This SEI message is applied to all subsequent decoded pictures of the current layer in output order, from the currently decoded picture to the end of the current CLVS.
[0295] If nnpfc_base_flag is 1, it indicates that the SEI message specifies the default NNPF. If nnpf_base_flag is 0, it indicates that the SEI message specifies an update based on the default NNPF.
[0296] The following constraints apply to the nnpfc_base_flag value.
[0297] If the NNPFC SEI message is the first NNPFC SEI message in decoding order with a specific nnpfc_id value within the current CLVS, the nnpfc_base_flag value is equal to 1.
[0298] If the NNPFC SEI message nnpfcB is not the first NNPFC SEI message in decoding order with a specific nnpfc_id value within the current CLVS, and the nnpfc_base_flag value is 1, the NNPFC SEI message is a repetition of the first NNPFC SEI message nnpfcA with the same nnpfc_id value in decoding order. That is, the payload content of nnpfcB is identical to the payload content of nnpfcA.
[0299] If nnpfc_base_flag is 0, the following applies.
[0300] This SEI message defines updates based on the previous default NNPF with the same nnpfc_id value in decoding order. Updates are not cumulative, and each update is applied to the default NNPF. The default NNPF is the NNPF specified in the first NNPFC SEI message in decoding order and has a specific nnpfc_id value within the current CLVS. The NNPF defined in this SEI message is obtained by applying the updates defined in this SEI message based on the default NNPF with the same nnpfc_id value.
[0301] This SEI message is about the currently decoded picture and all subsequent decoded pictures of the current layer (based on output order), up to the end of the current CLVS or up to the decoded picture following the currently decoded picture in output order within the current CLVS, and is associated with the subsequent NNPFC SEI message based on decoding order, where nnpfc_base_flag is 0 and there is a specific nnpfc_id value within the current CLVS (whichever is earlier).
[0302] If nnpfc_mode_idc is 0, it indicates that this SEI message contains an ISO / IEC 15938-17 bitstream specifying the default NNPF (if nnpfc_base_flag is 1) or is updated based on the default NNPF with the same nnpfc_id value (if nnpfc_base_flag is 0).
[0303] If nnpfc_base_flag is 1, and nnpfc_mode_idc is 1, it indicates that the base NNPF associated with the nnpfc_id value is a neural network in the format identified by a URI represented by nnpfc_uri and a tag URI identified by nnpfc_tag_uri.
[0304] When nnpfc_base_flag is 0, and nnpfc_mode_idc is 1, it indicates that updates to the base NNPF with the same nnpfc_id value are defined by the URI represented by nnpfc_uri and have a format identified by the tag URI nnpfc_tag_uri.
[0305] The nnpfc_mode_idc value is in the range from 0 to 1 in bitstreams compliant with this version of this document. Values of nnpfc_mode_idc from 2 to 255 (inclusive) are reserved for future use by ITU-T | ISO / IEC and do not exist in bitstreams compliant with this version of this document. Decoders compliant with this version of this document ignore NNPFC SEI messages where nnpfc_mode_idc is in the range from 2 to 255. If the nnpfc_mode_idc value is greater than 255, it does not exist in bitstreams compliant with this version of this document and is not reserved for future use.
[0306] nnpfc_reserved_zero_bit_a is equal to 0 in bitstreams following this version of this document. Decoders ignore NNPFC SEI messages where nnpfc_reserved_zero_bit_a is not 0.
[0307] nnpfc_tag_uri contains a tag URI with the syntax and semantics specified in IETF RFC 4151 and identifies the format and related information of a neural network used as an update to a base NNPF or a base NNPF having the same nnpfc_id value specified in nnpfc_uri.
[0308] Using nnpfc_tag_uri allows you to uniquely identify the format of neural network data specified in nnrpf_uri without a central registry.
[0309] If nnpfc_tag_uri is "tag:iso.org,2023:15938-17", it indicates that the neural network data identified by nnpfc_uri complies with ISO / IEC 15938-17.
[0310] nnpfc_uri contains a URI with the syntax and semantics specified in IETF Internet Standard 66, and identifies the neural network used as the base NNPF or the neural network used as an update to the base NNPF with the same nnpfc_id value.
[0311] If nnpfc_property_present_flag is 1, it indicates that syntax elements related to filter purpose, input format, output format, and complexity exist. If nnpfc_property_present_flag is 0, it indicates that syntax elements related to filter purpose, input format, output format, and complexity do not exist.
[0312] If nnpfc_base_flag is 1, then nnpfc_property_present_flag also becomes 1.
[0313] If nnpfc_property_present_flag is 0, the value of all syntax elements that may exist only when nnpfc_property_present_flag is 1 is inferred to be the same as the corresponding syntax element of the NNPFC SEI message containing the underlying NNPF that this SEI message provides updates for.
[0314] If the NNPFC SEI message nnpfcCurr is not the first NNPFC SEI message in decoding order with a specific nnpfc_id value within the current CLVS, and does not overlap with the first NNPFC SEI message with that specific nnpfc_id (i.e., when the nnpfc_base_flag value is 0), or when the nnpfc_property_present_flag value is 1, the following constraints apply.
[0315] The nnpfc_purpose value of an NNPFC SEI message is the same as the nnpfc_purpose value of the first NNPFC SEI message in decoding order that has that specific nnpfc_id value within the current CLVS.
[0316] In NNPFC SEI messages, the syntax element values after nnpfc_property_present_flag and before nnpfc_complexity_info_present_flag in the decoding order are the same as the corresponding syntax element values of the first NNPFC SEI message with the corresponding nnpfc_id value within the current CLVS.
[0317] In the first NNPFC SEI message with the corresponding nnpfc_id value within the current CLVS (indicated as nnpfcBase below), nnpfc_complexity_info_present_flag must be 0 or both 1 in the decoding order, and all of the following apply.
[0318] The nnpfc_parameter_type_idc of nnpfcCurr is the same as the nnpfc_parameter_type_idc of nnpfcBase.
[0319] If nnpfc_log2_parameter_bit_length_minus3 of nnpfcCurr exists, it is less than or equal to nnpfc_log2_parameter_bit_length_minus3 of nnpfcBase.
[0320] If nnpfc_num_parameters_idc of nnpfcBase is 0, then nnpfc_num_parameters_idc of nnpfcCurr also becomes 0.
[0321] Otherwise (if nnpfc_num_parameters_idc of nnpfcBase is greater than 0), nnpfc_num_parameters_idc of nnpfcCurr is greater than 0 and less than or equal to nnpfc_num_parameters_idc of nnpfcBase.
[0322] If nnpfc_num_kmac_operations_idc of nnpfcBase is 0, then nnpfc_num_kmac_operations_idc of nnpfcCurr is also 0.
[0323] Otherwise (if nnpfc_num_kmac_operations_idc of nnpfcBase is greater than 0), nnpfc_num_kmac_operations_idc of nnpfcCurr is greater than 0 and less than or equal to nnpfc_num_kmac_operations_idc of nnpfcBase.
[0324] If the nnpfc_total_kilobyte_size of nnpfcBase is 0, the nnpfc_total_kilobyte_size of nnpfcCurr also becomes 0.
[0325] Otherwise (if nnpfc_total_kilobyte_size of nnpfcBase is greater than 0), nnpfc_total_kilobyte_size of nnpfcCurr is greater than 0 and less than or equal to nnpfc_total_kilobyte_size of nnpfcBase.
[0326] nnpfc_num_input_pics_minus1 + 1 represents the number of pictures used as input to the NNPF. The value of nnpfc_num_input_pics_minus1 ranges from 0 to 63. If pictureRateUpsamplingFlag is 1, the value of nnpfc_num_input_pics_minus1 is greater than 0.
[0327] The variable numInputPics, which specifies the number of pictures used as inputs for NNPF, is derived as follows.
[0328] numInputPics = nnpfc_num_input_pics_minus1 + 1 (77)
[0329] If nnpfc_input_pic_output_flag[ i ] is 1, NNPF indicates that the corresponding output picture is generated for the i-th input picture. If nnpfc_input_pic_output_flag[ i ] is 0, NNPF indicates that the corresponding output picture is not generated for the i-th input picture. If nnpfc_num_input_pics_minus1 is 0, nnpfc_input_pic_output_flag
[0000] is inferred to be 1. If pictureRateUpsamplingFlag is 0 and nnpfc_num_input_pics_minus1 is greater than 0, nnpfc_input_pic_output_flag[ i ] is equal to 1 for at least one value of i in the range from 0 to nnpfc_num_input_pics_minus1.
[0330] If nnpfc_absent_input_pic_zero_flag is 1, it indicates that NNPF should represent input pictures not in the bitstream as a sample array with a sample value of 0. If nnpfc_absent_input_pic_flag is 0, it indicates that NNPF should represent input pictures not in the bitstream as the input picture closest to the output order in the bitstream.
[0331] nnpfc_out_sub_c_flag indicates the values of the outSubWidthC and outSubHeightC variables when chromaUpsamplingFlag is 1. When nnpfc_out_sub_c_flag is 1, it indicates that outSubWidthC is 1 and outSubHeightC is 1. When nnpfc_out_sub_c_flag is 0, it indicates that outSubWidthC is 2 and outSubHeightC is 1. If ChromaFormatIdc is 2 and nnpfc_out_sub_c_flag is present, the value of nnpfc_out_sub_c_flag becomes 1.
[0332] nnpfc_out_colour_format_idc specifies the color format of the NNPF output when colourizationFlag is 1, and consequently represents the values of the outSubWidthC and outSubHeightC variables. If nnpfc_out_colour_format_idc is 1, it indicates that the NNPF output color format is 4:2:0 and both outSubWidthC and outSubHeightC are 2. If nnpfc_out_colour_format_idc is 2, it indicates that the NNPF output color format is 4:2:2 and outSubWidthC is 2 and outSubHeightC is 1. If nnpfc_out_colour_format_idc is 3, it indicates that the NNPF output color format is 4:4:4 and both outSubWidthC and outSubHeightC are 1. The value of nnpfc_out_colour_format_idc is not 0.
[0333] If both chromaUpsamplingFlag and colourizationFlag are 0, outSubWidthC and outSubHeightC are inferred as follows: SubWidthC and SubHeightC.
[0334] nnpfc_pic_width_num_minus1 + 1 and nnpfc_pic_width_denom_minus1 + 1 represent the numerator and denominator for the resampling ratio of the NNPF output picture width relative to CroppedWidth, respectively. The value of (nnpfc_pic_width_num_minus1 + 1) χ (nnpfc_pic_width_denom_minus1 + 1) is in the range of 1 χ 16 to 16. If nnpfc_pic_width_num_minus1 and nnpfc_pic_width_denom_minus1 are absent, the values of nnpfc_pic_width_num_minus1 and nnpfc_pic_width_denom_minus1 are both inferred to be 0.
[0335] The variable nnpfcOutputPicWidth, which represents the width of the luminance sample array of the picture resulting from applying the NNPF identified by nnpfc_id to the input picture, is derived as follows.
[0336] nnpfcOutputPicWidth = Ceil(CroppedWidth *
[0337] ( nnpfc_pic_width_num_minus1 + 1 ) χ ( nnpfc_pic_width_denom_minus1 + 1 ) )
[0338] For bitstream compatibility, the value of nnpfcOutputPicWidth % outSubWidthC is equal to 0.
[0339] nnpfc_pic_height_num_minus1 + 1 and nnpfc_pic_height_denom_minus1 + 1 represent the numerator and denominator for the resampling ratio of the NNPF output picture height relative to CroppedHeight, respectively. The values of ( nnpfc_pic_height_num_minus1 + 1 ) χ ( nnpfc_pic_height_denom_minus1 + 1 ) range from 1 χ 16 to 16. If nnpfc_pic_height_num_minus1 and nnpfc_pic_height_denom_minus1 are absent, the values of nnpfc_pic_height_num_minus1 and nnpfc_pic_height_denom_minus1 are both inferred to be 0.
[0340] The variable nnpfcOutputPicHeight, which represents the height of the luminance sample array of the picture resulting from applying the NNPF identified by nnpfc_id to the input picture, is derived as follows.
[0341] nnpfcOutputPicHeight = Ceil(CroppedHeight *
[0342] ( nnpfc_pic_height_num_minus1 + 1 ) χ ( nnpfc_pic_height_denom_minus1 + 1 ) )
[0343] For bitstream compatibility, the value of nnpfcOutputPicHeight % outSubHeightC is equal to 0.
[0344] If nnpfc_pic_width_num_minus1, nnpfc_pic_width_denom_minus1, nnpfc_pic_height_num_minus1, nnpfc_pic_height_denom_minus1 exist, one or more of the following are true.
[0345] The value of nnpfcOutputPicWidth is not equal to CroppedWidth.
[0346] The value of nnpfcOutputPicHeight is not equal to CroppedHeight.
[0347] nnpfc_interpolated_pics[i] represents the number of interpolated pictures generated by NNPF between the i-th picture used as input to NNPF and the (i + 1)-th picture. The value of nnpfc_interpolated_pics[i] ranges from 0 to 63. The value of nnpfc_interpolated_pics[i] is greater than 0 for at least one of the i values in the range from 0 to nnpfc_num_input_pics_minus1 - 1.
[0348] The variable NumInpPicsInOutputTensor, which specifies the number of pictures in the output tensor of NNPF that have the corresponding input picture, InpIdx[ idx ], which specifies the input picture index of the idx-th picture in the output tensor of NNPF that has the corresponding input picture, and numOutputPics, which specifies the total number of pictures in the output tensor of NNPF, are derived as follows.
[0349] for( i = 0, numOutputPics = 0; i < numInputPics; i++ )
[0350] if(nnpfc_input_pic_output_flag[i]) {
[0351] InpIdx[ numOutputPics ] = i
[0352] numOutputPics++
[0353] }
[0354] NumInpPicsInOutputTensor = numOutputPics
[0355] if(pictureRateUpsamplingFlag)
[0356] for( i = 0; i <= numInputPics - 2; i++ )
[0357] numOutputPics += nnpfc_interpolated_pics[ i ]
[0358] If nnpfc_component_last_flag is 1, it indicates that the last dimension of the input tensor inputTensor for the NNPF and the output tensor outputTensor generated by the NNPF are used for the current channel. If nnpfc_component_last_flag is 0, it indicates that the third dimension of the input tensor inputTensor for the NNPF and the output tensor outputTensor generated by the NNPF are used for the current channel.
[0359] The first dimension of the input and output tensors is used for the batch index, which is the approach used in some neural network frameworks. The formula in the semantics of this SEI message uses a batch size with a batch index of 0, but determining the batch size used as input for neural network inference depends on the post-processing implementation.
[0360] For example, when nnpfc_inp_order_idc is 3 and nnpfc_auxiliary_inp_idc is 1, the input tensor has a total of 7 channels, including 4 luminance matrices, 2 chroma matrices, and 1 auxiliary input matrix. In this case, the DeriveInputTensors() process derives the 7 channels of the input tensor one by one, and when a specific channel among these is processed, that channel is referred to as the current channel during the process.
[0361] nnpfc_inp_format_idc indicates how to convert the sample values of the input picture into input values for the NNPF. If nnpfc_inp_format_idc is 0, the input values for the NNPF are real numbers, and the InpY() and InpC() functions are expressed as follows.
[0362] InpY(x) = x χ ( ( 1 << BitDepthY ) - 1 )
[0363] InpC(x)=x χ ( ( 1 << BitDepthC ) - 1 )
[0364] When nnpfc_inp_format_idc is 1, the input value of NNPF is an unsigned integer, and the InpY() and InpC() functions are expressed as follows.
[0365] shiftY = BitDepthY - inpTensorBitDepthY
[0366] if( inpTensorBitDepthY >= BitDepthY)
[0367] InpY(x) = x << ( inpTensorBitDepthY - BitDepthY )
[0368] otherwise
[0369] InpY(x) = Clip3(0, (1 << inpTensorBitDepthY ) - 1, (x + (1 << (shiftY - 1 ) ) ) >> shiftY )
[0370] shiftC = BitDepthC - inpTensorBitDepthC
[0371] If inpTensorBitDepthC >= BitDepthC
[0372] InpC(x) = x << ( inpTensorBitDepthC - BitDepthC )
[0373] otherwise
[0374] InpC(x) = Clip3(0, (1 << inpTensorBitDepthC ) - 1, (x + (1 << (shiftC - 1 ) ) ) >> shiftC )
[0375] The variable inpTensorBitDepthY is derived from the syntax element nnpfc_inp_tensor_luma_bitdepth_minus8 specified below. The variable inpTensorBitDepthC is derived from the syntax element nnpfc_inp_tensor_chroma_bitdepth_minus8 specified below.
[0376] If the value of nnpfc_inp_format_idc is greater than 1, it is reserved for future specifications by ITU-T | ISO / IEC and does not exist in bitstreams compliant with this version of this document. Decoders compliant with this version of this document ignore NNPFC SEI messages containing the reserved value of nnpfc_inp_format_idc.
[0377] If the value of nnpfc_auxiliary_inp_idc is greater than 0, it indicates that there is auxiliary input data in the input tensor of NNPF. If nnpfc_auxiliary_inp_idc is 0, it indicates that there is no auxiliary input data in the input tensor. If nnpfc_auxiliary_inp_idc is 1, it indicates that auxiliary input data is derived as specified in the formula (inpTensorBitDepthY = nnpfc_inp_tensor_luma_bitdepth_minus8 + 8).
[0378] The value of nnpfc_auxiliary_inp_idc is in the range from 0 to 1 in bitstreams compliant with this version of this document. Values of nnpfc_auxiliary_inp_idc from 2 to 255 (inclusive) are reserved for future use by ITU-T | ISO / IEC and do not exist in bitstreams compliant with this version of this document. Decoders compliant with this version of this document ignore NNPFC SEI messages where nnpfc_auxiliary_inp_idc is in the range from 2 to 255. If the value of nnpfc_auxiliary_inp_idc exceeds 255, it does not exist in bitstreams compliant with this version of this document and is not reserved for future use.
[0379] nnpfc_inp_order_idc represents a method of forming an input tensor for NNPF by sorting an array of samples from the input picture.
[0380] The nnpfc_inp_order_idc value is in the range of 0 to 3 in bitstreams compliant with this version of this document. The range of nnpfc_inp_order_idc values from 4 to 255 is reserved for future use by ITU-T | ISO / IEC and does not exist in bitstreams compliant with this version of this document. Decoders compliant with this version of this document ignore NNPFC SEI messages where the nnpfc_inp_order_idc value is in the range of 4 to 255. If the nnpfc_inp_order_idc value exceeds 255, the value does not exist in bitstreams compliant with this version of this document and is not reserved for future use.
[0381] If ChromaFormatIdc is not 1, nnpfc_inp_order_idc is not 3.
[0382] If ChromaFormatIdc is 0, nnpfc_inp_order_idc is not 0.
[0383] If chromaUpsamplingFlag is 1, nnpfc_inp_order_idc is not 0.
[0384] Table 7 (Description of nnpfc_inp_order_idc values) provides a description of the nnpfc_inp_order_idc values.
[0385] [Table 7]
[0386]
[0387] FIG. 25 illustrates the process of inducing a luma channel in a luma component according to the embodiments.
[0388] Figure 25 is an example of deriving 4 luminance channels (right) from the luminance component when nnpfc_inp_order_idc is 3.
[0389] nnpfc_inp_tensor_luma_bitdepth_minus8 + 8 represents the bit depth of the luminance sample values in the input integer tensor. The value of inpTensorBitDepthY is derived as follows.
[0390] inpTensorBitDepthY = nnpfc_inp_tensor_luma_bitdepth_minus8 + 8 (85)
[0391] For bitstream compatibility, the value of nnpfc_inp_tensor_luma_bitdepth_minus8 is in the range of 0 to 24 (inclusive).
[0392] nnpfc_inp_tensor_chroma_bitdepth_minus8 + 8 represents the bit depth of the chroma sample values in the input integer tensor. The value of inpTensorBitDepthC is derived as follows.
[0393] inpTensorBitDepthC = nnpfc_inp_tensor_chroma_bitdepth_minus8 + 8
[0394] For bitstream compatibility, the value of nnpfc_inp_tensor_chroma_bitdepth_minus8 is in the range from 0 to 24.
[0395] When nnpfc_auxiliary_inp_idc is 1, the variable strengthControlScaledVal is derived as follows.
[0396] for( i = 0; i < numInputPics; i++ )
[0397] if(nnpfc_inp_format_idc = = 1)
[0398] if( nnpfc_inp_order_idc = = 0 | | nnpfc_inp_order_idc = = 2 | |
[0399] nnpfc_inp_order_idc = = 3 )
[0400] strengthControlScaledVal[ i ] =
[0401] Floor ( StrengthControlVal[ i ] * ( ( 1 << inpTensorBitDepthY ) - 1 ) )
[0402] else if(nnpfc_inp_order_idc = = 1)
[0403] strengthControlScaledVal[ i ] =
[0404] Floor ( StrengthControlVal[ i ] * ( ( 1 << inpTensorBitDepthC ) - 1 ) )
[0405] otherwise
[0406] strengthControlScaledVal[i] = StrengthControlVal[i]
[0407] A patch is a rectangular array of samples extracted from a component of a picture (e.g., a luma or chroma component).
[0408] The DeriveInputTensors() process, which derives the input tensor inputTensor for given vertical sample coordinates cTop and horizontal sample coordinates cLeft, represents the top-left sample position of the sample patch contained in the input tensor and is defined as follows.
[0409] for( i = 0; i < numInputPics; i++ ) {
[0410] if(nnpfc_inp_order_idc = = 0)
[0411] for( yP = -nnpfc_overlap; yP < inpPatchHeight + nnpfc_overlap; yP++)
[0412] for( xP = -nnpfc_overlap; xP < inpPatchWidth + nnpfc_overlap; xP++ ) {
[0413] inpVal = InpY( InpSampleVal( cTop + yP, cLeft + xP, CroppedHeight,
[0414] CroppedWidth, CroppedYPic[ i ], 0 ) )
[0415] yPovlp = yP + nnpfc_overlap
[0416] xPovlp = xP + nnpfc_overlap
[0417] if( !nnpfc_component_last_flag )
[0418] inputTensor
[0000] [ i ]
[0000] [ yPovlp ][ xPovlp ] = inpVal
[0419] else
[0420] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0000] = inpVal
[0421] if( nnpfc_auxiliary_inp_idc = = 1 )
[0422] if( !nnpfc_component_last_flag )
[0423] inputTensor
[0000] [ i ]
[0001] [ yPovlp ][ xPovlp ] = strengthControlScaledVal[ i ]
[0424] else
[0425] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0001] = strengthControlScaledVal[ i ]
[0426] }
[0427] else if( nnpfc_inp_order_idc = = 1 )
[0428] for( yP = -nnpfc_overlap; yP < inpPatchHeight + nnpfc_overlap; yP++)
[0429] for( xP = -nnpfc_overlap; xP < inpPatchWidth + nnpfc_overlap; xP++ ) {
[0430] inpCbVal = InpC( InpSampleVal( cTop + yP, cLeft + xP, CroppedHeight / SubHeightC,
[0431] CroppedWidth / SubWidthC, CroppedCbPic[ i ], 1 ) )
[0432] inpCrVal = InpC( InpSampleVal( cTop + yP, cLeft + xP, CroppedHeight / SubHeightC,
[0433] CroppedWidth / SubWidthC, CroppedCrPic[ i ], 2 ) )
[0434] yPovlp = yP + nnpfc_overlap
[0435] xPovlp = xP + nnpfc_overlap
[0436] if( !nnpfc_component_last_flag ) {
[0437] inputTensor
[0000] [ i ]
[0000] [ yPovlp ][ xPovlp ] = inpCbVal
[0438] inputTensor
[0000] [ i ]
[0001] [ yPovlp ][ xPovlp ] = inpCrVal
[0439] } else {
[0440] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0000] = inpCbVal
[0441] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0001] = inpCrVal
[0442] }
[0443] if( nnpfc_auxiliary_inp_idc = = 1 )
[0444] if( !nnpfc_component_last_flag )
[0445] inputTensor
[0000] [ i ]
[0002] [ yPovlp ][ xPovlp ] = strengthControlScaledVal[ i ]
[0446] else
[0447] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0002] = strengthControlScaledVal[ i ]
[0448] }
[0449] else if( nnpfc_inp_order_idc = = 2 )
[0450] for( yP = -nnpfc_overlap; yP < inpPatchHeight + nnpfc_overlap; yP++)
[0451] for( xP = -nnpfc_overlap; xP < inpPatchWidth + nnpfc_overlap; xP++ ) {
[0452] yY = cTop + yP
[0453] xY = cLeft + xP
[0454] yC = yY / SubHeightC
[0455] xC = xY / SubWidthC
[0456] inpYVal = InpY( InpSampleVal( yY, xY, CroppedHeight,
[0457] CroppedWidth, CroppedYPic[ i ], 0 ) )
[0458] inpCbVal = InpC( InpSampleVal( yC, xC, CroppedHeight / SubHeightC,
[0459] CroppedWidth / SubWidthC, CroppedCbPic[ i ], 1 ) )
[0460] inpCrVal = InpC( InpSampleVal( yC, xC, CroppedHeight / SubHeightC,
[0461] CroppedWidth / SubWidthC, CroppedCrPic[ i ], 2 ) )
[0462] yPovlp = yP + nnpfc_overlap
[0463] xPovlp = xP + nnpfc_overlap
[0464] if( !nnpfc_component_last_flag ) {
[0465] inputTensor
[0000] [ i ]
[0000] [ yPovlp ][ xPovlp ] = inpYVal
[0466] inputTensor
[0000] [ i ]
[0001] [ yPovlp ][ xPovlp ] = inpCbVal
[0467] inputTensor
[0000] [ i ]
[0002] [ yPovlp ][ xPovlp ] = inpCrVal
[0468] } else {
[0469] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0000] = inpYVal
[0470] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0001] = inpCbVal
[0471] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0002] = inpCrVal
[0472] }
[0473] if( nnpfc_auxiliary_inp_idc = = 1 )
[0474] if( !nnpfc_component_last_flag )
[0475] inputTensor
[0000] [ i ]
[0003] [ yPovlp ][ xPovlp ] = strengthControlScaledVal[ i ]
[0476] else
[0477] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0003] = strengthControlScaledVal[ i ]
[0478] }
[0479] else if( nnpfc_inp_order_idc = = 3 )
[0480] for( yP = -nnpfc_overlap; yP < inpPatchHeight + nnpfc_overlap; yP++)
[0481] for( xP = -nnpfc_overlap; xP < inpPatchWidth + nnpfc_overlap; xP++ ) {
[0482] yTL = cTop + yP * 2
[0483] xTL = cLeft + xP * 2
[0484] yBR = yTL + 1
[0485] xBR = xTL + 1
[0486] yC = cTop / 2 + yP
[0487] xC = cLeft / 2 + xP
[0488] inpTLVal = InpY( InpSampleVal( yTL, xTL, CroppedHeight,
[0489] CroppedWidth, CroppedYPic[ i ], 0 ) )
[0490] inpTRVal = InpY( InpSampleVal( yTL, xBR, CroppedHeight,
[0491] CroppedWidth, CroppedYPic[ i ], 0 ) )
[0492] inpBLVal = InpY( InpSampleVal( yBR, xTL, CroppedHeight,
[0493] CroppedWidth, CroppedYPic[ i ], 0 ) )
[0494] inpBRVal = InpY( InpSampleVal( yBR, xBR, CroppedHeight,
[0495] CroppedWidth, CroppedYPic[ i ], 0 ) )
[0496] inpCbVal = InpC( InpSampleVal( yC, xC, CroppedHeight / 2,
[0497] CroppedWidth / 2, CroppedCbPic[ i ], 1 ) )
[0498] inpCrVal = InpC( InpSampleVal( yC, xC, CroppedHeight / 2,
[0499] CroppedWidth / 2, CroppedCrPic[ i ], 2 ) )
[0500] yPovlp = yP + nnpfc_overlap
[0501] xPovlp = xP + nnpfc_overlap
[0502] if( !nnpfc_component_last_flag ) {
[0503] inputTensor
[0000] [ i ]
[0000] [ yPovlp ][ xPovlp ] = inpTLVal
[0504] inputTensor
[0000] [ i ]
[0001] [ yPovlp ][ xPovlp ] = inpTRVal
[0505] inputTensor
[0000] [ i ]
[0002] [ yPovlp ][ xPovlp ] = inpBLVal
[0506] inputTensor
[0000] [ i ]
[0003] [ yPovlp ][ xPovlp ] = inpBRVal
[0507] inputTensor
[0000] [ i ]
[0004] [ yPovlp ][ xPovlp ] = inpCbVal
[0508] inputTensor
[0000] [ i ]
[0005] [ yPovlp ][ xPovlp ] = inpCrVal
[0509] } else {
[0510] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0000] = inpTLVal
[0511] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0001] = inpTRVal
[0512] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0002] = inpBLVal
[0513] inputTensor
[0000] [i][yPovlp][xPovlp]
[0003] = inpBRVal
[0514] inputTensor
[0000] [i][yPovlp][xPovlp]
[0004] = inpCbVal
[0515] inputTensor
[0000] [i][yPovlp][xPovlp]
[0005] = inpCrVal
[0516] }
[0517] if(nnpfc_auxiliary_inp_idc = = 1)
[0518] if(!nnpfc_component_last_flag)
[0519] inputTensor
[0000] [i]
[0006] [yPovlp][xPovlp] = strengthControlScaledVal[i]
[0520] else
[0521] inputTensor
[0000] [i][yPovlp][xPovlp]
[0006] = strengthControlScaledVal[i]
[0522] }
[0523] }
[0524] If nnpfc_out_format_idc is 0, the sample values output by NNPF are real numbers, and the range of values from 0 to 1 is linearly mapped to the range of unsigned integer values from 0 to (1 << bitDepth) - 1 for the desired bit depth bitDepth for subsequent post-processing or display.
[0525] If nnpfc_out_format_idc is 1, it indicates that the luminance sample value output by NNPF is an unsigned integer from 0 to ( 1 << outTensorBitDepthY ) - 1, and the chroma sample value output by NNPF is an unsigned integer from 0 to ( 1 << outTensorBitDepthC ) - 1.
[0526] nnpfc_out_format_idc values greater than 1 are reserved for future specifications by ITU-T | ISO / IEC and should not be included in bitstreams compliant with this version of this document. Decoders compliant with this version of this document ignore NNPFC SEI messages containing reserved nnpfc_out_format_idc values.
[0527] nnpfc_out_order_idc indicates the output order of samples generated by NNPF.
[0528] The nnpfc_out_order_idc value is in the range of 0 to 3 in bitstreams compliant with this version of this document. Values of nnpfc_out_order_idc from 4 to 255 (inclusive) are reserved for future use by ITU-T | ISO / IEC and do not exist in bitstreams compliant with this version of this document. Decoders compliant with this version of this document ignore NNPFC SEI messages where nnpfc_out_order_idc is in the range of 4 to 255. nnpfc_out_order_idc values greater than 255 do not exist in bitstreams compliant with this version of this document and are not reserved for future use.
[0529] If chromaUpsamplingFlag is 1, nnpfc_out_order_idc cannot be 0 or 3.
[0530] If colourizationFlag is 1, nnpfc_out_order_idc cannot be 0.
[0531] Table 8 (Description of nnpfc_out_order_idc values) provides a description of the nnpfc_out_order_idc values.
[0532] [Table 8]
[0533]
[0534] nnpfc_out_tensor_luma_bitdepth_minus8 + 8 represents the bit depth of the lumina sample values in the output integer tensor. The value of nnpfc_out_tensor_luma_bitdepth_minus8 ranges from 0 to 24. The value of outTensorBitDepthY is derived as follows.
[0535] outTensorBitDepthY = nnpfc_out_tensor_luma_bitdepth_minus8 + 8
[0536] nnpfc_out_tensor_chroma_bitdepth_minus8 + 8 represents the bit depth of the chroma sample values in the output integer tensor. The value of nnpfc_out_tensor_chroma_bitdepth_minus8 ranges from 0 to 24. The value of outTensorBitDepthC is derived as follows.
[0537] outTensorBitDepthC = nnpfc_out_tensor_chroma_bitdepth_minus8 + 8
[0538] If bitDepthUpsamplingFlag is 1, the value of nnpfc_out_format_idc must be 1 and satisfy one or more of the following conditions.
[0539] nnpfc_out_tensor_luma_bitdepth_minus8 exists and outTensorBitDepthY is greater than BitDepthY.
[0540] nnpfc_out_tensor_chroma_bitdepth_minus8 exists and outTensorBitDepthC is greater than BitDepthC.
[0541] If nnpfc_inp_tensor_luma_bitdepth_minus8, nnpfc_inp_tensor_chroma_bitdepth_minus8, nnpfc_out_tensor_luma_bitdepth_minus8, and nnpfc_out_tensor_chroma_bitdepth_minus8 exist and outTensorBitDepthY is greater than inpTensorBitDepthY, then outTensorBitDepthC cannot be less than inpTensorBitDepthC. If nnpfc_inp_tensor_luma_bitdepth_minus8, nnpfc_inp_tensor_chroma_bitdepth_minus8, nnpfc_out_tensor_luma_bitdepth_minus8, and nnpfc_out_tensor_chroma_bitdepth_minus8 exist and outTensorBitDepthC is greater than inpTensorBitDepthC, then outTensorBitDepthY cannot be less than inpTensorBitDepthY.
[0542] The StoreOutputTensors() process, which derives sample values of the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor outputTensor for a given vertical sample coordinate cTop and a horizontal sample coordinate cLeft specifying the top-left sample position of the sample patch included in the input tensor, is expressed as follows.
[0543] for( i = 0; i < numOutputPics; i++ ) {
[0544] if( nnpfc_out_order_idc = = 0 )
[0545] for( yP = 0; yP < outPatchHeight; yP++)
[0546] for( xP = 0; xP < outPatchWidth; xP++ ) {
[0547] yY = cTop * outPatchHeight / inpPatchHeight + yP
[0548] xY = cLeft * outPatchWidth / inpPatchWidth + xP
[0549] if ( yY < nnpfcOutputPicHeight && xY < nnpfcOutputPicWidth )
[0550] if( !nnpfc_component_last_flag )
[0551] FilteredYPic[ i ][ xY ][yY ] = outputTensor
[0000] [ i ]
[0000] [ yP ][ xP ]
[0552] else
[0553] FilteredYPic[ i ][ xY ][ yY ] = outputTensor
[0000] [ i ][ yP ][ xP ]
[0000] }
[0554] else if( nnpfc_out_order_idc = = 1 ) (91)
[0555] for( yP = 0; yP < outPatchCHeight; yP++)
[0556] for( xP = 0; xP < outPatchCWidth; xP++ ) {
[0557] xSrc = cLeft * horCScaling + xP
[0558] ySrc = cTop * verCScaling + yP
[0559] if ( ySrc < nnpfcOutputPicHeight / outSubHeightC &&
[0560] xSrc < nnpfcOutputPicWidth / outSubWidthC )
[0561] if( !nnpfc_component_last_flag ) {
[0562] FilteredCbPic[ i ][ xSrc ][ ySrc ] = outputTensor
[0000] [ i ]
[0000] [ yP ][ xP ]
[0563] FilteredCrPic[ i ][ xSrc ][ ySrc ] = outputTensor
[0000] [ i ]
[0001] [ yP ][ xP ]
[0564] } else {
[0565] FilteredCbPic[ i ][ xSrc ][ ySrc ] = outputTensor
[0000] [ i ][ yP ][ xP ]
[0000]
[0566] FilteredCrPic[ i ][ xSrc ][ ySrc ] = outputTensor
[0000] [ i ][ yP ][ xP ]
[0001]
[0567] }
[0568] }
[0569] else if( nnpfc_out_order_idc = = 2 )
[0570] for( yP = 0; yP < outPatchHeight; yP++)
[0571] for( xP = 0; xP < outPatchWidth; xP++ ) {
[0572] yY = cTop * outPatchHeight / inpPatchHeight + yP
[0573] xY = cLeft * outPatchWidth / inpPatchWidth + xP
[0574] yC = yY / outSubHeightC
[0575] xC = xY / outSubWidthC
[0576] yPc = ( yP / outSubHeightC ) * outSubHeightC
[0577] xPc = ( xP / outSubWidthC ) * outSubWidthC
[0578] if ( yY < nnpfcOutputPicHeight && xY < nnpfcOutputPicWidth )
[0579] if( !nnpfc_component_last_flag ) {
[0580] FilteredYPic[ i ][ xY ][ yY ] = outputTensor
[0000] [ i ]
[0000] [ yP ][ xP ]
[0581] FilteredCbPic[ i ][ xC ][ yC ] = outputTensor
[0000] [ i ]
[0001] [ yPc ][ xPc ]
[0582] FilteredCrPic[ i ][ xC ][ yC ] = outputTensor
[0000] [ i ]
[0002] [ yPc ][ xPc ]
[0583] } else {
[0584] FilteredYPic[ i ][ xY ][ yY ] = outputTensor
[0000] [ i ][ yP ][ xP ]
[0000]
[0585] FilteredCbPic[ i ][ xC ][ yC ] = outputTensor
[0000] [ i ][ yPc ][ xPc ]
[0001]
[0586] FilteredCrPic[ i ][ xC ][ yC ] = outputTensor
[0000] [ i ][ yPc ][ xPc ]
[0002]
[0587] }
[0588] }
[0589] else if( nnpfc_out_order_idc = = 3 )
[0590] for( yP = 0; yP < outPatchHeight; yP++ )
[0591] for( xP = 0; xP < outPatchWidth; xP++ ) {
[0592] ySrc = cTop / 2 * outPatchHeight / inpPatchHeight + yP
[0593] xSrc = cLeft / 2 * outPatchWidth / inpPatchWidth + xP
[0594] if ( ySrc < nnpfcOutputPicHeight / 2 &&
[0595] xSrc < nnpfcOutputPicWidth / 2 )
[0596] if( !nnpfc_component_last_flag ) {
[0597] FilteredYPic[ i ][ xSrc * 2 ][ ySrc * 2 ] = outputTensor
[0000] [ i ]
[0000] [ yP ][ xP ]
[0598] FilteredYPic[ i ][ xSrc * 2 + 1 ][ ySrc * 2 ] = outputTensor
[0000] [ i ]
[0001] [ yP ][ xP ]
[0599] FilteredYPic[ i ][ xSrc * 2 ][ ySrc * 2 + 1 ] = outputTensor
[0000] [ i ]
[0002] [ yP ][ xP ]
[0600] FilteredYPic[ i ][ xSrc * 2 + 1][ ySrc * 2 + 1 ] = outputTensor
[0000] [ i ]
[0003] [ yP ][ xP ]
[0601] FilteredCbPic[ i ][ xSrc ][ ySrc ] = outputTensor
[0000] [ i ]
[0004] [ yP ][ xP ]
[0602] FilteredCrPic[ i ][ xSrc ][ ySrc ] = outputTensor
[0000] [ i ]
[0005] [ yP ][ xP ]
[0603] } else {
[0604] FilteredYPic[ i ][ xSrc * 2 ][ ySrc * 2 ] = outputTensor
[0000] [ i ][ yP ][ xP ]
[0000]
[0605] FilteredYPic[ i ][ xSrc * 2 + 1 ][ ySrc * 2 ] = outputTensor
[0000] [ i ][ yP ][ xP ]
[0001]
[0606] FilteredYPic[ i ][ xSrc * 2 ][ ySrc * 2 + 1 ] = outputTensor
[0000] [ i ][ yP ][ xP ]
[0002]
[0607] FilteredYPic[ i ][ xSrc * 2 + 1][ ySrc * 2 + 1 ] = outputTensor
[0000] [ i ][ yP ][ xP ]
[0003]
[0608] FilteredCbPic[i][xSrc][ySrc] = outputTensor
[0000] [i][yP][xP]
[0004]
[0609] FilteredCrPic[i][xSrc][ySrc] = outputTensor
[0000] [i][yP][xP]
[0005]
[0610] }
[0611] }
[0612] }
[0613] If nnpfc_separate_colour_description_present_flag is 1, it indicates that a unique combination of color primary, transfer property, matrix factor, scaling, and offset values applied in relation to the matrix factor for the picture generated by NNPF is specified in the SEI message syntax structure. If nnpfc_separate_colour_description_present_flag is 0, it indicates that the combination of color primary, transfer property, matrix factor, scaling, and offset values applied in relation to the matrix factor for the picture generated by NNPF is the same as that specified in the VUI parameters of CLVS.
[0614] nnpfc_colour_primaries has the same meaning as the vui_colour_primaries syntax element, but with the following differences.
[0615] nnpfc_colour_primaries represents the color primary colors of the picture generated by applying the NNPF specified in the SEI message, rather than the color primary colors used in CLVS.
[0616] If nnpfc_colour_primaries is not present in the NNPFC SEI message, the value of nnpfc_colour_primaries is inferred to be the same as vui_colour_primaries.
[0617] nnpfc_transfer_characteristics has the same meaning as specified for the vui_transfer_characteristics syntax element, except for the following.
[0618] nnpfc_transfer_characteristics represents the transfer characteristics of the picture generated by applying the NNPF specified in the SEI message, rather than the transfer characteristics used in CLVS.
[0619] If nnpfc_transfer_characteristics is not present in the NNPFC SEI message, the value of nnpfc_transfer_characteristics is inferred to be the same as vui_transfer_characteristics.
[0620] nnpfc_matrix_coeffs describes the equations used to derive luminance and saturation signals from green, blue, red, or the primary colors Y, Z, and X. The meaning of this function applies to the picture generated by applying the NNPF specified in this SEI message, as specified in the MatrixCoefficients of Rec. ITU-T H.273 | ISO / IEC 23091-2, where BitDepthY and BitDepthC are equal to outTensorBitDepthY and outTensorBitDepthC, respectively.
[0621] If nnpfc_matrix_coeffs is not in the NNPFC SEI message, the value of nnpfc_matrix_coeffs is inferred to be the same as vui_matrix_coeffs.
[0622] nnpfc_matrix_coeffs cannot be 0 except when both of the following two conditions are true.
[0623] nnpfc_out_tensor_chroma_bitdepth_minus8 is equal to nnpfc_out_tensor_luma_bitdepth_minus8.
[0624] nnpfc_out_order_idc is 2, outSubHeightC is 1, and outSubWidthC is 1.
[0625] nnpfc_matrix_coeffs cannot be 8 unless one of the following conditions is true.
[0626] nnpfc_out_tensor_chroma_bitdepth_minus8 is equal to nnpfc_out_tensor_luma_bitdepth_minus8.
[0627] nnpfc_out_tensor_chroma_bitdepth_minus8 is equal to nnpfc_out_tensor_luma_bitdepth_minus8 + 1, nnpfc_out_order_idc is equal to 2, outSubHeightC is equal to 1, and outSubWidthC is equal to 1.
[0628] nnpfc_full_range_flag represents the scaling and offset values applied in relation to the matrix coefficients specified in nnpfc_matrix_coeffs. The meaning of this value is the same as that specified in the VideoFullRangeFlag parameter of Rec. ITU-T H.273 | ISO / IEC 23091-2. If nnpfc_full_range_flag is not present, the value is inferred to be 0.
[0629] If the value of nnpfc_chroma_loc_info_present_flag is 1, it indicates that the nnpfc_chroma_sample_loc_type_frame syntax element is present in the NNPFC SEI message. If the value of nnpfc_chroma_loc_info_present_flag is 0, it indicates that the nnpfc_chroma_sample_loc_type_frame syntax element is not present in the NNPFC SEI message. If colourizationFlag is 0 or nnpfc_out_colour_format_idc is not 1, the value of nnpfc_chroma_loc_info_present_flag is equal to 0.
[0630] If nnpfc_chroma_sample_loc_type_frame is not 6 and nnpfc_out_colour_format_idc is 1, it indicates the chroma sample location of the output picture. If nnpfc_chroma_sample_loc_type_frame is 6 and nnpfc_out_colour_format_idc is 1, it indicates that the chroma sample location is unknown, unspecified, or specified in another way not specified in this document. The value of nnpfc_chroma_sample_loc_type_frame ranges from 0 to 6.
[0631] nnpfc_overlap indicates the overlap in the number of horizontal and vertical samples of adjacent input tensors in NNPF. The nnpfc_overlap value ranges from 0 to 16,383.
[0632] If nnpfc_constant_patch_size_flag is 1, it indicates that NNPF exactly accepts the patch sizes specified in nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1 as input. If nnpfc_constant_patch_size_flag is 0, it indicates that NNPF accepts any patch size as input with a width of inpPatchWidth and a height of inpPatchHeight, wherein the width of the extended patch (i.e., the area overlapping with the patch) is equal to inpPatchWidth + 2 * nnpfc_overlap and this width is a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 * nnpfc_overlap, and the height of the extended patch is equal to inpPatchHeight + 2 * nnpfc_overlap and this width is a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 * nnpfc_overlap.
[0633] The value of nnpfc_patch_width_minus1 plus 1 represents the horizontal sample size of the patch size required for the NNPF input when nnpfc_constant_patch_size_flag is 1. The value of nnpfc_patch_width_minus1 ranges from 0 to Min(32,766, CroppedWidth - 1).
[0634] The value of nnpfc_patch_height_minus1 plus 1 represents the number of vertical samples of the patch size required for the NNPF input when nnpfc_constant_patch_size_flag is 1. The value of nnpfc_patch_height_minus1 ranges from 0 to Min(32,766, CroppedHeight - 1).
[0635] nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 * nnpfc_overlap represents the common divisor of all allowed values for the extended patch width required for the NNPF input when nnpfc_constant_patch_size_flag is 0. The value of nnpfc_extended_patch_width_cd_delta_minus1 is in the range from 0 to Min(32,766, CroppedWidth - 1).
[0636] nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 * nnpfc_overlap represents the common divisor of all allowable values of the extended patch height required for the NNPF input when nnpfc_constant_patch_size_flag is 0. The value of nnpfc_extended_patch_height_cd_delta_minus1 is in the range from 0 to Min(32,766, CroppedHeight - 1).
[0637] Set the inpPatchWidth and inpPatchHeight variables to the patch size width and patch size height, respectively.
[0638] If nnpfc_constant_patch_size_flag is 0, the following applies.
[0639] The inpPatchWidth and inpPatchHeight values are provided through external means not specified in this document or are set by the postprocessor itself.
[0640] The value of inpPatchWidth + 2 * nnpfc_overlap must be a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 * nnpfc_overlap, and inpPatchWidth is less than or equal to CroppedWidth. The value of inpPatchHeight + 2 * nnpfc_overlap must be a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 * nnpfc_overlap, and inpPatchHeight is less than or equal to CroppedHeight.
[0641] Otherwise (when nnpfc_constant_patch_size_flag is 1), the inpPatchWidth value is set to nnpfc_patch_width_minus1 + 1 and the inpPatchHeight value is set to nnpfc_patch_height_minus1 + 1.
[0642] The variables outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, and outPatchCHeight are derived as follows.
[0643] outPatchWidth = (nnpfcOutputPicWidth * inpPatchWidth) / CroppedWidth
[0644] outPatchHeight = (nnpfcOutputPicHeight * inpPatchHeight) / CroppedHeight
[0645] horCScaling = SubWidthC / outSubWidthC
[0646] verCScaling = SubHeightC / outSubHeightC
[0647] outPatchCWidth = outPatchWidth * horCScaling
[0648] outPatchCHeight = outPatchHeight * verCScaling
[0649] For bitstream conformance, outPatchWidth * CroppedWidth is equal to nnpfcOutputPicWidth * inpPatchWidth, and outPatchHeight * CroppedHeight is equal to nnpfcOutputPicHeight * inpPatchHeight.
[0650] nnpfc_padding_type represents the padding process when referring to sample locations outside the input picture boundaries, as described in Table 9 (Informative description of nnpfc_padding_type values). The values of nnpfc_padding_type range from 0 to 4 in bitstreams compliant with this version of this document. The range of nnpfc_padding_type values from 5 to 15 is reserved for future use by ITU-T | ISO / IEC and does not exist in bitstreams compliant with this version of this document. Decoders compliant with this revision of this document ignore NNPFC SEI messages where the nnpfc_padding_type value is between 5 and 15 (inclusive). nnpfc_padding_type values greater than 15 do not exist in bitstreams compliant with this revision of this document and are not reserved for future use.
[0651] [Table 9]
[0652]
[0653] nnpfc_luma_padding_val represents the lumina value to be used for padding when nnpfc_padding_type is 4. The value of nnpfc_luma_padding_val is in the range from 0 to (1 << BitDepthY) - 1.
[0654] nnpfc_cb_padding_val represents the Cb value to be used for padding when nnpfc_padding_type is 4. The value of nnpfc_cb_padding_val is in the range from 0 to (1 << BitDepthC) - 1.
[0655] nnpfc_cr_padding_val represents the Cr value to be used for padding when nnpfc_padding_type is 4. The nnpfc_cr_padding_val value is in the range of 0 to (1 << BitDepthC) - 1.
[0656] The InpSampleVal(y, x, picHeight, picWidth, croppedPic, cIdx) function takes the vertical sample position y, horizontal sample position x, picture height picHeight, picture width picWidth, sample array croppedPic, and component index cIdx (0 for Luma, 1 for Cb, 2 for Cr) as input and returns the sampleVal value derived as follows.
[0657] For the input to the InpSampleVal() function, vertical positions are listed before horizontal positions for compatibility with the input tensor rules of some inference engines.
[0658] if(nnpfc_padding_type = = 0)
[0659] if( y < 0 | | x < 0 | | y >= picHeight | | x >= picWidth )
[0660] sampleVal = 0
[0661] else
[0662] sampleVal = croppedPic[ x ][ y ] (98)
[0663] else if( nnpfc_padding_type = = 1 )
[0664] sampleVal = croppedPic[ Clip3( 0, picWidth - 1, x ) ][ Clip3( 0, picHeight - 1, y ) ]
[0665] else if( nnpfc_padding_type = = 2 )
[0666] sampleVal = croppedPic[ Reflect( picWidth - 1, x ) ][ Reflect( picHeight - 1, y ) ]
[0667] else if( nnpfc_padding_type = = 3 )
[0668] if( y >= 0 && y < picHeight )
[0669] sampleVal = croppedPic[ Wrap( picWidth - 1, x ) ][ y ]
[0670] else if( nnpfc_padding_type = = 4 )
[0671] if( y < 0 | | x < 0 | | y >= picHeight | | x >= picWidth )
[0672] sampleVal = ( cIdx = = 0 ? nnpfc_luma_padding_val :
[0673] ( cIdx = = 1 ? nnpfc_cb_padding_val : nnpfc_cr_padding_val ) )
[0674] else
[0675] sampleVal = croppedPic[ x ][ y ]
[0676] NNPF PostProcessingFilter() is the target NNPF derived from the semantics of the NNPFA SEI message. The following example process can be used with NNPF PostProcessingFilter() to generate a patch-filtered and / or interpolated image. This image contains the Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, respectively, as indicated in nnpfc_out_order_idc.
[0677] if( nnpfc_inp_order_idc = = 0 | | nnpfc_inp_order_idc = = 2 )
[0678] for( cTop = 0; cTop < CroppedHeight; cTop += inpPatchHeight )
[0679] for( cLeft = 0; cLeft < CroppedWidth; cLeft += inpPatchWidth ) {
[0680] DeriveInputTensors()
[0681] outputTensor = PostProcessingFilter( inputTensor )
[0682] StoreOutputTensors()
[0683] }
[0684] else if(nnpfc_inp_order_idc = = 1)
[0685] for( cTop = 0; cTop < CroppedHeight / SubHeightC; cTop += inpPatchHeight )
[0686] for( cLeft = 0; cLeft < CroppedWidth / SubWidthC; cLeft += inpPatchWidth ) {
[0687] DeriveInputTensors()
[0688] outputTensor = PostProcessingFilter( inputTensor )
[0689] StoreOutputTensors()
[0690] }
[0691] else if(nnpfc_inp_order_idc = = 3)
[0692] for( cTop = 0; cTop < CroppedHeight; cTop += inpPatchHeight * 2 )
[0693] for( cLeft = 0; cLeft < CroppedWidth; cLeft += inpPatchWidth * 2 ) {
[0694] DeriveInputTensors()
[0695] outputTensor = PostProcessingFilter( inputTensor )
[0696] StoreOutputTensors()
[0697] }
[0698] If present, the NNPF generated image with index i includes the sample arrays FilteredYPic[ i ], FilteredCbPic[ i ], and FilteredCrPic[ i ] derived by the above formula (cLeft = 0; cLeft < CroppedWidth / SubWidthC; cLeft += inpPatchWidth). The NNPF generated image does not contain overlapping regions.
[0699] The NNPF process consists of the process defined in the above formula (cLeft = 0; cLeft < CroppedWidth / SubWidthC; cLeft += inpPatchWidth) and the subsequent process of outputting the NNPF-generated images in index order. Here, all NNPF-generated images interpolated by the NNPF are output, and the NNPF-generated images corresponding to the input images for the NNPF are output as specified in the semantics of the NNPFA SEI message.
[0700] If nnpfc_complexity_info_present_flag is 1, it indicates that there is one or more syntax elements representing the complexity of the NNPF associated with nnpfc_id. If nnpfc_complexity_info_present_flag is 0, it specifies that there are no syntax elements representing the complexity of the NNPF associated with nnpfc_id.
[0701] If nnpfc_parameter_type_idc is 0, it indicates that the neural network uses only integer parameters. If nnpfc_parameter_type_flag is 1, it indicates that the neural network can use floating-point or integer parameters. If nnpfc_parameter_type_idc is 2, it indicates that the neural network uses only binary parameters. If nnpfc_parameter_type_idc is 3, it is reserved for future use by ITU-T | ISO / IEC and is not included in bitstreams compliant with this version of this document. Decoders compliant with this version of this document ignore NNPFC SEI messages where nnpfc_parameter_type_idc is 3.
[0702] If nnpfc_log2_parameter_bit_length_minus3 is 0, 1, 2, or 3, it indicates that the neural network does not use parameters with bit lengths greater than 8, 16, 32, and 64, respectively. If nnpfc_parameter_type_idc is present and nnpfc_log2_parameter_bit_length_minus3 is absent, the neural network does not use parameters with a bit length greater than 1.
[0703] nnpfc_num_parameters_idc represents the maximum number of neural network parameters for the NNPF in powers of 2048. If nnpfc_num_parameters_idc is 0, it indicates that the maximum number of neural network parameters is unknown. The value of nnpfc_num_parameters_idc ranges from 0 to 52 (inclusive). nnpfc_num_parameters_idc values greater than 52 are reserved for future use by ITU-T | ISO / IEC and do not exist in bitstreams compliant with this version of this document. Decoders compliant with this version of this document ignore NNPFC SEI messages where nnpfc_num_parameters_idc is greater than 52.
[0704] If the value of nnpfc_num_parameters_idc is greater than 0, the variable maxNumParameters is derived as follows.
[0705] maxNumParameters = ( 2 048 << nnpfc_num_parameters_idc ) - 1
[0706] The requirement for bitstream conformance is that the number of neural network parameters of NNPF must be less than or equal to maxNumParameters.
[0707] If nnpfc_num_kmac_operations_idc is greater than 0, it indicates that the maximum number of multiplicative-accumulator operations per NNPF sample is less than or equal to nnpfc_num_kmac_operations_idc * 1,000. If nnpfc_num_kmac_operations_idc is 0, it indicates that the maximum number of multiplicative-accumulator operations of the network is unknown. The value of nnpfc_num_kmac_operations_idc ranges from 0 to 2 32 It is in the range of -2(2).
[0708] If nnpfc_total_kilobyte_size is greater than 0, it indicates the total size (KB) required to store the neural network's uncompressed parameters. The total size (in bits) is a number greater than or equal to the sum of the bits used to store each parameter. nnpfc_total_kilobyte_size is the total size (in bits) divided by 8,000 and rounded. If nnpfc_total_kilobyte_size is 0, it indicates that the total size required to store the neural network's parameters is unknown. The value of nnpfc_total_kilobyte_size ranges from 0 to 2 32 It is in the range of -2(2).
[0709] If nnpfc_metadata_extension_num_bits is 0, it indicates that there is no nnpfc_reserved_metadata_extension. If nnpfc_metadata_extension_num_bits is greater than 0, it indicates the length (in bits) of the nnpfc_reserved_metadata_extension. In this version of this document, nnpfc_metadata_extension_num_bits is 0. Values for nnpfc_metadata_extension_num_bits in the range of 1 to 2,048 (inclusive) are reserved for future use by ITU-T | ISO / IEC and do not exist in bitstreams compliant with this version of this document. Decoders compliant with this version of this document accept all nnpfc_metadata_extension_num_bits values in the range of 0 to 2,048 (inclusive). If the value of nnpfc_metadata_extension_num_bits is greater than 2,048, it does not exist in the bitstream following this version of this document and is not reserved for future use.
[0710] The nnpfc_reserved_metadata_extension value does not exist in bitstreams following this version of this document. However, decoders following this version of this document ignore the existence and value of nnpfc_reserved_metadata_extension. If nnpfc_reserved_metadata_extension exists, the length (in bits) of nnpfc_metadata_extension is equal to nnpfc_metadata_extension_num_bits.
[0711] nnpfc_reserved_zero_bit_b is equal to 0 in bitstreams following this version of this document. Decoders ignore NNPFC SEI messages where nnpfc_reserved_zero_bit_b is not 0.
[0712] nnpfc_payload_byte[ i ] contains the i-th byte of a bitstream compliant with ISO / IEC 15938-17. The byte sequence for all current values of i, nnpfc_payload_byte[ i ], is a complete bitstream compliant with ISO / IEC 15938-17.
[0713] FIG. 26 shows the syntax of a neural network post-filter activation SEI (Supplemental enhancement information) message according to embodiments.
[0714] Referring to FIG. 26, the neural network post-filter activation (NNPFA) SEI message semantics are described.
[0715] The NNPFA SEI message enables or disables post-processing filtering for a series of pictures using a target neural network post-processing filter (NNPF) identified by nnpfa_target_id and nnpfa_target_base_flag. For a specific picture with the NNPF enabled, the target NNPF is derived as follows.
[0716] If nnpfa_target_base_flag is 1, the target NNPF is a base NNPF where nnpfc_id and nnpfa_target_id are the same.
[0717] Otherwise (when nnpfa_target_base_flag is 0), the target NNPF is the NNPF specified in the last NNPFC SEI message with nnpfc_id equal to nnpfa_target_id, which precedes the first VCL NAL unit of the current picture in the decoding order, and is not a repetition of the NNPFC SEI message containing the base NNPF.
[0718] Multiple NNPFA SEI messages may exist for the same picture. For example, this occurs when NNPF is used for different purposes or for filtering different color components.
[0719] nnpfa_target_id represents a target NNPF specified by one or more NNPFC SEI messages that are associated with the current picture and whose nnpfc_id is equal to nnpfa_target_id. The value of nnpfa_target_id is from 0 to 2 32 It is in the range up to -2.
[0720] An NNPFA SEI message with a specific value of nnpfa_target_id does not exist in the current PU unless one or both of the following conditions are met.
[0721] There is an NNPFC SEI message in the current CLVS that is ahead of the current PU in decoding order, where nnpfc_id is the same as a specific value of nnpfa_target_id.
[0722] Currently, there is an NNPFC SEI message in the PU where nnpfc_id is the same as a specific value of nnpfa_target_id.
[0723] If PU contains both an NNPFC SEI message with a specific value of nnpfc_id and an NNPFA SEI message where nnpfa_target_id is the same as a specific value of nnpfc_id, the NNPFC SEI message comes before the NNPFA SEI message in the decoding order.
[0724] If nnpfa_cancel_flag is 1, it indicates that the persistence of the target NNPF set by the previous NNPFA SEI message with the same nnpfa_target_id as the current SEI message is canceled. That is, the target NNPF is no longer used unless activated by another NNPFA SEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag of 0. If nnpfa_cancel_flag is 0, it indicates that nnpfa_target_base_flag, nnpfa_persistence_flag, and nnpfa_num_output_entries follow.
[0725] If nnpfa_target_base_flag is 1, it indicates that the target NNPF is a base NNPF where nnpfc_id and nnpfa_target_id are the same. If nnpfa_target_base_flag is 0, it indicates that the target NNPF is an NNPF where the nnpfc_id of the last NNPFC SEI message preceding the first VCL NAL unit of the current picture in the decoding order is the same as nnpfa_target_id, and it is not a repetition of the NNPFC SEI message containing the base NNPF.
[0726] nnpfa_persistence_flag indicates the persistence of the target NNPF for the current layer.
[0727] If nnpfa_persistence_flag is 0, it indicates that the target NNPF can only be used for post-processing filtering on the current picture.
[0728] If nnpfa_persistence_flag is 1, it indicates that the target NNPF can be used for post-processing filtering on the current picture and all subsequent pictures of the current layer until one or more of the following conditions are met.
[0729] A new CLVS of the current layer starts. The bitstream ends. The picture of the current layer associated with the NNPFA SEI message, which has the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag as 1, is output after the current picture in output order.
[0730] The target NNPF is not applied to subsequent pictures of the current layer associated with an NNPFA SEI message that has the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag is 1.
[0731] nnpfcTargetPictures is defined as the set of pictures associated with the last NNPFC SEI message, which precedes the current NNPFA SEI message in the decoding order and has an nnpfc_id equal to nnpfa_target_id. nnpfaTargetPictures is defined as the set of pictures for which the target NNPF is activated by the current NNPFA SEI message. For bitstream conformance, all pictures included in nnpfaTargetPictures must also be included in nnpfcTargetPictures.
[0732] nnpfa_num_output_entries indicates the number of nnpfa_output_flag[ i ] syntax elements in the NNPFA SEI message. The value of nnpfa_num_output_entries ranges from 0 to NumInpPicsInOutputTensor (inclusive).
[0733] If nnpfa_output_flag[ i ] is 1, it indicates that the NNPF generated picture corresponding to the input picture with index InpIdx[ i ] is output by the NNPF process activated by this NNPFA SEI message, where the NNPF process is specified in the semantics of the NNPFC SEI message. If nnpfa_output_flag[ i ] is 0, it indicates that the NNPF generated picture corresponding to the input picture with index InpIdx[ i ] is not output by the NNPF process activated by this NNPFA SEI message. If nnpfa_num_output_entries is less than NumInpPicsInOutputTensor, nnpfa_output_flag[ i ] is inferred to be equal to 1 for each i value in nnpfa_num_output_entries within the range NumInpPicsInOutputTensor - 1 (inclusive).
[0734] FIG. 27 shows the syntax of a neural network post-pillar group characteristic SEI message according to embodiments.
[0735] Referring to Fig. 27, the semantics of the Neural-network post-filter group characteristics (NNPFGC) SEI message are explained.
[0736] The NNPFGC SEI message represents a neural network post-filter (NNPF) group. If the NNPF group defines an NNPF cascade or defines NNPF groups of NNPF or NNPF cascades that substitute for each other, it is indicated in the SEI message. If an NNPF group of an NNPF cascade is used for a specific picture, it is indicated via the neural network post-filter group activation (NNPFGA) SEI message.
[0737] nnpfgc_id contains an identification number that can be used to identify an NNPF group. The nnpfgc_id value is 0 to 232 It is in the range of -2 (inclusive). The values of nnpfgc_id are 256 to 511 (inclusive) and 231 to 2 32 -2 (inclusive) is reserved for future use by ITU-T | ISO / IEC. Decoders complying with this version of this document must have an nnpfgc_id in the range of 256 to 511 (inclusive) or 2 31 ~2 32 If an NNPFGC SEI message is found in the range of -2 (inclusive), that SEI message is ignored. The nnpfgc_id value must not be the same as the nnpfgc_id value of an NNPFGC SEI message in the same CLVS. If the nnpfgc_id value of NNPFGC SEI message nnpfgcSeiA is the same as the nnpfgc_id value of another NNPFGC SEI message nnpfgcSeiB in the same CLVS, then nnpfgcSeiA and nnpfgcSeiB are identical.
[0738] If nnpfgc_grouping_type is 0, this SEI message specifies a stepped neural network post-processing filter group.
[0739] If nnpfgc_grouping_type is 1, it indicates that the NNPF or NNPF groups identified by nnpfgc_member_id[ i ] are interchangeable, and the postprocessor must select and apply only one of them.
[0740] If nnpfgc_grouping_type is 2, it specifies the NNPF groups intended for this SEI message to be used jointly, and indicates that they are alternately activated so that at most only one of these NNPFs is activated for all pictures.
[0741] If nnpfgc_grouping_type is 3, it indicates that the NNPF or NNPF group identified by nnpfgc_member_id[ i ] is intended to be used in parallel.
[0742] If nnpfgc_grouping_type is 4, it indicates that the NNPF or NNPF group identified by nnpfgc_member_id[ i ] is optional. That is, it may or may not be applied in the postprocessor.
[0743] The nnpfgc_grouping_type value is in the range from 0 to 255. The nnpfgc_grouping_type value in the range from 5 to 255 is reserved for future specification by ITU-T | ISO / IEC and does not exist in bitstreams compliant with this version of this document. Decoders compliant with this version of this document ignore NNPFGC SEI messages where nnpfgc_grouping_type is in the range from 5 to 255.
[0744] nnpfgc_purpose has the same meaning as nnpfc_purpose, but differs in that it specifies the meaning for the NNPF group defined in this SEI message, rather than the NNPF defined in the NNPFC SEI message.
[0745] nnpfgc_num_members_minus2 + 2 represents the number of NNPF or NNPF groups within the NNPF group defined by this SEI message.
[0746] nnpfgc_member_id[ i ] represents the i-th member of the NNPF group defined by this SEI message as follows.
[0747] If there is an NNPF with nnpfc_id equal to nnpfgc_member_id[ i ] defined in CLVS, the i-th member of the NNPF group defined by this SEI message is an NNPF with nnpfc_id equal to nnpfgc_member_id[ i ].
[0748] - Otherwise (if there is no NNPF defined in CLVS with nnpfc_id equal to nnpfgc_member_id[ i ]), the i-th member of the NNPF group defined by this SEI message is the NNPF group with nnpfgc_id equal to nnpfgc_member_id[ i ].
[0749] If the value of nnpfgc_member_id[ i ] refers to the nnpfgc_id value of the NNPFGC SEI message nnpfgcSei, the nnpfgc_grouping_type of the NNPFGC SEI message nnpfgcSei is 0 for bitstream conformance. If nnpfgc_grouping_type is 0 or 2, there is an NNPF with the same nnpfgc_id value as nnpfgc_member_id[ i ] defined in CLVS for bitstream conformance. If nnpfgc_grouping_type is 1, 3, or 4, for bitstream conformance, there is an NNPF with the same nnpfgc_id value as nnpfgc_member_id[ i ] or an NNPF group with the same nnpfgc_id value as nnpfgc_member_id[ i ] defined in CLVS.
[0750] When nnpfgc_grouping_type is 0, NNPFs with nnpfgc_id values equal to nnpfgc_member_id[ i ] are cascaded in increasing order of i because they are activated by NNPFGA SEI messages with nnpfga_target_id equal to nnpfgc_id.
[0751] nnpfgc_complexity_info_present_flag, nnpfgc_parameter_type_idc, nnpfgc_log2_parameter_bit_length_minus3, nnpfgc_num_parameters_idc, nnpfgc_num_kmac_operations_idc, and nnpfgc_total_kilobyte_size have the semantics of nnpfc_complexity_info_present_flag, nnpfc_parameter_type_idc, nnpfc_log2_parameter_bit_length_minus3, nnpfc_num_parameters_idc, nnpfc_num_kmac_operations_idc, and nnpfc_total_kilobyte_size, respectively, but differ in that these semantics are specified for the NNPF defined in this SEI message rather than the NNPF defined in the NNPFC SEI message. If nnpfgc_grouping_type is 1, nnpfgc_complexity_info_present_flag is equal to 0.
[0752] FIG. 28 shows the syntax of a neural network post-filter group activation SEI message according to embodiments.
[0753] Referring to Fig. 28, the semantics of the Neural-network post-filter group activation (NNPFGA) SEI message are described.
[0754] The NNPFGA SEI message enables or disables post-processing filtering of a picture set using the target neural network post-processing filter group (NNPFG) among the NNPF groups identified by nnpfga_target_id. The nnpfgc_grouping_type of the identified NNPF group is 0 (cascade) or 1 (alternative). If nnpfgc_grouping_type is 1, each member of the group has the same number of input pictures and NNPF output pictures. For a specific picture where the NNPFG is enabled, the target NNPFG precedes the first VCL NAL unit of that picture in the decoding order, and the NNPF of the target NNPFG is defined by the NNPFC SEI message where the nnpfgc_id of the target NNPFG is the same as the nnpfgc_member_id[ i ] value of the target NNPFG and exists in the current picture unit or precedes the current picture in the decoding order.
[0755] To use this SEI message, the following variables must be defined.
[0756] - Input picture width and height in luma samples. Here, denoted as InitCroppedWidth[idx] and InitCroppedHeight[idx], respectively, are the width and height of candidate input pictures with indices idx ranging from 0 to numCandInputPics - 1 (inclusive) that can be used as inputs to NNPFG.
[0757] - Luma sample array InitCroppedYPic[idx] and chroma sample array InitCroppedCbPic[idx] and InitCroppedCrPic[idx] (if any), and the width and height of candidate input pictures with indices idx ranging from 0 to numCandInputPics - 1 (inclusive). Can be used as inputs for NNPFG.
[0758] - Bit depth for the luma sample array of the candidate input picture BitDepthY.
[0759] - BitDepthC for the chroma sample array (if any) of the candidate input picture
[0760] - Chroma format indicator displayed as ChromaFormatIdc
[0761] - When nnpfc_auxiliary_inp_idc is 1, the filtering strength control value array StrengthControlVal[idx] contains real numbers in the range of 0 to 1 (inclusive) for candidate input pictures whose index idx is in the range of 0 to numCandInputPics - 1.
[0762] The candidate input picture with index 0 corresponds to the picture for which NNPFG is activated by this NNPFGA SEI message. Candidate input pictures with index i in the range (inclusive) from 1 to numCandInputPics-1 come before the candidate input picture with index i-1 in the output order. Assume candInputPicList[0] is a list of candidate input pictures output in reverse order.
[0763] nnpfga_target_id is specified by the NNPFGC SEI message and is associated with the current picture, and nnpfgc_id represents the target NNPFG that is the same as nnpfga_target_id.
[0764] The value of nnpfga_target_id ranges from 0 to 2 32 It is in the range (inclusive) up to -2.
[0765] An NNPFGA SEI message with a specific nnpfga_target_id value does not exist in the current PU unless there is an NNPFGC SEI message in the current PU or in a PU that precedes the current PU in decoding order within the current PU or current CLVS where nnpfgc_id is equal to a specific value of nnpfgc_target_id and nnpfgc_grouping_type is 0.
[0766] If PU contains both an NNPFGC SEI message with a specific nnpfgc_id value and an NNPFGA SEI message where nnpfga_target_id is equal to a specific value of nnpfgc_id, the NNPFGC SEI message comes before the NNPFGA SEI message in the decoding order.
[0767] If nnpfga_cancel_flag is 1, it indicates that the persistence of the target NNPFG set by a previous NNPFGA SEI message with the same nnpfga_target_id as the current SEI message is canceled. In other words, the target NNPFG is no longer used unless it is activated by another NNPFGA SEI message with the same nnpfga_target_id as the current SEI message and nnpfga_cancel_flag of 0. If nnpfga_cancel_flag is 0, it indicates that the target NNPFG is activated and available for use.
[0768] nnpfga_persistence_flag indicates the persistence of the target NNPFG for the current layer.
[0769] If nnpfga_persistence_flag is 0, it indicates that the target NNPFG can only be used for post-processing filtering on the current picture.
[0770] If nnpfga_persistence_flag is 1, it indicates that the target NNPFG can be used for post-processing filtering on the current picture and all subsequent pictures of the current layer in the output order until one or more of the following conditions are true.
[0771] - A new CLVS of the current layer starts.
[0772] - The bitstream ends.
[0773] - The picture of the current layer associated with the NNPFGA SEI message having the same nnpfga_target_id as the current SEI message, which follows the current picture in output order.
[0774] Note - Target NNPFG does not apply to subsequent pictures of the current layer associated with an NNPFGA SEI message having the same nnpfga_target_id as the current SEI message.
[0775] nnpfgcTargetPictures is defined as the set of pictures to which the last NNPFGC SEI message, whose nnpfgc_id is equal to nnpfga_target_id, belongs, which comes before the current NNPFGA SEI message in decoding order. nnpfgaTargetPictures is defined as the set of pictures to which the target NNPFG is activated by the current NNPFGA SEI message. It is a bitstream conformance requirement that all pictures included in nnpfgaTargetPictures must also be included in nnpfgcTargetPictures.
[0776] nnpfga_num_filters_minus2 + 2 represents the number of NNPFs of the NNPFG that this SEI message activates. The value of nnpfga_num_filters_minus2 is equal to the value of nnpfgc_num_members_minus2 of the NNPFGC SEI message where nnpfgc_id is the same as nnpfga_target_id.
[0777] If nnpfga_target_base_flag[ i ] is 1, it indicates that the i-th NNPF of the target NNPFG is the base NNPF where nnpfgc_id is nnpfgc_member_id[ i ] in the NNPFGC SEI message where nnpfgc_id is nnpfga_target_id. If nnpfga_target_base_flag[ i ] is 0, it indicates that the i-th NNPF of the target NNPFG is the NNPF specified in the last NNPFC SEI message where nnpfgc_id is nnpfgc_member_id[ i ] in the NNPFGC SEI message where nnpfgc_id is nnpfga_target_id, comes before the first VCL NAL unit of the current picture in the decoding order, and is not a repetition of the NNPFC SEI message containing the base NNPF.
[0778] If nnpfga_input_all_pics_flag[ i ] is 1, it indicates that the input picture for the i-th NNPF is selected without skipping from the list of candidate input pictures candInputPicList[ i ]. If nnpfga_input_all_pics_flag[ i ] is 0, the input picture for the i-th NNPF is selected from the list of candidate input pictures (candInputPicList[ i ]), and some candidate input pictures are skipped.
[0779] nnpfga_num_input_pics_minus1[ i ] represents the number of input pictures for the i-th NNPF in the target NNPFG. If nnpfga_num_input_pics_minus1[ i ] exists, it is equal to nnpfc_num_input_pics_minus1 for the NNPF that has nnpfgc_member_id[ i ] in the NNPFGC SEI message where nnpfgc_id is equal to nnpfga_target_id. If it does not exist, nnpfga_num_input_pics_minus1[ i ] is inferred to be equal to nnpfc_num_input_pics_minus1 for the NNPF that has nnpfgc_id nnpfgc_member_id[ i ] in the NNPFGC SEI message where nnpfgc_id is equal to nnpfga_target_id.
[0780] nnpfga_input_pic_skip_count[ i ][ j ] represents the j-th picture count to skip from the candidate input picture list candInputPicList[ i ] when selecting an input picture for the NNPF activated by the i-th loop item. If nnpfga_input_pic_skip_count[ i ][ j ] is not present, it is assumed to be 0 for all j values in the range from 0 to nnpfga_num_input_pics_minus1[ i ]. The variable numCandInputPics, representing the number of candidate input pictures for the NNPFG, is derived as follows.
[0781] numCandInputPics = 0
[0782] for( j = 0; j <= nnpfga_num_input_pics_minus1
[0000] ; j++ )
[0783] numCandInputPics += 1 + nnpfga_input_pic_skip_count
[0000] [ j ]
[0784] If the range of candInputPicList[ m ] is from 1 to nnpfga_num_filters_minus2 + 1 (inclusive), it becomes a picture list in reverse output order formed in decreasing order of n in the range from 0 to m-1 (inclusive). This includes each picture output from the NNPF process of the n-th loop item that does not already have a corresponding picture in candInputPicList[ m ], and finally includes each picture in candInputPicList
[0000] that does not already have a corresponding picture in candInputPicList[ m ]. For candidate input pictures candInputPicList[ m ][ idx ], where m is in the range from 1 to nnpfga_num_filters_minus2 + 1, and n is the NNPF output picture of the nth NNPF process, if n is less than m, the width and height of the candidate input picture are equal to the nnpfcOutputPicWidth and nnpfcOutputPicHeight of the NNPF output picture, respectively.
[0785] The input picture inputPicList[ m ] list for the NNPF of the m-th loop entry is derived as follows.
[0786] for( k = 0, candIdx = 0; k <= nnpfga_num_input_pics_minus1[ m ]; k++, candIdx++ ) {
[0787] candIdx += nnpfga_input_pic_skip_count[ m ][ k ]
[0788] inputPicList[ m ][ k ] = candInputPicList[ m ][ candIdx ]
[0789] }
[0790] For bitstream compliance, candIdx must not exceed the number of pictures in candInputPicList[ m ].
[0791] For bitstream compatibility, the pictures in inputPicList[ m ] have the same width, height, bit depth, and chroma format for values of m in the range from 1 to nnpfga_num_filters_minus2 + 1.
[0792] To interpret NNPFC SEI messages where nnpfc_id is nnpfgc_member_id[ i ] from NNPFGC SEI messages where nnpfgc_id is nnpfga_target_id, the following variable is assigned to the i-th loop entry.
[0793] - The BitDepthY, BitDepthC, and ChromaFormatIdc variables are used as provided for the interpretation of this SEI message.
[0794] - CroppedWidth and CroppedHeight are set in luma samples to be equal to the width and height of the picture in inputPicList[ i ], respectively.
[0795] - For each input picture k in the range from 0 to nnpfga_num_input_pics_minus1[ i ], the following applies.
[0796] - If CroppedYPic[ k ], CroppedCbPic[ k ] and CroppedCrPic[ k ] exist, they are set to be the same as the corresponding sample array of inputPicList[ i ][ k ].
[0797] - In NNPFGC SEI messages where nnpfgc_id is equal to nnpfga_target_id, for NNPFs where nnpfc_id is equal to nnpfgc_member_id[ i ], if nnpfc_auxiliary_inp_idc is equal to 1, the following applies.
[0798] - It is a bitstream conformance requirement that for all idx values in the range from 0 to nnpfga_num_input_pics_minus1[ i ], inputPicList[ i ][ k ] must be equal to candInputPicList
[0000] [ idx ]. numCandInputPics - 1 (inclusive).
[0799] - StrengthControlVal[ k ] is set to be the same as InitStrengthControlVal[ idx ].
[0800] nnpfga_num_output_entries[ i ] represents the number of nnpfga_output_flag[ i ][ j ] syntax elements present in NNPFGA SEI messages. The value of nnpfga_num_output_entries[ i ] is in the range from 0 to NumInpPicsInOutputTensor for NNPFGC SEI messages where nnpfc_id is equal to nnpfgc_member_id[ i ] and nnpfgc_id is equal to nnpfga_target_id.
[0801] If nnpfga_output_flag[ i ][ j ] is 1, it indicates that the NNPF-generated picture corresponding to the input picture with index InpIdx[ j ] derived for the i-th NNPF of the target NNPFG is output by the NNPF process activated by this loop entry. Here, the NNPF process is specified in the semantics of the NNPFC SEI message. If nnpfga_output_flag[ i ][ j ] is 0, the input picture with index InpIdx[ j ] derived for the i-th NNPF of the target NNPFG
[0802] Indicates that the NNPF-generated picture corresponding to is not output by the NNPF process enabled by this loop entry. If nnpfga_num_output_entries[ i ] is less than NumInpPicsInOutputTensor derived for the i-th NNPF of the target NNPFG, nnpfga_output_flag[ i ][ j ] is inferred to be 1 for each i value in the range NumInpPicsInOutputTensor - 1 (inclusive) in nnpfga_num_output_entries[ i ].
[0803] NnpfgaOutputPicList, which is a list of pictures output in output order by the NNPF processes of NNPFG, is initially empty and is formed in decreasing order of n in the range from 0 to nnpfga_num_filters_minus2 + 1, and contains each picture output by the NNPF process of the nth loop item for which there is no picture already corresponding to NnpfgaOutputPicList.
[0804] FIG. 29 shows source picture timing information according to embodiments.
[0805] Source Picture Timing Information (SPTI) SEI messages indicate the temporal distance between the source picture associated with the corresponding decoded output picture prior to encoding. For example, in the case of content captured by a camera, the temporal distance between source pictures is the difference between the time the image sensor was exposed to generate the source picture associated with the currently decoded picture and the time the image sensor was exposed to generate the source picture associated with the previously decoded picture in the output order. The information provided by the SPTI SEI message applies only to all subsequent pictures in the current layer in the output order, starting from the picture in the current layer of the access unit containing the SPTI SEI message.
[0806] If spti_cancel_flag is 1, it indicates that the persistence of the previous SPTI SEI message in the output order applied to the current layer is canceled. If spti_cancel_flag is 0, it indicates that source picture timing information follows.
[0807] spti_persistence_flag indicates the persistence of SPTI SEI messages for the current layer.
[0808] If spti_persistence_flag is 0, it indicates that SPTI SEI messages are applied only to the currently decoded picture.
[0809] If spti_persistence_flag is 1, it indicates that SPTI SEI messages apply only to the currently decoded picture and persist to all subsequent pictures of the current layer in output order until one or more of the following conditions are met. If spti_persistence_flag is 1, it indicates that it applies to multiple sublayers.
[0810] - A new CLVS of the current layer starts.
[0811] - The bitstream ends.
[0812] - The picture in the current layer of the AU associated with the SPTI SEI message is output after the current picture in the output order.
[0813] If spti_source_timing_equals_output_timing_flag is 1, it indicates that the timing of the source picture is the same as the timing of the corresponding output picture decoded. If spti_source_timing_equals_output_timing_flag is 0, it indicates that the timing of the source picture may be different from the timing of the corresponding output picture decoded.
[0814] If spti_source_timing_equals_output_timing_flag is 1 and there is a picture timing SEI message for the current picture, the source picture timing can be determined through the information passed in the picture timing SEI message.
[0815] If spti_source_type_present_flag is 1, it indicates that the syntax element spti_source_type exists in the SEI message. If spti_source_type_present_flag is 0, it indicates that the syntax element spti_source_type does not exist in the SEI message.
[0816] spti_source_type represents the timing relationship between the source picture specified in Table 10 (Interpretation of spti_source_type) and the corresponding decoded output picture. Here, if (spti_source_type & bitMask) is not 0, it indicates that the timing relationship has an interpretation related to the bitMask value in Table 10 (Interpretation of spti_source_type). If spti_source_type is greater than 0 and (spti_source_type & bitMask) is 0, the interpretation related to the bitMask value is not applied to SPTI SEI messages. If spti_source_type is 0, the timing relationship can be specified by the application.
[0817] The value of spti_source_type is in the range from 0 to 127 in bitstreams compliant with this version of this document. Values for spti_source_type from 128 to 255 (inclusive) are reserved for future use by ITU-T | ISO / IEC and do not exist in bitstreams compliant with this version of this document. Decoders compliant with this version of this document ignore SPTI SEI messages with spti_source_type in the range of 128 to 255.
[0818] [Table 10]
[0819]
[0820] The values of (spti_source_type & 0x04) and (spti_source_type & 0x08) are 0 (for example, spti_source_type must not represent fast imaging and time-lapse imaging simultaneously).
[0821] spti_time_scale represents the number of time units that elapse in one second. The value of spti_time_scale cannot be 0. For example, the spti_time_scale of a time coordinate system measuring time using a 27 MHz clock is 27,000,000.
[0822] spti_num_units_in_elemental_interval represents the number of time units of a clock operating at a spti_time_scale Hz frequency corresponding to the specified element source picture interval of consecutive pictures in output order in CLVS. The value of spti_num_units_in_elemental_interval is not 0.
[0823] The element source picture interval in seconds, represented by the ElementalSourcePictureInterval variable, is equal to the quotient obtained by dividing spti_num_units_in_elemental_interval by spti_time_scale. For example, to represent the element source picture interval as 0.04 seconds, spti_time_scale could be 27,000,000 and spti_num_units_in_elemental_interval could be 1,080,000.
[0824] The method of representing the element source picture interval is similar to the time_scale in Rec. ITU-T H.266 | ISO / IEC 23090-3, where spti_time_scale is similar to the time_scale in that syntax and spti_num_units_in_elemental_interval is similar to the num_units_in_tick in that syntax, and thus the ElementalSourcePictureInterval variable is similar to the ClockTick variable in Rec. ITU-T H.266 | ISO / IEC 23090-3.
[0825] spti_max_sublayers_minus_1 + 1 represents the maximum number of temporal sublayers to which picture interval scale factor (spti_sublayer_interval_scale_factor[ i ]) and synthesis flag (spti_sublayer_synthesized_picture_flag[ i ]) information is signaled. If spti_max_sublayers_minus_1 is not present, it is inferred to be the same as TemporalId.
[0826] If spti_sublayer_interval_scale_factor[ i ] is present, it represents the scale factor used to determine the source picture interval of the corresponding picture in the CLVS with TemporalId i. This picture is relative to the previous output picture with TemporalId less than or equal to i. A value of 0 can be used to indicate that the source picture corresponding to the currently decoded output picture is the same as the source picture corresponding to the previous decoded output picture with TemporalId less than or equal to i.
[0827] The displayed source picture interval associated with the output picture with TemporalId i is represented by the SourcePictureInterval[ i ] variable in seconds compared to the previous output picture with TemporalId less than or equal to i, and is derived as follows.
[0828] SourcePictureInterval[ i ] = ElementalSourcePictureInterval * spti_sublayer_interval_scale_factor[ i ] *
[0829] ( 1 - 2 * temporalReversalFlag )
[0830] If spti_source_type_present_flag is 1, the variable temporalReversalFlag is equal to ( spti_source_type & 0x10 )? 1 : 0. Otherwise (i.e., if spti_source_type_present_flag is 0), the variable temporalReversalFlag is equal to 0.
[0831] When calculating SourcePictureInterval[i], ElementalSourcePictureInterval is multiplied by spti_sublayer_interval_scale_factor[i], so the same value of SourcePictureInterval[i] can be expressed in various ways by applying a scale factor to the spti_time_scale value and applying the same scale factor to spti_num_units_in_elemental_interval or spti_sublayer_interval_scale_factor[i]. There is no assumption that the common scale factor has been removed or that the value of spti_sublayer_interval_scale_factor[i] is equal to 1 for the highest value of i. The reason for allowing the same value to be expressed in various ways is, at least in part, to allow spti_time_scale to be selected to match other timing-related factors used in the system environment, such as the 27 MHz clock speed used in some multimedia communication systems.
[0832] If spti_sublayer_synthesized_picture_flag[ i ] is present, if it is 1, it indicates that the decoded output picture belonging to the i-th temporal sublayer has been composited and does not match the original source picture that was not modified. If spti_sublayer_synthesized_picture_flag[ i ] is 0, it does not provide this indication. If it is absent, the value of spti_sublayer_synthesized_picture_flag[ i ] is inferred to be 0.
[0833] If the TemporalId of an SPTI SEI message is greater than 0 and the SPTI SEI message persists for one or more pictures with a low TemporalId, the encoder can include the information of the SPTI SEI message in one or more SPTI SEI messages with a low TemporalId to prevent information loss when pictures of the lower time layer are lost or removed.
[0834] FIGS. 30a and FIGS. 30b show object mask information SEI messages according to embodiments.
[0835] The Object Mask Information (OMI) SEI message provides object mask information for an object mask picture in an auxiliary picture layer associated with the primary picture layer where the SEI message exists (currently referred to as the primary picture layer). If an OMI SEI message exists, it exists in the primary picture layer. A single primary picture layer may be associated with one or more auxiliary picture layers. For each associated auxiliary picture layer containing an object mask picture with nuh_layer_id equal to sdi_layer_id[ i ], the value of sdi_aux_id[ i ] is equal to AUX_OBJECT_MASK for all i values in the range from 0 to sdi_max_layers_minus1.
[0836] To use this SEI message, the following variables must be defined.
[0837] - The cropped picture width and picture height in Luma sample units are indicated as CroppedWidth and CroppedHeight, respectively.
[0838] - Fit Crop Window Left Offset, ConfWinLeftOffset
[0839] - Fit Crop Window Top Offset, ConfWinTopOffset
[0840] Chroma format indicator displayed as -ChromaFormatIdc
[0841] The SubWidthC and SubHeightC variables are derived from ChromaFormatIdc.
[0842] If omi_cancel_flag is 1, it indicates that the SEI message cancels the persistence of the previous object mask information SEI message in the same layer in the output order. If omi_cancel_flag is 0, it indicates that the object mask information follows.
[0843] omi_persistence_flag indicates the persistence of the object mask information provided in this SEI message. If omi_persistence_flag is 0, it indicates that the object mask information is applied only to the current picture. If omi_persistence_flag is 1, it indicates that the object mask information is applied to the current picture and all subsequent pictures on the same layer in output order until one or more of the following conditions are true.
[0844] - A new CLVS of the current layer starts.
[0845] - The bitstream ends.
[0846] - The picture in the current layer of the PU containing the object mask information SEI message is output after the current picture in the output order.
[0847] If sdi_aux_id[ i ] for one or more of the i values in CVS does not contain an SDI SEI message such as AUX_OBJECT_MASK, the OMI SEI message is ignored.
[0848] If AU contains both an SDI SEI message and an OMI SEI message such as AUX_OBJECT_MASK for one or more of the i values, the SDI SEI message comes before the OMI SEI message in the decoding order.
[0849] omi_num_aux_pic_layer represents the number of auxiliary picture layers associated with the current primary picture layer. For bitstream conformance, the value of omi_num_aux_pic_layer must be equal to numAuxLayer, where the numAuxLayer variable is derived as follows.
[0850] omiPrimaryLayerId is represented as the nuh_layer_id value in NAL units containing SEI messages.
[0851] numAuxLayer = 0;
[0852] for( i = 0; i <= sdi_max_layers_minus1; i++ )
[0853] if( sdi_aux_id[ i ] = = AUX_OBJECT_MASK )
[0854] for( j = 0; j <= sdi_num_associated_primary_layers_minus1[ i ]; j++ )
[0855] if( sdi_layer_id[ sdi_associated_primary_layer_idx[ i ][ j ] ] = = omiPrimaryLayerId )
[0856] numAuxLayer++;
[0857] omi_mask_id_length_minus1 + 1 represents the length (in bits) of the omi_mask_id[ i ][ j ] syntax element.
[0858] omi_mask_sample_value_length_minus8 + 8 represents the length (in bits) of the omi_aux_sample_value[ i ][ j ] syntax element. The value of omi_mask_sample_value_length_minus8 is in the range of 0 to 8.
[0859] If omi_mask_confidence_info_present_flag is 1, it indicates that there is a syntax element of omi_mask_confidence[ i ][ j ]. If omi_mask_confidence_info_present_flag is 0, it indicates that there is no syntax element of omi_mask_confidence[ i ][ j ].
[0860] omi_mask_confidence_length_minus1 + 1 represents the length (in bits) of the omi_mask_confidence[ i ][ j ] syntax element.
[0861] If omi_mask_depth_info_present_flag is 1, it indicates that the omi_mask_depth[ i ][ j ] syntax element exists. If omi_mask_depth_info_present_flag is 0, it indicates that the omi_mask_depth[ i ][ j ] syntax element does not exist.
[0862] omi_mask_depth_length_minus1 + 1 represents the length (in bits) of the omi_mask_depth[ i ][ j ] syntax element.
[0863] For bitstream conformance, the values of omi_num_aux_pic_layer, omi_mask_id_length_minus1, omi_mask_sample_value_length_minus8, omi_mask_confidence_info_present_flag, omi_mask_confidence_length_minus1 (if present), omi_mask_depth_info_present_flag and omi_mask_depth_length_minus1 (if present) are identical in all object_mask_info() syntax structures within CLVS.
[0864] If omi_mask_label_info_present_flag is 1, it indicates that the omi_mask_label_language_present_flag and omi_mask_label[ i ][ j ] syntax elements are present. If omi_mask_label_info_present_flag is 0, it indicates that the omi_mask_label_language_present_flag and omi_mask_label[ i ][ j ] syntax elements are not present.
[0865] If omi_mask_label_language_present_flag is 1, it indicates that the omi_mask_label_language syntax element exists, and if omi_mask_label_language_present_flag is 0, it indicates that the omi_mask_label_language syntax element does not exist.
[0866] omi_bit_equal_to_zero is equal to 0.
[0867] omi_mask_label_language contains language tags specified in IETF RFC 5646 and a null termination byte such as 0x00. The length of the omi_mask_label_language syntax element is 255 bytes or less, excluding the null termination byte. If this element is omitted, the language of the label is not specified.
[0868] If omi_mask_pic_update_flag[ i ] is 1, it indicates that the object mask information of the object mask picture in the i-th secondary picture layer associated with the current primary picture layer can be updated. If omi_mask_pic_update_flag[ i ] is 0, it indicates that the mask information of the object mask picture in the i-th secondary picture layer associated with the current primary picture layer will not be changed. If omi_mask_pic_update_flag[ i ] is 0, the persistence mechanism is used. That is, the object mask information is inherited from the last OMI SEI message in the same layer in the decoding order, and this message signals the mask information for the object mask picture in the i-th secondary picture layer associated with the current primary picture layer.
[0869] omi_num_mask_in_pic_update[ i ] represents the number of object masks in the object mask picture of the i-th auxiliary picture layer associated with the current primary picture layer. omi_num_mask_in_pic_update[ i ] is in the range from 0 to (1<<(omi_mask_id_length_minus1 + 1)) - 1.
[0870] omi_mask_id[ i ][ j ] represents the identifier of the j-th object mask in the object mask picture of the i-th auxiliary picture layer associated with the current primary picture layer. The length of the omi_mask_id[ i ][ j ] syntax element is omi_mask_id_length_minus1 + 1 bit.
[0871] The variable maskId[ i ][ j ], which specifies the j-th object mask identifier of the i-th auxiliary picture layer object mask picture associated with the current primary picture layer, is derived as follows.
[0872] for(i = 0; i < omi_num_aux_pic_layer; i++)
[0873] for(j = 0; j < omi_num_mask_in_pic_update[ i ]; j++)
[0874] maskId[ i ][ j ] = omi_mask_id[ i ][ j ] + (1<<(omi_mask_id_length_minus1 + 1))*i
[0875] omi_aux_sample_value[ i ][ j ] represents the sample value within the j-th object mask area of the object mask picture in the i-th auxiliary picture layer associated with the current primary picture layer.
[0876] If omi_mask_cancel[ i ][ j ] is 1, cancels the persistence range of the j-th object mask of the object mask picture in the i-th secondary picture layer associated with the current primary picture layer. If omi_mask_cancel[ i ][ j ] is 0, it indicates that the j-th object mask information of the object mask picture in the i-th secondary picture layer associated with the current primary picture layer is passed as a signal.
[0877] As a requirement for bitstream conformance, when omi_mask_id[ i ][ j ] with a specific value in the current CLVS is parsed for the first time, the value of the corresponding omi_mask_cancel[ i ][ j ] is equal to 0.
[0878] If omi_mask_bounding_box_present_flag[ i ][ j ] is equal to 1, it indicates that the syntax elements omi_mask_top[ i ][ j ], omi_mask_left[ i ][ j ], omi_mask_width[ i ][ j ], and omi_mask_height[ i ][ j ] exist. If omi_mask_bounding_box_present_flag[ i ][ j ] is 0, it indicates that the syntax elements omi_mask_top[ i ][ j ], omi_mask_left[ i ][ j ], omi_mask_width[ i ][ j ], and omi_mask_height[ i ][ j ] do not exist.
[0879] omi_mask_top[ i ][ j ], omi_mask_left[ i ][ j ], omi_mask_width[ i ][ j ], and omi_mask_height[ i ][ j ] represent the top-left corner coordinates, width, and height of the bounding box of the j-th object mask in the cropped decoded object mask picture of the i-th secondary picture layer associated with the current primary picture layer, based on the conformity cropping window specified in the active SPS, respectively.
[0880] The value of omi_mask_left[ i ][ j ] must be in the range from 0 to ( CroppedWidth / SubWidthC - 1 ), where CroppedWidth and SubWidthC are associated with the object mask picture of the i-th secondary picture layer associated with the current primary picture layer. If no value is found, the value of omi_mask_left[ i ][ j ] is inferred to be 0.
[0881] The value of omi_mask_top[ i ][ j ] must be in the range from 0 to ( CroppedHeight / SubHeightC - 1 ), where CroppedHeight and SubHeightC are associated with the object mask picture of the i-th secondary picture layer associated with the current primary picture layer. If no value is found, the value of omi_mask_top[ i ][ j ] is inferred to be 0.
[0882] The value of omi_mask_width[ i ][ j ] is in the range from 0 to (CroppedWidth / SubWidthC - omi_mask_left[ i ][ j ] ). If there is no value, the value of omi_mask_width[ i ][ j ] is inferred as (CroppedWidth / SubWidthC - omi_mask_left[ i ][ j ] ).
[0883] The value of omi_mask_height[ i ][ j ] is in the range from 0 to (CroppedHeight / SubHeightC - omi_mask_top[ i ][ j ] ). If there is no value, the value of omi_mask_height[ i ][ j ] is inferred as (CroppedHeight / SubWidthC - omi_mask_top[ i ][ j ] ).
[0884] The identified object mask is within a bounding box (bounding box) containing a luminance sample with horizontal coordinates of SubWidthC * (ConfWinLeftOffset + omi_mask_left[ i ][ j ] ) to SubWidthC * (ConfWinLeftOffset + omi_mask_left[ i ][ j ] + omi_mask_width[ i ][ j ] ) - 1 (inclusive) and vertical coordinates of SubHeightC * (ConfWinTopOffset + omi_mask_top[ i ][ j ] ) to SubHeightC * (ConfWinTopOffset + omi_mask_top[ i ][ j ] + omi_mask_height[ i ][ j ] ) - 1 (inclusive).
[0885] The variable pI[ i ] [ x ][ y ] is the decoded value of the sample at the relative sample position (x, y) in the cropped object mask picture of the i-th auxiliary picture layer associated with the current primary picture layer. The next process is to determine the mask area in the auxiliary picture.
[0886] for( i = 0; i < omi_num_aux_pic_layer; i++ )
[0887] for( j = 0; j < omi_num_mask_in_pic_update[ i ]; j++ )
[0888] if( pI[ i ][ x ][ y ] = = omi_aux_sample_value [ i ][ j ]
[0889] && x >= omi_mask_left[ i ][ j ]
[0890] && x < omi_mask_left[ i ][ j ] + omi_mask_width[ i ][ j ]
[0891] && y >= omi_mask_top[ i ][ j ]
[0892] && y < omi_mask_top[ i ][ j ] + omi_mask_height[ i ][ j ] )
[0893] the sample at location (x, y) is associated with the object mask with the identifier of maskId[ i ][ j ]
[0894] omi_mask_confidence[ i ][ j ] represents the confidence associated with the j-th object mask of the i-th auxiliary picture layer, which is associated with the current primary picture layer, in units of 2 - ( omi_mask_confidence_length_minus1 + 1 ). Therefore, a higher value of omi_mask_confidence[ i ][ j ] indicates higher confidence. The length of the omi_mask_confidence[ i ][ j ] syntax element is omi_mask_confidence_length_minus1 + 1 bits.
[0895] omi_mask_depth[ i ][ j ] represents the object depth associated with the object mask of the i-th auxiliary picture layer associated with the current primary picture layer and the j-th object mask of the picture. A smaller omi_mask_depth value indicates a shorter distance to the object. The length of the omi_mask_depth[ i ][ j ] syntax element is omi_mask_depth_length_minus1 + 1 bit.
[0896] omi_mask_label[ i ][ j ] represents the contents of the label associated with the j-th object mask of the i-th auxiliary picture layer object mask picture associated with the current primary picture layer. The length of the omi_mask_label[ i ][ j ] syntax element is 255 bytes or less, excluding the null termination byte.
[0897] FIG. 31 shows an SEI processing order SEI message according to embodiments.
[0898] SEI Processing Order (SPO): SEI messages convey information indicating the preferred processing order determined by the encoder (i.e., the content creator) for the various types of SEI messages that may exist in CVS.
[0899] The meaning of SPO SEI messages utilizes the concept of SEI message types. SEI messages with different payloadType values are considered to be of different types of SEI messages. Additionally, SEI messages with the same payloadType value but distinguished by the syntax element values of the SEI payload are also considered to be of different types. This distinction based on the syntax element values of the SEI payload is performed by comparing the transmitted value using the po_sei_prefix_data_bit[ i ][ j ] syntax element (if present) or the value transmitted as an SEI message within a processing order nested SEI message (if present). For example, Neural Network Post-Proof Filter (NNPFC) SEI messages can be distinguished by having different nnpfc_id values.
[0900] If both po_sei_wrapping_flag[ i ] and po_sei_prefix_flag[ i ] of the i-th SEI message seiA in all SPO SEI messages are 0, then no other SEI message seiB is included in the same SPO SEI message or other SPO SEI messages in the current CVS, provided that all of the following conditions are true.
[0901] - The value of po_sei_payload_type[ i ] in seiB is the same as the value in seiA.
[0902] - The value of seiB's po_sei_wrapping_flag[ i ] is 0.
[0903] - The value of po_sei_prefix_flag[ i ] of seiB is 1.
[0904] If an SPO SEI message with a specific po_id value exists in an access unit of CVS, that SPO SEI message with the corresponding po_id value exists in the first access unit of CVS in terms of decoding order. The number of SEI messages and the payloadType code indicated within each SPO SEI message with the same po_id value are maintained in decoding order from the current access unit to the end of CVS (in terms of output order).
[0905] An SPO SEI message may contain one or more SEI prefix marks for a specific payloadType. Each SEI prefix mark is a bit string that follows the SEI payload syntax of the corresponding payloadType value and includes multiple complete syntax elements starting from the first syntax element of the SEI payload. These SEI prefix marks provide sufficient information to determine the specific processing order for SEI message types that have the same payloadType value but different preferred processing orders.
[0906] po_id contains an identification number that identifies the SPO SEI message.
[0907] The processing chain consists of a list of SEI message types identified by the SPO SEI message according to the preferred processing order specified in the SPO SEI message.
[0908] Each SEI message type in the processing chain indicated by the SPO SEI message is identified by the syntax elements po_sei_payload_type[ i ], po_sei_wrapping_flag[ i ], po_sei_processing_order[ i ], and if present, po_num_bits_in_prefix_indication_minus1[ i ] and po_prefix_data_bit[ i ][ j ].
[0909] SEI message types do not need to belong to any processing chain and can belong to multiple processing chains identified by SPO SEI messages having different po_id values.
[0910] Each SEI message of the SEI message type identified within the SPO SEI message has the same persistence range as when the corresponding SEI message is transmitted outside the SPO SEI message and is not identified within the SPO SEI message.
[0911] Processing chains can be interchangeable; that is, at most one processing chain may be selected to be applied. Alternatively, they can be complementary; that is, two or more processing chains may be selected and applied individually, with each processing chain producing a single output.
[0912] po_num_sei_messages_minus2 + 2 represents the number of SEI message types with a preferred processing order in SPO SEI messages.
[0913] If po_sei_wrapping_flag[ i ] is 1, it indicates that there must be at least one processing order nested SEI message satisfying both of the following two constraints.
[0914] - pon_target_po_id[ j ] with all values of j is equal to po_id.
[0915] - In the processing order nested SEI message, there is a k-th loop entry, so the payloadType of the k-th nested SEI message is equal to po_sei_payload_type[ i ] and pon_processing_order[ k ] is equal to po_sei_processing_order[ i ].
[0916] If po_sei_wrapping_flag[ i ] is 0, the SEI message with payloadType po_sei_payload_type[ i ] (and if po_sei_prefix_flag[ i ] is 1, the prefix data matching the value of po_sei_prefix_data_bit[ i ][ j ]) is outside the processing order nested SEI message.
[0917] When po_sei_wrapping_flag[ i ] is 1, SEI messages can be passed within nested SEI messages in the processing order, thus preventing decoders that do not process SPO SEI messages from misinterpreting the SEI message. Therefore, when po_sei_wrapping_flag[ i ] is 0, if unintended results may be generated in the corresponding decoder, po_sei_wrapping_flag[ i ] is used.
[0918] po_sei_importance_flag[ i ] represents the importance determined by the encoder for the SEI message type with index i.
[0919] If the decoding system cannot interpret or does not support the features indicated in the SEI message where po_sei_importance_flag[ i ] is 1, the entire SPO SEI message is ignored.
[0920] po_sei_payload_type[ i ] represents the payloadType value of the i-th type of the SEI message.
[0921] If po_sei_prefix_flag[ i ] is 1, it indicates that po_num_bits_in_prefix_indication_minus1[ i ] and some po_sei_prefix_data_bit[ i ][ j ] syntactic elements are present. If po_sei_prefix_flag[ i ] is 0, it indicates that these syntactic elements are absent.
[0922] SeiProcessingOrderSeiList is configured to consist of payloadType values 3, 4, 5, 19, 137, 142, 144, 147, 148, 149, 165, 177, 210, 211. For each i in the range from 0 to po_num_sei_messages_minus2 + 1, the po_sei_payload_type[ i ] value is equal to the value of SeiProcessingOrderSeiList.
[0923] po_sei_processing_order[ i ] indicates the default processing order for SEI messages of the i-th type, for which default processing order information is provided in the SPO SEI message. For two different integer values of m and n, if po_sei_processing_order[ m ] is less than po_sei_processing_order[ n ], it indicates that the SEI message type associated with index m must be processed before the SEI message type associated with index n; if po_sei_processing_order[ m ] is equal to po_sei_processing_order[ n ], it indicates that there is no preferred processing order between the SEI message types associated with index m and n (e.g., they may represent different attributes applicable to both at that stage or alternative processes that can be applied, or one may represent an attribute and the other a process).
[0924] If i is greater than 0, po_sei_processing_order[ i ] is greater than or equal to po_sei_processing_order[ i - 1 ].
[0925] If present, po_num_bits_in_prefix_indication_minus1[ i ] and po_sei_prefix_data_bit[ i ][ j ] have the same semantics as the num_bits_in_prefix_indication_minus1[ i ] and sei_prefix_data_bit[ i ][ j ] syntax elements of the SEI prefix indication SEI message, and prefix_sei_payload_type is replaced with po_sei_payload_type[ i ].
[0926] If there are two or more SPO SEI messages with a specific po_id value in CVS, the value of po_num_sei_messages_minus2 and the values of po_sei_wrapping_flag[ i ], po_sei_prefix_flag[ i ], po_sei_importance_flag[ i ], po_sei_payload_type[ i ], and po_sei_processing_order[ i ] for each i value are identical to other SPO SEI messages with the same po_id value in CVS.
[0927] po_byte_alignment_bit_equal_to_one is equal to 1.
[0928] FIG. 32 shows a processing order nesting SEI message according to embodiments.
[0929] A Proceeding Order Nested (PON) SEI message contains one or more SEI messages that must be applied only as part of the processing chain identified by the associated SEI processing order SEI message and must not be applied in a manner that conflicts with the processing chain identified by the associated SEI processing order SEI message.
[0930] A SEI message included in a PON SEI message is called a PON nested SEI message.
[0931] An encoder may include multiple PON SEI messages within the same access unit. For example, the first PON SEI message of an access unit may include a PON nested SEI message applicable to multiple processing chains and one or more other PON SEI messages of the same access unit applicable to only a single processing chain.
[0932] As a requirement of bitstream conformity, the meaning and effect of SEI messages other than PON nested SEI messages must not depend on PON nested SEI messages. The consequences of this constraint include the following specific restrictions, wherein the associated SEI message is considered to be an SEI message that affects the meaning or effect of a specific SEI message.
[0933] If there exists a neural network post-filter characteristic SEI message with a specific value of nnpfc_id, which is a PON nested SEI message, then the associated neural network post-filter activation SEI message where nnpfa_target_id is the same as the corresponding nnpfc_id value is also a PON nested SEI message.
[0934] - If nnpfa_persistence_flag is 1 and there is a Neural Network Post-Filter Activation (NNPFA) SEI message with a specific value of nnpfa_target_id that is not a PON overlapping SEI message, then the next picture in the same CLVS that has an NNPFA SEI message with the same nnpfa_target_id value (if any) in output order does not have an associated NNPFA SEI message that is a PON overlapping SEI message.
[0935] - If fg_characteristics_persistence_flag is 1 and there is a film grain characteristic SEI message that is not a PON overlapping SEI message, there is no associated film grain characteristic SEI message that is a PON overlapping SEI message in the same CLVS.
[0936] - If fp_arrangement_persistence_flag is 1 and there is a frame packing array SEI message that is not a PON nested SEI message, then there is no associated frame packing array SEI message in the same CLVS where fp_arrangement_cancel_flag is 1 or has the same fp_arrangement_id value. This message is a PON nested SEI message.
[0937] - If the content color volume SEI message with ccv_persistence_flag 1 is not a PON nested SEI message, there is no associated frame packing array SEI message that is a PON nested SEI message in the same CLVS.
[0938] If erp_persistence_flag is 1 and there is an isotropic projection SEI message that is not a PON nested SEI message, there is no associated isotropic projection SEI message that is a PON nested SEI message in the same CLVS.
[0939] - If gcmp_persistence_flag is 1 and there is a generalized cubemap project SEI message that is not a PON nested SEI message, there is no associated generalized cubemap project SEI message that is a PON nested SEI message in the same CLVS.
[0940] - If sphere_rotation_persistence_flag is 1 and there is a spherical rotation SEI message that is not a PON overlapping SEI message, there is no associated spherical rotation SEI message that is a PON overlapping SEI message in the same CLVS.
[0941] - If rwp_persistence_flag is 1 and there is a local packing SEI message that is not a PON nested SEI message, there is no associated local packing SEI message that is a PON nested SEI message in the same CLVS.
[0942] - If omni_viewport_persistence_flag is 1 and there is a forward viewport SEI message that is not a PON nested SEI message, there is no associated forward viewport SEI message that is a PON nested SEI message in the same CLVS.
[0943] - If sari_persistence_flag is 1 and there is a sample aspect ratio SEI message that is not a PON nested SEI message, there is no associated sample aspect ratio SEI message that is a PON nested SEI message in the same CLVS.
[0944] - If there is an annotated local SEI message that is not a PON nested SEI message, there is no associated annotated local SEI message that is a PON nested SEI message in the same CLVS.
[0945] If an alpha channel information SEI message exists that is not a PON nested SEI message, then an associated alpha channel information SEI message that is a PON nested SEI message does not exist in the same CLVS.
[0946] - If a display direction SEI message exists that is not a PON nested SEI message, there is no associated display direction SEI message that is a PON nested SEI message in the same CLVS.
[0947] - If there is a color transformation indicator SEI message with colour_transform_persistence_flag of 1 that is not a PON nested SEI message, there is no associated color transformation indicator SEI message in the same CLVS that is a PON nested SEI message with colour_transform_cancel_flag of 1 or has the same colour_transform_id value.
[0948] pon_num_po_ids_minus1 + 1 represents the number of SEI processing sequence SEI messages associated with this PON SEI message.
[0949] pon_target_po_id[ i ] represents the po_id of the i-th associated SEI processing sequence SEI message.
[0950] pon_num_seis_minus1 + 1 represents the number of PON nested SEI messages included in this PON SEI message.
[0951] pon_processing_order[ i ] indicates the position of the i-th processing order nested SEI message within the processing order defined in the associated SEI processing order SEI message. If i is greater than 0, pon_processing_order[ i ] is greater than or equal to pon_processing_order[ i - 1 ].
[0952] Each associated SEI processing sequence SEI message must have at least one value of i in the range from 0 to pon_num_seis_minus1, and the associated SEI processing sequence SEI message has an item k in which all of the following are true.
[0953] - po_sei_processing_order[ k ] is equal to pon_processing_order[ i ].
[0954] - po_sei_payload_type[ k ] is equal to the payloadType value of the i-th PON nested SEI message.
[0955] - When po_sei_prefix_flag[ k ] is 1, po_sei_prefix_data_bit[ k ][ j ] for j in the range from 0 to po_num_bits_in_prefix_indication_minus1[ k ] contains the same content as po_num_bits_in_prefix_indication_minus1[ k ] of the SEI message payload of the i-th PON nested SEI message plus 1 initial bit.
[0956] The i-th PON nested SEI message is applied as the k-th loop item of the associated SEI processing sequence SEI message.
[0957] Processing of the processing chain
[0958] Processing chains are interchangeable. That is, the decoding system can select and apply at most one processing chain at a time.
[0959] The decoding system can select and apply the processing chain as follows.
[0960] First, decode the bitstream, set the PoPicList to a list of decoded images truncated in output order generated from the bitstream decoding, and select a processing chain.
[0961] - (Option 1: Picture-wise, zigzag, width-first for a single filter) For each SEI message type of the selected processing chain, the following are applied in the non-decreasing order of the corresponding po_sei_processing_order[ i ] values.
[0962] - If the SEI message associated with the i-th SEI message type persists for picA or the NNPF that generated picA for a picture activated by a previous process in the processing chain, the following applies to each picture picA in PoPicList in the output order.
[0963] - If picA is not a cropped decoded picture, the following exception applies to SEI message interpretation.
[0964] - Interface variables for SEI message interpretation are derived from picA instead of syntactic elements representing the attributes of the corresponding truncated decoded picture.
[0965] The meaning of the SEI message, or the meaning of the SEI message, and the meaning of the associated NNPFC SEI message if the SEI message is an NNPFA SEI message, is applied to the pictures in PoPicList instead of the truncated decoded picture.
[0966] - If the i-th SEI message type exists in SpoProcessingList, the process included in the SEI message is executed, the corresponding picture is replaced with the corresponding processed picture (if any) generated as a result of the process, and another picture (if any) generated as a result of the process is inserted into PoPicList to maintain the output order, thereby updating PoPicList.
[0967] (Option 2: Filter-by-filter filtering, jagging compression, depth-first for a picture) The following is applied iteratively to each picture picA in PoPicList in the output order. If a set of SEI messages associated with an SEI message type in the SpoProcessingList of the selected processing chain persists for picA, the following is applied.
[0968] - The following are applied to each set of SEI messages in the non-decreasing order of the corresponding po_sei_processing_order[ i ] values.
[0969] - If the current SEI message is not the first message in the SEI message set, the following exceptions apply to SEI message interpretation.
[0970] - Interface variables for interpreting SEI messages are derived from the pictures in the updated PoPicList instead of syntax elements representing the attributes of the corresponding truncated decoded picture.
[0971] - The meaning of the SEI message, or the meaning of the SEI message, and the meaning of the associated NNPFC SEI message if the SEI message is an NNPFA SEI message, is applied to the picture in PoPicList instead of the decoded cropped picture.
[0972] The process included in the SEI message is repeatedly called for each picture in picA and PoPicList in output order. These pictures are the interpolated or extrapolated pictures, or their corresponding ones, generated by applying the process included in the previous SEI message to picA. Whenever a process is called, PoPicList is updated to maintain output order by replacing the picture with the corresponding processed picture (if any) generated as a result of that process and inserting another picture (if any) into PoPicList.
[0973] FIG. 33 shows the syntax of an encoder optimization information SEI message according to embodiments.
[0974] Encoder optimization information SEI messages are used to indicate whether the video is optimized for human viewing or machine analysis, and what type of optimization was applied during the preprocessing or encoding process.
[0975] If eoi_cancel_flag is 1, it indicates that the persistence of the encoder optimization information SEI message included in the previous PU in the output order is canceled. If eoi_cancel_flag is 0, it indicates that the optimization information applied during the preprocessing or encoding process is applied next.
[0976] eoi_persistence_flag indicates the persistence of the optimization information provided in this SEI message. If eoi_persistence_flag is 0, it indicates that the optimization information is applied only to the current picture. If eoi_persistence_flag is 1, it indicates that the optimization information is applied to the current picture and all subsequent pictures in the current layer in output order until one or more of the following conditions are met.
[0977] - A new CLVS of the current layer starts.
[0978] - The bitstream ends.
[0979] - The picture of the current layer associated with the encoder optimization information SEI message is output after the current picture in the output order.
[0980] If eoi_for_human_viewing_idc is 3, it indicates that human viewing is included in the purpose of the applied optimization. If eoi_for_human_viewing_idc is 2, it indicates that the video is suitable but not specifically optimized for human viewing. If eoi_for_human_viewing_idc is 1, it indicates that the video is not suitable for human viewing. If eoi_for_human_viewing_idc is 0, it indicates that it is unknown whether the video is suitable for human viewing.
[0981] If eoi_for_machine_analysis_idc is 3, it indicates that machine analysis is included in the purpose of the applied optimization. If eoi_for_machine_analysis_idc is 2, it indicates that the video is suitable but not specifically optimized for machine analysis. If eoi_for_machine_analysis_idc is 1, it indicates that the video is not suitable for machine analysis. If eoi_for_machine_analysis_idc is 0, it indicates that it is unknown whether the video is suitable for machine analysis.
[0982] As a requirement for bitstream conformity, the values of eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc are not both 1. eoi_type indicates the type of optimization method specified in Table x1. Here, if (eoi_type & bitMask) is not 0, it indicates that an optimization type using the bitMask value from Table 11 (Definition of eoi_type) has been applied. If eoi_type is greater than 0 and (eoi_type & bitMask) is 0, an optimization type using the bitMask value is not applied. If eoi_type is 0, the optimization determined by the application is used.
[0983] [Table 11]
[0984]
[0985] EoiTemporalQualityFlag, EoiSpatialQualityFlag, and EoiPrivacyProtectionFlag specify whether eoi_type represents an optimization type including object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, and privacy protection optimization, and are derived as follows.
[0986] EoiObjectBasedFlag = ( ( eoi_type & 0x01 ) > 0 ) ? 1:0
[0987] EoiTemporalResamplingFlag = ( ( eoi_type & 0x02 ) > 0 ) ? 1:0
[0988] EoiSpatialResamplingFlag = ( ( eoi_type & 0x04 ) > 0 ) ? 1 : 0 (xx)
[0989] EoiTemporalQualityFlag = ( ( eoi_type & 0x08 ) > 0 ) ? 1:0
[0990] EoiSpatialQualityFlag = ( (eoi_type & 0x10 ) > 0 ) ? 1:0
[0991] EoiPrivacyProtectionFlag = ( (eoi_type & 0x20 ) > 0 ) ? 1:0
[0992] For example, if a specific top-temporal sublayer is encoded with coarse quantization that makes quality variation unpleasant for human viewers but does not affect machine task performance, you can set eoi_for_human_viewing_flag and eoi_for_machine_analaysis_flag to 0 and 1, respectively, and set eoi_type to the value that sets EoiTemporalQualityFlag to 1.
[0993] If eoi_persistence_flag is 0, EoiTemporalResamplingFlag is 0 and EoiTemporalQualityFlag is 0 according to the bitstream conformity requirements.
[0994] eoi_object_based_idc, if present, represents the object-based optimization type specified in Table 12 (Definition of eoi_object_based_idc). Here, if (eoi_object_based_idc & bitMask) is not 0, it indicates that the object-based optimization type associated with the bitMask value in Table 12 has been applied. If eoi_object_based_idc is greater than 0 and (eoi_object_based_idc & bitMask) is 0, the object-based optimization type associated with the bitMask value has not been applied. If eoi_object_based_idc is 0, the object-based optimization type defined by the application is applied. The value of eoi_object_based_idc ranges from 0 to 7 in bitstreams compliant with this version of this specification. Values for eoi_object_based_idc from 8 to 65,535 are reserved by ITU-T for future use. Defined according to ISO / IEC and does not exist in bitstreams compliant with this version of this specification. If the eoi_object_based_idc value is between 8 and 65,535 (inclusive), decoders compliant with this version of this specification ignore eoi_object_based_idc.
[0995] [Table 12]
[0996]
[0997] If eoi_temporal_resampling_type_flag is 0, it indicates that temporal resampling optimization is a subsampling operation. If eoi_temporal_resampling_type_flag is 1, it indicates that temporal resampling optimization is an upsampling operation.
[0998] If eoi_num_int_pics is greater than 0, it indicates that the number of pictures excluded between each pair of coded pictures in the output order (when eoi_temporal_resampling_type_flag is 0) or added for encoding between each pair of source pictures within the persistence of this SEI message (when eoi_temporal_resampling_type_flag is 1) by the encoding system is constant. If eoi_temporal_resampling_type_flag is 0 and eoi_num_int_pics is greater than 0, eoi_num_int_pics indicates the number of pictures excluded between each pair of coded pictures in the output order by the encoding system. If eoi_temporal_resampling_type_flag is 1 and eoi_num_int_pics is greater than 0, eoi_num_int_pics indicates the number of pictures added between each pair of source pictures for encoding by the encoding system.
[0999] If eoi_num_int_pics is 0, it indicates that within the persistence of this SEI message, the number of pictures excluded between each pair of coded pictures by the encoding system in the output order (when eoi_temporal_resampling_type_flag is 0) or added between each pair of source pictures for encoding (when eoi_temporal_resampling_type_flag is 1) is unknown or variable.
[1000] The eoi_num_int_pics values are in the range from 0 to 63.
[1001] If eoi_spatial_resampling_type_flag is 0, it indicates that spatial resampling optimization is a subsampling operation. If eoi_spatial_resampling_type_flag is 1, it indicates that spatial resampling optimization is an upsampling operation.
[1002] If eoi_privacy_protection_type_idc is present, it indicates the privacy protection optimization type specified in Table 13 (Definition of eoi_privacy_protection_type_idc).
[1003] [Table 13]
[1004]
[1005] eoi_privacy_protected_info_type, if present, indicates the type of information protected as specified in Table 14(). Here, if eoi_privacy_protected_info_type is greater than 0 and (eoi_privacy_protected_info_type & bitMask) is not 0, it indicates that the information type with the bitMask value in Table 14 (Definition of eoi_privacy_protection_info_type) is protected. If eoi_privacy_protected_info_type is 0, the information of the type defined by the application is protected. The value of eoi_privacy_protection_info_type is in the range of 0 to 7 in bitstreams compliant with this version of this specification. Values of 8 to 255 for eoi_privacy_protected_info_type (inclusive) are reserved for future use by ITU-T | ISO / IEC and do not exist in bitstreams compliant with this version of this specification. If the eoi_privacy_protected_info_type value is in the range of 8 to 255, decoders compliant with this version of this specification ignore eoi_privacy_protected_info_type.
[1006] [Table 14]
[1007]
[1008] FIG. 34 shows the syntax of a text description information SEI message according to embodiments.
[1009] Text description information SEI messages provide text descriptions for one or more pictures.
[1010] txt_descr_id represents the identifier value of this text description information SEI message. The txt_descr_id value ranges from 1 to 16383. The value 0 is reserved.
[1011] If txt_cancel_flag is 1, it indicates that the persistence of the previous text description SEI message with the same txt_descr_id is canceled in the output order applied to the current layer. If txt_cancel_flag is 0, it indicates that the text description comes later.
[1012] txt_persistence_flag indicates the persistence of text information description SEI messages for the current layer.
[1013] If txt_persistence_flag is 0, it indicates that text description information is applied only to the currently decoded picture.
[1014] If txt_persistence_flag is 1, it indicates that the text description information SEI message is applied to the currently decoded picture and persists in output order for all subsequent pictures of the current layer until one or more of the following conditions are met.
[1015] - A new CLVS of the current layer starts.
[1016] - The bitstream ends.
[1017] - The picture in the current hierarchy of the AU associated with the text description information SEI message with the same txt_descr_id is output after the current picture in the output order.
[1018] txt_descr_purpose represents the purpose of the text description SEI specified in Table 15 (Definition of txt_descr_purpose). The value of text_descr_purpose is in the range from 0 to 5. Values for text_descr_purpose in the range from 6 to 255 are reserved for future use by ITU-T | ISO / IEC and do not exist in bitstreams compliant with this version of this specification. Decoders compliant with this version of this specification accept all text_descr_purpose values in the range from 0 to 255.
[1019] [Table 15]
[1020]
[1021] txt_num_strings_minus1 + 1 represents the number of items in txt_descr_string_lang[ i ] and the following txt_descr_string[ i ].
[1022] txt_descr_string_lang[ i ] represents the language of txt_descr_string[ i ]. The language of txt_descr_string[ i ] is specified by language tags defined in IETF RFC 5646. The length of txt_descr_string_lang[ i ] is from 0 to 49 (inclusive).
[1023] txt_descr_string[ i ] represents the i-th text description information string interpreted as the value specified in txt_descr_purpose.
[1024] If txt_descr_purpose is 0, the interpretation of the information contained in txt_descr_string is defined by the application.
[1025] If txt_descr_purpose is 1, txt_descr_string[ i ] represents copyright information related to a picture within the persistence range defined by txt_cancel_flag and txt_persistence_flag.
[1026] If txt_descr_purpose is 2, txt_descr_string[ i ] represents AI display information related to a picture within the persistence scope of this SEI message if it is not a null string.
[1027] Note: If txt_descr_purpose is 2, the string may contain information about machine learning-based processing, the use of the decoded picture, or other aspects related to the picture.
[1028] If txt_descr_purpose is 3, txt_descr_string[ i ] represents a plain text label description associated with a picture within the persistence range defined by txt_cancel_flag and txt_persistence_flag.
[1029] When txt_descr_purpose is 4, txt_descr_string[ i ] represents content advisory rating information compliant with the U.S. and Canadian Rating Region Tables (RRT) related to pictures within the persistence range defined by txt_cancel_flag and txt_persistence_flag.
[1030] When txt_descr_purpose is 5, txt_descr_string[ i ] contains a tag URI with the syntax and semantics specified in IETF RFC 4151, which identifies CLVS.
[1031] FIGS. 35a, 35b, 35c, and 35d show the syntax of a generative face video SEI message according to the embodiments.
[1032] A GFV (Generative Face Video) SEI message contains face parameters and represents a face parameter transformation network denoted as TranslatorNN(). This network can be used to convert face parameters of various formats contained in the SEI message into a specific face parameter format supported by the decoding system. A face picture generation neural network denoted as GenerativeNN() can be used to generate an output picture using the face parameters converted into a specific format and the previously decoded output picture.
[1033] If a picture unit contains a GFV SEI message with a specific gfv_id value and gfv_base_pic_flag of 1, the picture of that picture unit is called the base picture for that specific gfv_id value.
[1034] If a picture unit contains a GFV SEI message with a specific gfv_id value and gfv_base_pic_flag of 0, and the picture unit does not contain a GFV SEI message with the specific gfv_id value and gfv_base_pic_flag of 1, the picture of the picture unit is said to be the driving picture for the specific gfv_id value.
[1035] If a picture unit contains a GFV SEI message with a specific gfv_id value, gfv_base_pic_flag of 0, and gfv_drive_pic_fusion_flag of 1, and there is no GFV SEI message with the specific gfv_id value and gfv_base_pic_flag of 1 in the picture unit, the picture in the picture unit is said to be a fusion picture for the specific gfv_id value.
[1036] Note 1 - Face parameters can be checked in the source picture before encoding.
[1037] Note 2 - The previously decoded output picture input to GenerativeNN() can be a base picture (a decoded output picture that provides a reference texture to generate a face picture) and optionally a picture that can be fused by GenerativeNN() to enhance background textures and face details. If the current picture is not a base picture, a face picture can be generated based on the previously decoded base picture, face parameters passed through the GFV SEI message, and optionally the current decoded picture for fusion using a GFV SEI message.
[1038] To use this SEI message, define the following variables:
[1039] The width and height of the input and output images (pictures) in luma samples are indicated as CroppedWidth and CroppedHeight, respectively.
[1040] The chroma sample arrays baseCroppedYPic for the luminance sample array and baseCroppedCbPic and baseCroppedCrPic for the decoded output image correspond to the source base image and are denoted as BasePicture.
[1041] The lumina sample array driveCroppedYPic and the chroma sample arrays driveCroppedCbPic and driveCroppedCrPic for the decoded output image correspond to the source drive image and are denoted as DrivePicture.
[1042] Bit depth BitDepthY for the luminance sample arrays of the input and output images.
[1043] BitDepthC for the chroma sample array (if any) of the input and output images.
[1044] Chroma format indicator displayed as ChromaFormatIdc.
[1045] The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc specified in Table 2.
[1046] gfv_id contains an identification number that can be used to identify face feature information and specify a neural network that can be used as TranslatorNN(). gfv_id values range from 0 to 232-2. gfv_id values from 256 to 511 and gfv_id values from 231 to 232-2 are reserved for future use by ITU-T | ISO / IEC. Decoders compliant with this version of this document ignore GFV SEI messages when they encounter a message in which gfv_id is in the range of 256 to 511 or 231 to 232-2.
[1047] gfv_cnt represents the number of GFV SEI message instances for this gfv_id value within the picture unit.
[1048] The gfv_cnt of the first GFV SEI message in decoding order with a specific gfv_id value within a picture unit is equal to 0. If the gfv_cnt assigned to currGfvCnt is greater than 0, then a GFV SEI message with the same gfv_id value and a gfv_cnt equal to currGfvCnt - 1 must exist in the same picture unit and precede the current GFV SEI message in decoding order.
[1049] The gfv_cnt value is in the range from 0 to 65,535.
[1050] If gfv_base_pic_flag is 1, it indicates that the currently decoded output image corresponds to the base image. If gfv_base_pic_flag is 0, it indicates that the currently decoded output image does not correspond to the base image or that this SEI message does not specify a syntax element of the base image. If gfv_base_pic_flag is absent, it is inferred to be 0.
[1051] If the GFV SEI message is the first GFV SEI message with a specific gfv_id value within the current CLVS in the decoding order, the gfv_base_pic_flag value is equal to 1.
[1052] If the gfv_base_pic_flag of a GFV SEI message with a specific gfv_id value is 1, the base picture for that specific gfv_id value (the currently cropped decoded picture) comes after the currently decoded picture in output order for the currently decoded picture and all subsequent decoded pictures of the current layer, either up to the end of the current CLVS or excluding the decoded picture within the current CLVS, and is associated with the GFV SEI message with that specific gfv_id value and gfv_base_pic_flag 1 (whichever is earlier).
[1053] If gfv_nn_present_flag is 1, it indicates that a neural network that can be used as TranslatorNN() is included in or specified in the SEI message. If gfv_nn_present_flag is 0, it indicates that a neural network that can be used as TranslatorNN() is not included in or specified in the SEI message. If gfv_nn_present_flag is not present, it is inferred to be 0.
[1054] If a GFV SEI message with gfv_cnt of 0 exists in the first picture unit of CLVS in the decoding order, gfv_nn_present_flag exists and is equal to 1.
[1055] If gfv_nn_present_flag is 0 and TranslatorNN is referenced in the semantics of a GFV SEI message, the following constraints apply.
[1056] - When gfv_cnt is 0, there must be at least one GFV SEI message in the previous picture unit in the output order of the current CLVS, and this message has the same gfv_id value as the current GFV SEI message and gfv_nn_present_flag is equal to 1.
[1057] - Otherwise (when gfv_cnt is greater than 0), there must be at least one GFV SEI message in the current CLVS that exists in the current picture unit or the previous picture unit in output order, and has the same gfv_id value as the current GFV SEI message and a gfv_nn_present_flag value equal to 1.
[1058] If gfv_nn_present_flag is 0 and TranslatorNN is referenced in the semantics of this SEI message, the following is applied to derive an applicable TranslatorNN.
[1059] - If gfv_cnt is greater than 0 and there is one or more previous GFV SEI messages with gfv_nn_present_flag of 1 that have the same gfv_id value as the current GFV SEI message in decoding order at the current picture level, the applicable TranslatorNN is defined as the last previous GFV SEI message with gfv_nn_present_flag of 1 that has the same gfv_id value as the current GFV SEI message in decoding order at the current picture level.
[1060] - Otherwise, the applicable TranslatorNN is defined as the GFV SEI message existing in the picture unit puB prior to the last in output order in the current CLVS, which has the same gfv_id value as the current GFV SEI message and gfv_nn_present_flag as 1. If there are multiple such GFV SEI messages in the picture unit puB that have the same gfv_id value as the current GFV SEI message and gfv_nn_present_flag as 1, the TranslatorNN is defined as the last of these GFV SEI messages in the decoding order.
[1061] gfv_nn_base_flag, gfv_nn_mode_idc, gfv_nn_alignment_zero_bit_a, gfv_nn_tag_uri, gfv_nn_uri, gfv_nn_alignment_zero_bit_b, and gfv_nn_payload_byte[ i ] represent neural networks that can be used as TranslatorNN(). gfv_nn_base_flag, gfv_nn_mode_idc, gfv_nn_alignment_zero_bit_a, gfv_nn_tag_uri, gfv_nn_uri, gfv_nn_alignment_zero_bit_b, and gfv_nn_payload_byte[ i ] have the same syntax and semantics as nnpfc_base_flag, nnpfc_mode_idc, nnpfc_alignment_zero_bit_a, nnpfc_tag_uri, nnpfc_uri, nnpfc_alignment_zero_bit_b, and nnpfc_payload_byte[ i ], respectively.
[1062] If one of the following conditions is true, the GFV SEI message has the same SEI payload content.
[1063] - GFV SEI messages exist in the same picture unit, gfv_cnt is 0, gfv_nn_base_flag exists, and gfv_id and gfv_nn_base_flag values are the same.
[1064] - GFV SEI messages exist in the same picture unit, gfv_cnt is greater than 0, and gfv_id is the same.
[1065] If gfv_chroma_key_info_present_flag is 1, it indicates that syntax elements gfv_chroma_key_value_present_flag[ c ] and gfv_chroma_key_thr_present_flag[ i ] exist and that syntax elements gfv_chroma_key_value[ c ] and gfv_chroma_key_thr_value[ i ] may exist. If gfv_chroma_key_info_present_flag is 0, it indicates that syntax elements gfv_chroma_key_value_present_flag[ c ], gfv_chroma_key_thr_present_flag[ i ], gfv_chroma_key_value[ c ], and gfv_chroma_key_thr_value[ i ] do not exist.
[1066] If gfv_chroma_key_value_present_flag[ c ] is 1, it indicates that the syntax element gfv_chroma_key_value[ c ] exists. If gfv_chroma_key_present_flag[ c ] is 0, it indicates that the syntax element gfv_chroma_key_value[ c ] does not exist.
[1067] The ChromaKeyDefaultValueFlag variable is set to !(gfv_chroma_key_value_present_flag
[0000] | | gfv_chroma_key_value_present_flag
[0001] | | gfv_chroma_key_value_present_flag
[0002] ).
[1068] gfv_chroma_key_value[ c ] represents the chroma key value corresponding to the c-th color component as follows.
[1069] - If ChromaKeyDefaultValueFlag is 1, the variable GfvChromaKeyValue[ c ] is specified as follows.
[1070] - GfvChromaKeyValue
[0000] is set to 50.
[1071] - GfvChromaKeyValue
[0001] is set to 220.
[1072] - GfvChromaKeyValue
[0002] is set to 100.
[1073] - Otherwise, if ChromaKeyDefaultValueFlag is 0, the following applies.
[1074] - If gfv_chroma_key_value_present_flag[ c ] is 1, GfvChromaKeyValue[ c ] is set to be the same as the value of gfv_chroma_key_value[ c ].
[1075] - Otherwise, if gfv_chroma_key_value_present_flag[ c ] is 0, GfvChromaKeyValue[ c ] is not specified in this specification.
[1076] If gfv_chroma_key_thr_present_flag[ i ] is 1, it indicates that the syntax element gfv_chroma_thr_value[ i ] exists. If gfv_chroma_key_thr_present_flag[ i ] is 0, it indicates that gfv_chroma_key_thr_value[ i ] does not exist.
[1077] If gfv_chroma_key_thr_value[ i ] exists, it represents the i-th chroma key threshold. If it does not exist, the value of gfv_chroma_key_thr_value[ i ] is inferred as follows.
[1078] If i is 0, gfv_chroma_key_thr_value
[0000] is set to 48.
[1079] Otherwise, if i is 1, gfv_chroma_key_thr_value
[0001] is set to 75.
[1080] Note 3 - The syntax elements gfv_chroma_key_value_present[ c ], gfv_chroma_key_value[ c ], and gfv_chroma_key_thr_value[ i ] can be used to determine the transparency indicator for the fusion of the generated face picture and background picture. For example, the transparency indicator (alpha[ x ][ y ]) for the picture sample value is denoted as I[ c ][ x ][ y ], bitDepth[ c ] is such that bitDepth
[0000] is equal to BitDepthY, bitDepth
[0001] and bitDepth
[0002] are equal to BitDepthC, and the GfvChromaKey[ c ] value for the sample coordinates x, y and color component c can be determined as follows.
[1081] d[ x ][ y ]= 0
[1082] for( c = 0; c < 3; c++ )
[1083] if(gfv_chroma_key_value_present[ c ] | | ChromaKeyDefaultValueFlag )
[1084] d[ x ][ y ] += ( I[ c ][ x ][ y ] / ( 1 << (bitDepth[ c ] - 8) ) - GfvChromaKeyValue[ c ] )2
[1085] if( ( d[ x ][ y ] < gfv_chroma_key_thr_value
[0000] )
[1086] alpha[ x ][ y ] = 0
[1087] else if ( ( d[ x ][ y ] > gfv_chroma_key_thr_value
[0001] )
[1088] alpha[ x ][ y ] = 1
[1089] else
[1090] alpha[ x ][ y ] = ( d[ x ][ y ] - gfv_chroma_key_thr_value
[0000] ) χ
[1091] ( gfv_chroma_key_thr_value
[0001] - gfv_chroma_key_thr_value
[0000] )
[1092] If the value of alpha[ x ][ y ] is 0, it indicates transparency. If the value of alpha[ x ][ y ] is 1, it indicates opacity. Intermediate values of alpha[ x ][ y ] indicate translucency.
[1093] If gfv_drive_pic_fusion_flag is present, if it is 1, it indicates that the currently decoded image corresponding to the driving image that can be used for fusion can be input to GenerativeNN(). If gfv_drive_pic_fusion_flag is 0, it indicates that the currently decoded image should not be input to GenerativeNN().
[1094] Note 4 - If the gfv_drive_pic_fusion_flag value is 1, it may indicate that, for example, the currently decoded image can be used to enhance facial details or process background changes.
[1095] Note 5 - When gfv_base_pic_flag is 0 and gfv_drive_pic_fusion_flag is 1, the GFV process receives three inputs: the base image, keypoint features, and / or matrices included in the GFV SEI message. Then, it outputs the image generated by GenerativeNN(), which is the currently decoded image, which is the fusion image.
[1096] Note 6 - When gfv_base_pic_flag is 0 and gfv_drive_pic_fusion_flag is 0, the GFV process takes two inputs, the base image and the keypoint and / or matrix features passed in the GFV SEI message, and outputs an image generated by GenerativeNN().
[1097] Note 7 - When gfv_base_pic_flag is 1, the GFV process outputs the truncated decoded image directly.
[1098] If gfv_base_pic_flag of the GFV SEI message is 0 and gfv_drive_pic_fusion_flag is 0, the GFV SEI message applies only to the currently decoded picture.
[1099] If the gfv_base_pic_flag of a GFV SEI message with a specific gfv_id value is 0 and the gfv_drive_pic_fusion_flag is 1, the fusion picture for that gfv_id value (the currently cropped decoded picture) comes after the currently decoded picture in output order, excluding decoded pictures within the current CLVS or up to the end of the current CLVS, for the currently decoded picture and all subsequent decoded pictures of the current layer, and is associated with the GFV SEI message with that gfv_id value (whichever is earlier).
[1100] When gfv_cnt of GFV SEI message gfvSeiA with a specific gfv_id value is greater than 0, and gfv_base_pic_flag of GFV SEI message gfvSeiB with the same gfv_id value in the same picture unit is 1 (i.e., when the currently decoded picture is the base picture), gfv_drive_pic_fusion_flag of GFV SEI message gfvSeiA becomes 0.
[1101] If gfv_low_confidence_face_parameter_flag is 1, it indicates that the face parameter was derived with low confidence. If gfv_low_confidence_face_parameter_flag is 0, it indicates that confidence information for the face parameter is not specified.
[1102] If gfv_coordinate_present_flag is 1, it indicates that there is coordinate information for the keypoint. If gfv_coordinate_present_flag is 0, it indicates that there is no coordinate information for the keypoint.
[1103] The bitstream conformance requirement is that for all i from 0 to gfv_num_matrix_types_minus1, the value of gfv_coordinate_present_flag is 1 when gfv_matrix_type_idx[ i ] is 0 or 1.
[1104] If gfv_kps_pred_flag is 1, it indicates that syntax elements gfv_coordinate_dx_abs[ i ], gfv_coordinate_dy_abs[ i ] and gfv_coordinate_dz_abs[ i ] exist and syntax elements gfv_coordinate_dx_sign_flag[ i ], gfv_coordinate_dy_sign_flag[ i ] and gfv_coordinate_dz_sign_flag[ i ] may exist. If gfv_kps_pred_flag is 0, it indicates that syntax elements gfv_coordinate_x_abs[ i ], gfv_coordinate_y_abs[ i ], gfv_coordinate_z_abs[ i ] exist and syntax elements gfv_coordinate_x_sign_flag[ i ], gfv_coordinate_y_sign_flag[ i ], gfv_coordinate_z_sign_flag[ i ] may exist.
[1105] If gfv_coordinate_present_flag is 1, gfv_base_pic_flag is 0, and gfv_kps_pred_flag is 1, then there is a previous GFV SEI message in the current CLVS with gfv_base_pic_flag 1 that has the same gfv_id as the current GFV SEI message in the decoding order.
[1106] gfv_coordinate_precision_factor_minus1 + 1 represents the precision of the key point coordinates displayed in the SEI message. The value of gfv_coordinate_precision_factor_minus1 ranges from 0 to 31. When gfv_coordinate_present_flag is 1, gfv_base_pic_flag is 0, and gfv_kps_pred_flag is 1, the value of gfv_coordinate_precision_factor_minus1 is inferred to be the same as the gfv_coordinate_precision_factor_minus1 of the previous GFV SEI message that has the same gfv_id as the current GFV SEI message in the decoding order and gfv_base_pic_flag is 1.
[1107] gfv_num_kps_minus1 + 1 represents the number of key points. The value of gfv_num_kps_minus1 is in the range from 0 to 210 - 1. When gfv_coordinate_present_flag is 1, gfv_base_pic_flag is 0, and gfv_kps_pred_flag is 1, the value of gfv_num_kps_minus1 is inferred to be the same as the gfv_num_kps_minus1 of the previous GFV SEI message that has the same gfv_id as the current GFV SEI message in the decoding order and gfv_base_pic_flag is 1.
[1108] If gfv_coordinate_z_present_flag is 1, it indicates that there is z-axis coordinate information for the keypoint. If gfv_coordinate_z_present_flag is 0, it indicates that there is no z-axis coordinate information for the keypoint. If gfv_coordinate_present_flag is 1, gfv_base_pic_flag is 0, and gfv_kps_pred_flag is 1, the coordinate_z_present_flag value is inferred to be the same as the coordinate_z_present_flag of the previous GFV SEI message that has the same gfv_id as the current GFV SEI message and gfv_base_pic_flag is 1 in the decoding order.
[1109] gfv_coordinate_z_max_value_minus1 plus 1 represents the maximum absolute value of the keypoint's z-axis coordinate. The value of gfv_coordinate_z_max_value_minus1 is in the range from 0 to 216-1. When gfv_coordinate_present_flag is 1, gfv_base_pic_flag is 0, and gfv_kps_pred_flag is 1, the value of gfv_coordinate_z_max_value_minus1 is inferred to be the same as when the previous GFV SEI message in the decoding order has gfv_id and gfv_base_pic_flag is 1 and gfv_coordinate_z_max_value_minus1 is 1.
[1110] gfv_coordinate_x_abs[ i ] is used to derive the x-axis coordinates of the i-th keypoint.
[1111] gfv_coordinate_x_sign_flag[ i ] specifies the sign of the x-axis coordinate of the i-th keypoint. If gfv_coordinate_x_sign_flag[ i ] is not present, it is inferred to be 0.
[1112] gfv_coordinate_y_abs[ i ] is used to derive the y-axis coordinates of the i-th keypoint.
[1113] gfv_coordinate_y_sign_flag[ i ] specifies the sign of the y-axis coordinate of the i-th keypoint. If gfv_coordinate_y_sign_flag[ i ] is not present, it is inferred to be 0.
[1114] gfv_coordinate_z_abs[ i ] is used to derive the z-axis coordinates of the i-th keypoint.
[1115] gfv_coordinate_z_sign_flag[ i ] specifies the sign of the z-axis coordinate of the i-th keypoint. If gfv_coordinate_z_sign_flag[ i ] is not present, it is inferred to be 0.
[1116] gfv_coordinate_dx_abs[ i ] specifies the difference value used to derive the x-axis coordinate of the i-th keypoint.
[1117] gfv_coordinate_dx_sign_flag[ i ] specifies the sign of the x-axis coordinate difference value of the i-th keypoint. If gfv_coordinate_dx_sign_flag[ i ] is not present, it is inferred to be 0.
[1118] gfv_coordinate_dy_abs[ i ] specifies the difference value used to derive the y-axis coordinate of the i-th keypoint.
[1119] gfv_coordinate_dy_sign_flag[ i ] specifies the sign of the y-axis coordinate difference value of the i-th keypoint. If gfv_coordinate_yd_sign_flag[i] is not present, it is inferred to be 0.
[1120] gfv_coordinate_dz_abs[ i ] specifies the difference value used to derive the z-axis coordinate of the i-th keypoint.
[1121] gfv_coordinate_dz_sign_flag[ i ] specifies the sign of the z-axis coordinate difference value of the i-th keypoint. If gfv_coordinate_dz_sign_flag[ i ] is not present, it is inferred to be 0.
[1122] If gfv_coordinate_z_max_value_minus1 exists, the CroppedDepth variable is set to gfv_coordinate_z_max_value_minus1 + 1. Otherwise, CroppedDepth is set to 0.
[1123] If gfv_kps_pred_flag is 1, the variables coordinateDeltaX[ i ], coordinateDeltaY[ i ], and coordinateDeltaZ[ i ], representing the delta x-axis, delta y-axis, and delta z-axis coordinates of the i-th keypoint respectively, are derived as follows:
[1124] coordinateDeltaX[ i ] = ( 1 - 2 * gfv_coordinate_dx_sign_flag[ i ] ) * gfv_coordinate_dx_abs[ i ] /
[1125] ( 1 << ( gfv_coordinate_precision_factor_minus1 + 1 ) )
[1126] coordinateDeltaY[ i ] = ( 1 - 2 * gfv_coordinate_dy_sign_flag [ i ] ) * gfv_coordinate_dy_abs[ i ] χ
[1127] ( 1 << ( gfv_coordinate_precision_factor_minus1 + 1 ) )
[1128] if(gfv_coordinate_z_present_flag)
[1129] coordinateDeltaZ[ i ] = ( 1 - 2 * gfv_coordinate_dz_sign_flag[ i ] ) * gfv_coordinate_dz_abs[ i ] /
[1130] ( 1 << ( gfv_coordinate_precision_factor_minus1 + 1 ) )
[1131] The variables coordinateX[ i ], coordinateY[ i ], and, when gfv_coordinate_z_present_flag is equal to 1, coordinateZ[ i ] indicating the x-axis coordinate, y-axis coordinate and z-axis coordinate of the i-th keypoint, respectively, are derived as follows:
[1132] If gfv_kps_pred_flag is equal to 0, the following applies:
[1133] coordinateX[ i ] = ( 1 - 2 * gfv_coordinate_x_sign_flag[ i ] ) * gfv_coordinate_x_abs[ i ] /
[1134] ( 1 << ( gfv_coordinate_precision_factor_minus1 + 1 ) )
[1135] coordinateY[ i ] = ( 1 - 2 * gfv_coordinate_y_sign_flag[ i ] ) * gfv_coordinate_y_abs[ i ] /
[1136] ( 1 << ( gfv_coordinate_precision_factor_minus1 + 1 ) )
[1137] if (gfv_coordinate_z_present_flag )
[1138] coordinateZ[ i ] = ( 1 - 2 * gfv_coordinate_z_sign_flag[ i ] ) * gfv_coordinate_z_abs[ i ] /
[1139] ( 1 << ( gfv_coordinate_precision_factor_minus1 + 1 ) )
[1140] Otherwise (gfv_kps_pred_flag is equal to 1), the following applies:
[1141] if( gfv_base_pic_flag ) {
[1142] coordinateX[ i ] = (( i > 0 ) ? coordinateX[ i - 1 ] : 0 ) + coordinateDeltaX[ i ]
[1143] coordinateY[ i ] = (( i > 0 ) ? coordinateY[ i - 1 ] : 0 ) + coordinateDeltaY[ i ]
[1144] if (gfv_coordinate_z_present_flag )
[1145] coordinateZ[ i ] = (( i > 0 ) ? coordinateZ[ i - 1 ] : 0 ) + coordinateDeltaZ[ i ]
[1146] } else if( gfv_cnt = = 0 ) {
[1147] coordinateX[ i ] = BaseKpCoordinateX[ i ] + coordinateDeltaX[ i ]
[1148] coordinateY[ i ] = BaseKpCoordinateY[ i ] + coordinateDeltaY[ i ]
[1149] if (gfv_coordinate_z_present_flag )
[1150] coordinateZ[ i ] = BaseKpCoordinateZ[ i ] + coordinateDeltaZ[ i ]
[1151] } else {
[1152] coordinateX[ i ] = PrevKpCoordinateX[ i ] + coordinateDeltaX[ i ]
[1153] coordinateY[ i ] = PrevKpCoordinateY[ i ] + coordinateDeltaY[ i ]
[1154] coordinateZ[ i ] = PrevKpCoordinateZ[ i ] + coordinateDeltaZ[ i ]
[1155] }
[1156] The following applies for derivation of the variables BaseKpCoordinateX[ i ], BaseKpCoordinateY[ i ], BaseKpCoordinateZ[ i ], PrevKpCoordinateX[ i ], PrevKpCoordinateY[ i ], and PrevKpCoordinateZ[ i ]:
[1157] if( gfv_base_pic_flag ) {
[1158] PrevKpCoordinateX[ i ] = BaseKpCoordinateX[ i ] = coordinateX[ i ]
[1159] PrevKpCoordinateY[ i ] = BaseKpCoordinateY[ i ] = coordinateY[ i ]
[1160] if (gfv_coordinate_z_present_flag )
[1161] PrevKpCoordinateZ[ i ] = BaseKpCoordinateZ[ i ] = coordinateZ[ i ]
[1162] } else {
[1163] PrevKpCoordinateX[i] = coordinateX[i]
[1164] PrevKpCoordinateY[i] = coordinateY[i]
[1165] PrevKpCoordinateZ[i] = coordinateZ[i]
[1166] }
[1167] If gfv_matrix_present_flag is 1, it indicates that there are matrix parameters. If gfv_matrix_present_flag is 0, it indicates that there are no matrix parameters. If gfv_coordinate_present_flag is 0, gfv_matrix_present_flag is equal to 1.
[1168] If gfv_matrix_pred_flag is 1, it indicates that syntax elements gfv_matrix_element_int[ i ][ j ][ k ][ m ] and gfv_matrix_element_dec[ i ][ j ][ k ][ m ] exist and that syntax element gfv_matrix_element_sign_flag[ i ][ j ][ k ][ m ] may exist. If gfv_matrix_pred_flag is 0, it indicates that syntax elements gfv_matrix_delta_element_int[ i ][ j ][ k ][ m ] and gfv_matrix_delta_element_dec[ i ][ j ][ k ][ m ] exist and that syntax element gfv_matrix_delta_element_sign_flag[ i ][ j ][ k ][ m ] may exist. If gfv_matrix_pred_flag is missing, it is inferred as 0.
[1169] If gfv_matrix_present_flag is 1, gfv_base_pic_flag is 0, and gfv_matrix_pred_flag is 1, then there is a previous GFV SEI message in the current CLVS with gfv_base_pic_flag 1 that has the same gfv_id as the current GFV SEI message in the decoding order.
[1170] gfv_matrix_element_precision_factor_minus1 plus 1 indicates the precision of the matrix element signaled in the SEI message. The value of gfv_matrix_element_precision_factor_minus1 ranges from 0 to 31. When gfv_matrix_present_flag is 1, gfv_base_pic_flag is 0, and gfv_matrix_pred_flag is 1, the value of gfv_matrix_element_precision_factor_minus1 is inferred to be the same as the gfv_matrix_element_precision_factor_minus1 of the previous GFV SEI message that has the same gfv_id as the current GFV SEI message in decoding order and gfv_base_pic_flag is 1.
[1171] The value of gfv_num_matrix_types_minus1 plus 1 indicates the number of matrix types signaled in the SEI message. The value of gfv_num_matrix_types_minus1 is in the range from 0 to 26 - 1. The bitstream conformance requirement is that when gfv_matrix_pred_flag is 1 and gfv_base_pic_flag is 0, the value of gfv_num_matrix_types_minus1 must be equal to the value of gfv_num_matrix_types_minus1 in each of the previous GFV SEI messages in the current CLVS, which has the same gfv_id value as the gfv_id value of the current SEI and gfv_base_pic_flag is 1, in the decoding order. When gfv_matrix_present_flag is 1, gfv_base_pic_flag is 0, and gfv_matrix_pred_flag is 1, the value of gfv_matrix_type_num_minus1 is inferred to be as follows: it is the gfv_matrix_type_num_minus1 of the previous GFV SEI message in the decoding order, uses the same gfv_id as the current GFV SEI message, and gfv_base_pic_flag is 1.
[1172] gfv_matrix_type_idx[ i ] represents the index of the i-th matrix type specified in the table below. The value of gfv_matrix_type_idx[ i ] is in the range of 0 to 63 (inclusive). In bitstreams compliant with this version of this specification, the value of gfv_matrix_type_idx[ i ] is in the range of 0 to 31 (inclusive). Decoders compliant with this version of this specification must allow gfv_matrix_type_idx[ i ] to be greater than 31 to appear in the bitstream, and the decoder ignores all information for the i-th type matrix where gfv_matrix_type_idx[ i ] is greater than 31.
[1173] Specification of gfv_matrix_type_idx[ i ]
[1174] [Table 16]
[1175]
[1176] Note 8 - The undefined matrxi type is used to represent matrxi types instead of affine translation matrices, covariance matrices, rotation matrices, translation matrices, and compact feature matrices. Users can use this to extend matrix types.
[1177] If gfv_num_matrices_equal_to_num_kps_flag[ i ] ] is 1, it indicates that the number of matrices of the i-th matrix type is equal to gfv_num_kps_minus1 + 1. If gfv_num_matrices_equal_to_num_kps_flag[ i ] ] ] ] ] ] ] ] ] indicates that the number of matrices of the i-th matrix type is not equal to gfv_num_kps_minus1 + 1. If gfv_matrix_present_flag is 1, gfv_base_pic_flag is 0, gfv_matrix_pred_flag is 1, gfv_matrix_type_idx[ i ] is 0 or 1, and gfv_coordinate_present_flag is 1, then the next value gfv_num_matrices_equal_to_num_kps_flag[ i ] is inferred to be the same as gfv_num_matrices_equal_to_num_kps_flag[ i ] in the previous GFV SEI message in the decoding order, gfv_id is the same as the current GFV SEI message, and gfv_base_pic_flag is 1. Otherwise, if gfv_num_matrices_equal_to_num_kps_flag[ i ] is not present, the value is inferred to be 0.
[1178] gfv_num_matrices_info[ i ] provides information for deriving the number of matrices of the i-th matrix type. The value of gfv_num_matrices_info[ i ] is in the range from 0 to 210-1. If gfv_matrix_present_flag is 1, gfv_base_pic_flag is 0, gfv_matrix_pred_flag is 1, gfv_matrix_type_idx[ i ] is 0 or 1, and gfv_coordinate_present_flag is 0 or gfv_num_matrix_equal_to_num_kps_flag[ i ] is 0, then the value of gfv_num_matrices_info[ i ] is inferred to be the same as gfv_num_matrices_info[ i ] in the previous GFV SEI message in the decoding order, and gfv_base_pic_flag is 1, having the same gfv_id as the current GFV SEI message.
[1179] gfv_matrix_width_minus1[ i ] plus 1 represents the matrix width of the i-th matrix type. The value of gfv_matrix_width_minus1[ i ] is in the range from 0 to 2¹⁰-1 (inclusive). If gfv_matrix_present_flag is 1, gfv_matrix_pred_flag is 0, gfv_matrix_pred_flag is 1, and gfv_matrix_type_idx[ i ] is 2 or 3 or greater than or equal to 7, the value of gfv_matrix_width_minus1[ i ] is inferred to be the same as gfv_matrix_width_minus1[ i ] in the previous GFV SEI message in the decoding order, and gfv_id and gfv_base_pic_flag are 1, which are the same as the current GFV SEI message.
[1180] The value obtained by adding 1 to gfv_matrix_height_minus1[ i ] represents the matrix height of the i-th matrix type. The value of gfv_matrix_height_minus1[ i ] is in the range from 0 to 210-1. If gfv_matrix_present_flag is 1, gfv_base_pic_flag is 0, gfv_matrix_pred_flag is 1, and gfv_matrix_type_idx[ i ] is 2 or 3 or greater than or equal to 7, then the value of gfv_matrix_height_minus1[ i ] is inferred to be the same as gfv_matrix_height_minus1[ i ] in the previous GFV SEI message in the decoding order, has the same gfv_id as the current GFV SEI message, and gfv_base_pic_flag is 1.
[1181] If gfv_matrix_for_3D_space_flag[ i ] is 1, it indicates that the matrix of type i is a matrix defined in 3D space. If gfv_matrix_for_3D_space_flag[ i ] is 0, it indicates that the matrix of type i is a matrix defined in 2D space. If gfv_matrix_present_flag is 1, gfv_base_pic_flag is 0, gfv_matrix_pred_flag is 1, gfv_matrix_type_idx[ i ] is 4, 5 or 6, and gfv_coordinate_present_flag is 0, the value of gfv_matrix_for_3D_space_flag[ i ] is inferred to be the same as gfv_matrix_for_3D_space_flag[ i ] (if present) in the previous GFV SEI message in the decoding order, and gfv_base_pic_flag is 1, having the same gfv_id as the current GFV SEI message.
[1182] If gfv_matrix_width_minus1[ i ] is missing, it is inferred as follows.
[1183] If gfv_matrix_type_idx[ i ] is 0, 1 or 4, and either coordinate_z_present_flag or gfv_matrix_for_3D_space_flag[ i ] exists and is equal to 1, then gfv_matrix_width_minus1[ i ] is inferred to be equal to 2.
[1184] Otherwise, if gfv_matrix_type_idx[ i ] is 0, 1 or 4, and either coordinate_z_present_flag or gfv_matrix_for_3D_space_flag[ i ] exists and is equal to 0, then gfv_matrix_width_minus1[ i ] is inferred to be equal to 1.
[1185] Otherwise (if gfv_matrix_type_idx[ i ] is equal to 5 or 6), it is inferred to be equal to gfv_matrix_width_minus1[ i ] 0.
[1186] If gfv_matrix_height_minus1[ i ] is missing, it is inferred as follows.
[1187] If gfv_matrix_type_idx is 0, 1, 4, 5 or 6 and one of gfv_coordinate_z_present_flag and gfv_matrix_for_3D_space_flag[ i ] exists and is equal to 1, then gfv_matrix_height_minus1[ i ] is inferred to be equal to 2.
[1188] Otherwise (where gfv_matrix_type_idx is 0, 1, 4, 5, or 6 and either gfv_coordinate_z_present_flag or gfv_matrix_for_3D_space_flag[ i ] is 0), gfv_matrix_height is inferred to be equal to 1. The variables matrixWidth[ i ] and matrixHeight[ i ], representing the width and height of a matrix of the i-th matrix type, are derived as follows.
[1189] if( gfv_matrix_pred_flag ) {
[1190] matrixWidth[i] = BaseMatrixWidth[i]
[1191] matrixHeight[i] = BaseMatrixHeight[i]
[1192] } else {
[1193] matrixWidth[i] = gfv_matrix_width_minus1[i] + 1
[1194] matrixHeight[i] = gfv_matrix_height_minus1[i] + 1
[1195] }
[1196] if( gfv_base_pic_flag ) {
[1197] BaseMatrixWidth[i] = matrixWidth[i]
[1198] BaseMatrixHeight[i] = matrixHeight[i]
[1199] }
[1200] The value obtained by adding 1 to gfv_num_matrices_minus1[ i ] represents the number of matrices of the i-th matrix type. The value of gfv_num_matrices_minus1[ i ] ranges from 0 to 2¹⁰⁴ 1. When gfv_matrix_present_flag is 1, gfv_base_pic_flag is 0, gfv_matrix_pred_flag is 1, and gfv_matrix_type_idx[ i ] is greater than or equal to 7, the value of gfv_num_matrices_minus1[ i ] is inferred to be the same as gfv_num_matrices_minus1[ i ] in the previous GFV SEI message in the decoding order. This message has the same gfv_id as the current GFV SEI message and gfv_base_pic_flag is 1.
[1201] The variable numMatrices[ i ], representing the number of matrices of the i-th matrix type, is derived as follows.
[1202] if( gfv_matrix_pred_flag )
[1203] numMatrices[i] = BaseNumMatrices[i]
[1204] else if( gfv_matrix_type_idx[ i ] = = 0 | | gfv_matrix_type_idx[ i ] = = 1 ) {
[1205] if(gfv_coordinate_present_flag)
[1206] numMatrices[ i ] = gfv_num_matrices_equal_to_num_kps_flag[ i ] ? gfv_num_kps_minus1 + 1:
[1207] ( gfv_num_matrices_info [ i ] < gfv_num_kps_minus1 ? gfv_num_matrices_info [ i ] + 1:
[1208] gfv_num_matrices_info [ i ] + 2 )
[1209] else
[1210] numMatrices[i] = gfv_num_matrices_info[i] + 1
[1211] } else if( gfv_matrix_type_idx[ i ] >= 2 && gfv_matrix_type_idx[ i ] < 7 )
[1212] numMatrices[ i ] = 1
[1213] else
[1214] numMatrices[ i ] = gfv_num_matrices_minus1[ i ] + 1
[1215] if( gfv_base_pic_flag )
[1216] BaseNumMatrices[i] = numMatrices[i]
[1217] When gfv_matrix_pred_flag is 1 and gfv_base_pic_flag is 0, the bitstream conformance requirement is that when i is in the range from 0 to gfv_num_matrix_types_minus1 (inclusive), the values of numMatrices[i], matrixWidth[i], and matrixHeight[i] must have the same gfv_id value as the gfv_id value of the current SEI, and when gfv_base_pic_flag is 1, the values of numMatrices[i], matrixWidth[i], and matrixHeight[i] must each be equal to the values of numMatrices[i], matrixWidth[i], and matrixHeight[i], respectively, in each previous GFV SEI message in the decoding order from the current CLVS where gfv_base_pic_flag is 1.
[1218] gfv_matrix_element_int[ i ][ j ][ k ][ m ] represents the integer part of the matrix element value at position (m, k) of the j-th matrix of type i. The values of gfv_matrix_element_int[ i ][ j ][ k ][ m ] range from 0 to 2³² - 2.
[1219] gfv_matrix_element_dec[ i ][ j ][ k ][ m ] represents the fractional part of the matrix element value at position (m, k) of the j-th matrix of the i-th matrix type. The length of gfv_matrix_element_dec[ i ][ j ][ k ][ m ] is gfv_matrix_element_precision_factor_minus1 + 1 bit.
[1220] gfv_matrix_element_sign_flag[ i ][ j ][ k ][ m ] represents the sign of the matrix element at position (m, k) of the j-th matrix of the i-th matrix type. If gfv_matrix_element_sign_flag[ i ][ j ][ k ][ m ] is not present, it is inferred to be 0.
[1221] gfv_matrix_delta_element_int[ i ][ j ][ k ][ m ] represents the integer part of the difference value of the matrix element at position (m, k) of the j-th matrix of the i-th matrix type.
[1222] gfv_matrix_delta_element_dec[ i ][ j ][ k ][ m ] represents the fractional part of the difference value of the matrix element at position (m, k) of the j-th matrix of the i-th matrix type.
[1223] gfv_matrix_delta_element_sign_flag[ i ][ j ][ k ][ m ] represents the sign of the difference value of the matrix element at position (m, k) of the j-th matrix of the i-th matrix type. If gfv_matrix_element_sign_flag[ i ][ j ][ k ][ m ] is not present, it is inferred to be 0.
[1224] When gfv_matrix_pred_flag is 1, the variable matrixElementDeltaVal[ i ][ j ][ k ][ m ], which represents the difference value of the matrix element at position (m, k) of the j-th matrix of the i-th matrix type, is derived as follows.
[1225] matrixElementDeltaVal[ i ][ j ][ k ][ m ] = (1-2*gfv_matrix_delta_element_sign_flag[ i ][ j ][ k ][ m ])*(gfv_matrix_delta_element_int[ i ][ j ][ k ][ m ]+(gfv_matrix_delta_element_dec[ i ][ j ][ k ][ m ]χ(1< <gfv_matrix_element_precision_factor_minus1+1))
[1226] The variable matrixElementVal[ i ][ j ][ k ][ m ], representing the value of the matrix element at position (m, k) of the j-th matrix of the i-th matrix type, is derived as follows.
[1227] If gfv_matrix_pred_flag is 0, the following applies.
[1228] matrixElementVal[ i][ j ][ k ][ m ] = (1-2*gfv_matrix_element_sign_flag[ i ][ j ][ k ][ m ])*(gfv_matrix_element_int[ i ][ j ][ k ][ m ]+(gfv_matrix_element_dec[ i ][ j ][ k ][ m ] / (1<< gfv_matrix_element_precision_factor_minus1+1))
[1229]
[1230] if( gfv_base_pic_flag )
[1231] BaseMatrixElementVal[ i][ j ][ k ][ m ] = matrixElementVal[ i][ j ][ k ][ m ]
[1232] Otherwise (gfv_matrix_pred_flag is equal to 1), the following applies:
[1233] if( gfv_cnt = = 0 )
[1234] matrixElementVal[ i][ j ][ k ][ m ] = BaseMatrixElementVal[ i][ j ][ k ][ m ] +
[1235] matrixElementDeltaVal[ i][ j ][ k ][ m ]
[1236] else
[1237] matrixElementVal[ i][ j ][ k ][ m ] = PrevMatrixElementVal[ i][ j ][ k ][ m ] +
[1238] matrixElementDeltaVal[ i][ j ][ k ][ m ]
[1239] The following applies:
[1240] if( gfv_base_pic_flag )
[1241] PrevMatrixElementVal[ i][ j ][ k ][ m ] = BaseMatrixElementVal[ i][ j ][ k ][ m ] =
[1242] matrixElementVal[i][j][k][m]
[1243] else
[1244] PrevMatrixElementVal[ i][ j ][ k ][ m ] = matrixElementVal[ i][ j ][ k ][ m ]
[1245] For a specific gfv_id value, the following process is used in the order of gfv_cnt to create a video picture in which gfv_base_pic_flag is 0 for each GFV SEI message and has a unique gfv_cnt value within the picture unit.
[1246] DeriveSigParam()
[1247] TranslatorNN(sigKeyPoint, sigMatrix)
[1248] DeriveInputTensors()
[1249] if( gfv_base_pic_flag = = 0 && gfv_drive_pic_fusion_flag = = 0 ) {
[1250] if( ChromaFormatIdc == 0 )
[1251] GenerativeNN( inputBaseY, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint,
[1252] inputDriveMatrix, CroppedWidth, CroppedHeight, CroppedDepth )
[1253] else
[1254] GenerativeNN( inputBaseY, inputBaseCb, inputBaseCr, inputBaseKeyPoint, inputBaseMatrix,
[1255] inputDriveKeyPoint, inputDriveMatrix, CroppedWidth, CroppedHeight, CroppedDepth )
[1256] } else if( gfv_base_pic_flag = = 0 && gfv_drive_pic_fusion_flag = = 1 ) {
[1257] if( ChromaFormatIdc = = 0 )
[1258] GenerativeNN( inputBaseY, inputDriveY, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint,
[1259] inputDriveMatrix, CroppedWidth, CroppedHeight, CroppedDepth )
[1260] else
[1261] GenerativeNN( inputBaseY, inputBaseCb, inputBaseCr, inputDriveY, inputDriveCb, inputDriveCr,
[1262] inputBaseKeyPoint, inputBaseMatrix,, inputDriveKeyPoint, inputDriveMatrix,
[1263] CroppedWidth, CroppedHeight, CroppedDepth )
[1264] }
[1265] StoreOutputTensors( )
[1266] The process DeriveSigParam( ) for deriving the inputs of TranslatorNN( ) is specified as follows:
[1267] The keypoint coordinate array sigKeyPoint and the matrix sigMatrix are derived as follows:
[1268] if( gfv_coordinate_present_flag )
[1269] for( i = 0; i <= gfv_num_kps_minus1; i++ ) {
[1270] sigKeyPoint[ i ]
[0000] = coordinateX[ i ]
[1271] sigKeyPoint[ i ]
[0001] = coordinateY[ i ]
[1272] if( gfv_coordinate_z_present_flag )
[1273] sigKeyPoint[ i ]
[0002] = coordinateZ[ i ]
[1274] }
[1275] else
[1276] for( i = 0; i <= gfv_num_kps_minus1; i++ ) {
[1277] sigKeyPoint[ i ]
[0000] = 0
[1278] sigKeyPoint[ i ]
[0001] = 0
[1279] if ( gfv_coordinate_z_present_flag )
[1280] sigKeyPoint[ i ]
[0002] = 0
[1281] }
[1282] if( gfv_matrix_present_flag )
[1283] for ( i = 0; i <= gfv_num_matrix_types_minus1; i++ )
[1284] for ( j = 0; j < numMatrices[ i ]; j++ )
[1285] for( k = 0; k < matrixHeight [ i ]; k++ )
[1286] for( l = 0;l < matrixWidth [ i ]; l++)
[1287] sigMatrix[ i ][ j ][ k ][ l ] = matrixElementVal[ i ][ j][ k][ l ]
[1288] else
[1289] for( i = 0; i <= gfv_num_matrix_types_minus1; i++ )
[1290] for ( j = 0; j < numMatrices[ i ]; j++ )
[1291] for( k = 0; k < matrixHeight [ i ]; k++ )
[1292] for( l = 0;l < matrixWidth [ i ]; l++)
[1293] sigMatrix[i][j][k][l] = 0
[1294] TranslatorNN() is a process that generates an output picture by converting various formats of face parameters included in SEI messages into fixed formats of face parameters that are input to the generation network.
[1295] The input to TranslatorNN() is as follows.
[1296] sigKeyPoint and sigMatrix
[1297] The output of TranslatorNN() is as follows.
[1298] convKeyPoint and convNumKeyPoint
[1299] convMatrix and convNumMatrix, convMatrixWidth, convMatrixHeight
[1300] The DeriveInputTensors() process that derives the input to GenerativeNN() is specified as follows.
[1301] When gfv_base_pic_flag is 1, the BasePicture input tensors inputBaseY, inputBaseCb, and inputBaseCr are derived as follows.
[1302] for(x = 0; x < CroppedWidth; x++ )
[1303] for ( y = 0; y < CroppedHeight; y++ )
[1304] inputBaseY[ x ][ y ] = InpY( baseCroppedYPic[ x ][ y ] )
[1305] if( ChromaFormatIdc != 0 )
[1306] for(x = 0; x < CroppedWidth / SubWidthC; x++ )
[1307] for ( y = 0; y < CroppedHeight / SubHeightC; y++ ) {
[1308] inputBaseCb[ x][ y ] = InpC( baseCroppedCbPic[ x][ y ] )
[1309] inputBaseCr[ x][ y ] = InpC( baseCroppedCrPic[ x ][ y ] )
[1310] }
[1311] When gfv_base_pic_flag is 0 and gfv_drive_pic_fusion_flag is 1, the DrivePicture luminance sample arrays inputDriveY, inputDriveCb, and input DriveCr are derived as follows.
[1312] for(x = 0; x< CroppedWidth; x++ )
[1313] for ( y = 0; y < CroppedHeight; y++ )
[1314] inputDriveY[x][y] = InpY( driveCroppedYPic[x][y])
[1315] if( ChromaFormatIdc != 0 )
[1316] for(x = 0; x< CroppedWidth / SubWidthC; x++ )
[1317] for( y = 0; y < CroppedHeight / SubHeightC; y++ ) {
[1318] InputDriveCb[ x][ y ] = InpC( driveCroppedCbPic[ x][ y ] )
[1319] InputDriveCr[ x ][ y ] = InpC( driveCroppedCrPic[ x ][ y ] )
[1320] }
[1321] If gfv_base_pic_flag is equal to 0, the keypoint coordinate array inputDriveKeyPoint and matrix inputDriveMatrix of the current picture are derived as follows.
[1322] for( i = 0; i < = convNumKeyPoint; i++ ) {
[1323] inputDriveKeyPoint[ i ]
[0000] = convKeyPoint[ i ]
[0000]
[1324] inputDriveKeyPoint[i]
[0001] = convKeyPoint[i]
[0001]
[1325] inputDriveKeyPoint[i]
[0002] = convKeyPoint[i]
[0002]
[1326] }
[1327] for( j = 0; j < convNumMatrix; j++ )
[1328] for( k = 0; k < convMatrixHeight; k++ )
[1329] for( m = 0; m < convMatrixWidth; m++ )
[1330] inputDriveMatrix[j][k][m] = convMatrix[j][k][m]
[1331] When gfv_base_pic_flag is 1, the keypoint coordinate array inputBaseKeyPoint and matrix inputBaseMatrix for the base picture are derived as follows.
[1332] for( i = 0; i <= convNumKeyPoint; i++ ) {
[1333] inputBaseKeyPoint[ i ]
[0000] = convKeyPoint[ i ]
[0000]
[1334] inputBaseKeyPoint[i]
[0001] = convKeyPoint[i]
[0001]
[1335] inputBaseKeyPoint[i]
[0002] = convKeyPoint[i]
[0002]
[1336] }
[1337] for( j = 0; j < convNumMatrix; j++ )
[1338] for( k = 0; k < convMatrixHeight; k++ )
[1339] for( l = 0; l < convMatrixWidth; l++ )
[1340] inputBaseMatrix[j][k][l] = convMatrix[j][k][l]
[1341] Here, the InpY() and InpC() functions are specified as follows.
[1342] InpY(x) = x / ( ( 1 << BitDepthY ) - 1 )
[1343] InpC(x) = x / ( ( 1 << BitDepthC ) - 1 )
[1344] GenerativeNN() is a process that generates sample values for the output image corresponding to the driving image. It is called only when gfc_base_pic_flag is 0. The input and output values of GenerativeNN() are real numbers.
[1345] The input to GenerativeNN() is as follows.
[1346] When gfv_base_pic_flag is 0, gfv_drive_pic_fusion_flag is 0, and ChromaFormatIdc is 0, inputBaseY, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix, CroppedWidth, CroppedHeight, CroppedDepth.
[1347] gfv_base_pic_flag가 0이고, gfv_drive_pic_fusion_flag가 0이고, ChromaFormatIdc가 0이 아닌 경우: inputBaseY, inputBaseCb, inputBaseCr, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix, CroppedWidth, CroppedHeight, CroppedDepth.
[1348] gfv_base_pic_flag가 0이고, gfv_drive_pic_fusion_flag가 1이고, gfv_chroma_key_info_present_flag가 0이고, ChromaFormatIdc가 0인 경우: inputBaseY, inputDriveY, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix, CroppedWidth, CroppedHeight, CroppedDepth.
[1349] gfv_base_pic_flag가 0이고, gfv_drive_pic_fusion_flag가 1이고, gfv_chroma_key_info_present_flag가 0이고, ChromaFormatIdc가 0이 아닌 경우: inputBaseY, inputBaseCb, inputBaseCr, inputDriveY, inputDriveCb, inputDriveCr, inputBaseKeyPoint, inputBaseMatrix,, inputDriveKeyPoint, inputDriveMatrix, CroppedWidth, CroppedHeight 및 CroppedDepth.
[1350] If gfv_base_pic_flag is 0, gfv_drive_pic_fusion_flag is 1, gfv_chroma_key_info_present_flag is 1, and ChromaFormatIdc is 0: inputBaseY, inputDriveY, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix, CroppedWidth, CroppedHeight, CroppedDepth, gfv_chroma_key_thr_value
[0000] , gfv_chroma_key_thr_value
[0001] , GfvChromaKeyValue
[0000] .
[1351] If gfv_base_pic_flag is 0, gfv_drive_pic_fusion_flag is 1, gfv_chroma_key_info_present_flag is 0, and ChromaFormatIdc is non-0: inputBaseY, inputBaseCb, inputBaseCr, inputDriveY, inputDriveCb, inputDriveCr, inputBaseKeyPoint, inputBaseMatrix, inputDriveKeyPoint, inputDriveMatrix, CroppedWidth, CroppedHeight, CroppedDepth, gfv_chroma_key_thr_value
[0000] , gfv_chroma_key_thr_value
[0001] , and if specified, GfvChromaKeyValue
[0000] , GfvChromaKeyValue
[0001] , GfvChromaKeyValue
[0002] .
[1352] The output of GenerativeNN() is as follows.
[1353] Luma sample array genY
[1354] If ChromaFormatIdc is not 0, two chroma sample arrays genCb and genCr.
[1355] The StoreOutputTensors() process for deriving output is specified as follows.
[1356] When gfv_base_pic_flag is 0, the output sample arrays outYPic[x][y], outCbPic[x][y], and outCrPic[x][y] are derived as follows.
[1357] for(x = 0; x < CroppedWidth; x++ )
[1358] for( y = 0; y < CroppedHeight; y++ )
[1359] outputYPic[ x ][ y ] = OutY( genY[ x ][ y ] )
[1360] if( ChromaFormatIdc != 0 )
[1361] for(x = 0; x < CroppedWidth / SubWidthC; x++ )
[1362] for( y = 0; y < CroppedHeight / SubHeightC; y++ ) {
[1363] outputCbPic[x][y] = OutC(genCb[x][y])
[1364] outputCrPic[x][y] = OutC(genCr[x][y])
[1365] }
[1366] If gfv_base_pic_flag is equal to 1, the output sample arrays outYPic[ x ][ y ], outCbPic[ x ][ y ], and outCrPic[ x ][ y ] are derived as follows (each output picture derived from the StoreOutputTensors() process is called a GFV-generated picture).
[1367] for(x = 0; x< CroppedWidth; x++ )
[1368] for( y = 0; y< CroppedHeight; y++ )
[1369] outputYPic[ x ][ y ] = baseCroppedYPic[ x ][ y ]
[1370] if( ChromaFormatIdc != 0 )
[1371] for(x = 0; x< CroppedWidth / SubWidthC; x++ )
[1372] for( y = 0; y< CroppedHeight / SubHeightC; y++ ) {
[1373] outputCbPic[x][y] = baseCroppedCbPic[x][y]
[1374] outputCrPic[x][y] = baseCroppedCbPic[x][y]
[1375] }
[1376] Here, the OutY() and OutC() functions are specified as follows.
[1377] OutY( x ) = Clip3( 0, ( 1 << BitDepthY ) - 1 , x * ( ( 1 << BitDepthY ) - 1 )
[1378] OutC( x ) = Clip3( 0, ( 1 << BitDepthC ) - 1 , x * ( ( 1 << BitDepthC ) - 1 )
[1379] The output order of GFV-generated pictures corresponding to picture-unit GFV SEI messages having the same gfv_id value and different gfv_cnt values must be sorted in ascending order of gfv_cnt values. For two pictures picA and picB, where picA precedes picB in the output order, the GFV-generated picture associated with picA that corresponds to a GFV SEI message with a specific gfv_id value precedes the GFV-generated picture associated with picB that corresponds to a GFV SEI message with a specific gfv_id value in the output order.
[1380] Regarding the technical problems to be solved in the embodiments,
[1381] JVET-AI2032). The current design includes an instance count (i.e., gfv_cnt) signal that specifies the order of GFV SEI messages for a specific ID (i.e., gfv_id) within a Picture Unit (PU). The first GFV SEI message for a specific ID within the PU must have a value of 0.
[1382] The gfv_cnt value for a GFV SEI message of a specific ID within a PU must be unique. This is to ensure that when there are two GFV SEI messages within the same PU with the same gfv_id value and the same gfv_cnt value, the decoder can easily understand that the second SEI message is merely a duplicate / repetition of the first message and can be ignored. However, this does not apply in the current design, as the meaning of gfv_cnt is specified as follows.
[1383] gfv_cnt specifies the number of GFV SEI message instances for this gfv_id value within the picture unit.
[1384] The gfv_cnt of the first GFV SEI message in decoding order with a specific gfv_id value within a picture unit must be equal to 0. If the gfv_cnt assigned to currGfvCnt is greater than 0, then a GFV SEI message with the same gfv_id value and a gfv_cnt equal to currGfvCnt - 1 must exist in the same picture unit and take precedence over the current GFV SEI message in decoding order.
[1385] The gfv_cnt value is in the range from 0 to 65535.
[1386] The mechanism for determining whether a GFV SEI message is a repetition of another GFV SEI message in the same PU is as follows.
[1387] If any of the following conditions are true, the GFV SEI message must have the same SEI payload content.
[1388] - GFV SEI messages exist in the same picture unit, gfv_cnt is 0, gfv_nn_base_flag exists, and gfv_id and gfv_nn_base_flag values are the same.
[1389] - GFV SEI messages exist in the same picture unit, gfv_cnt is greater than 0, and gfv_id is the same.
[1390] The result of determining above whether a GFV SEI message within the PU is a repetition of another GFV SEI message leads to the conclusion that while there may exist two GFV SEI messages with identical gfv_id and gfv_cnt values, the second message is not a repetition of the first message. For example, if two GFV SEI messages have the same gfv_id value and both have a gfv_cnt of 0, their gfv_nn_base_flag values are different.
[1391] It is desirable to clarify the GFV SEI message design in terms of count number (i.e., syntax element gfv_cnt) signaling. The following clarification is proposed.
[1392] - Option 1: Clearly defines that the gfv_cnt value is unique within the same PU as the same gfv_id value.
[1393] - Option 2: If gfv_cnt is 0 and two GFV SEI messages with different payloads are allowed, you must specify that the gfv_nn_base_flag value of at least the first SEI must be 1.
[1394] The method and apparatus according to the embodiments provide a solution to the aforementioned problem. Each item of the embodiment may be applied individually or in combination.
[1395] 1. Specify that the instance count value (i.e., the syntax element gfv_cnt) is unique for a specific GFV SEI message ID within the same picture unit.
[1396] 2. If two GFV SEI messages within the same picture unit have the same ID (i.e., the same gfv_id value) and both messages have the same instance count value (i.e., gfv_cnt is 0), the second SEI is a repetition of the first SEI (i.e., both messages have the same payload).
[1397] 3. You can additionally specify that the decoder may ignore or remove repetitive GFV SEI messages within the same picture unit.
[1398] 4. The signaling design of GFV SEI messages can be further modified so that GFV SEI messages where gfv_cnt is greater than 0 can include a filter for generative faces, and the following applies.
[1399] - If a GFV SEI message with gfv_cnt greater than 0 contains a filter, that filter must not be the default filter.
[1400] - The first GFV SEI message with the same ID in the same picture unit (i.e., when gfv_cnt is 0) must contain a default filter.
[1401] 5. Or, two GFV SEI messages with the same gfv_id value in the same picture unit must both have a gfv_cnt value of 0 but different gfv_nn_base_flag values (i.e., one SEI contains the base filter and the other SEI contains the updated filter). The message containing the base filter must take precedence over the message containing the updated filter.
[1402] The embodiments can be described based on the descriptions in the VSEI and VVC standard documents.
[1403] Below, Example 1:
[1404] If two or more GFV SEI messages in the same picture unit have the same gfv_id and gfv_cnt values, those messages must have the same SEI payload content.
[1405] Note: If two or more GFV SEI messages in the same picture unit have the same gfv_id and gfv_cnt values, the second GFV SEI message and the remaining messages may be ignored as they are duplicates of the first message.
[1406] Below, Example 2:
[1407] FIG. 36 shows the syntax of a generative face video SEI message according to the embodiments.
[1408] gfv_cnt specifies the number of GFV SEI message instances for this gfv_id value within the picture unit.
[1409] The gfv_cnt of the first GFV SEI message in decoding order with a specific gfv_id value within a picture unit must be equal to 0. If the gfv_cnt assigned to currGfvCnt is greater than 0, a GFV SEI message with the same gfv_id value and a gfv_cnt equal to currGfvCnt - 1 exists in the same picture unit and takes precedence over the current GFV SEI message in decoding order.
[1410] The gfv_cnt value must be in the range from 0 to 65535.
[1411] If gfv_nn_present_flag is 1, it indicates that a neural network that can be used as TranslatorNN() is included or displayed in the SEI message. If gfv_nn_present_flag is 0, it indicates that a neural network that can be used as TranslatorNN() is not included or displayed in the SEI message. If gfv_nn_present_flag is absent, it is inferred to be 0.
[1412] If a GFV SEI message with gfv_cnt of 0 exists in the first picture unit of CLVS in the decoding order, gfv_nn_present_flag must exist and be equal to 1.
[1413] If gfv_nn_present_flag is 0 and TranslatorNN is referenced in the semantics of a GFV SEI message, the following constraints apply.
[1414] - When gfv_cnt is 0, there must be at least one GFV SEI message in the previous picture unit in the output order of the current CLVS, and this message must have the same gfv_id value as the current GFV SEI message and gfv_nn_present_flag must be equal to 1.
[1415] - Otherwise (if gfv_cnt is greater than 0), at least one GFV SEI message must exist in the current picture unit or the previous picture unit in output order. Or, if it is the previous picture unit in output order within the current CLVS, has the same gfv_id value as the current GFV SEI message, and gfv_nn_present_flag is 1.
[1416] If gfv_nn_present_flag is 0 and TranslatorNN is referenced in the semantics of this SEI message, the following is applied to derive an applicable TranslatorNN.
[1417] - If gfv_cnt is greater than 0 and there is one or more previous GFV SEI messages with gfv_nn_present_flag of 1 that have the same gfv_id value as the current GFV SEI message in decoding order at the current picture level, the applicable TranslatorNN is defined as the last previous GFV SEI message with gfv_nn_present_flag of 1 that has the same gfv_id value as the current GFV SEI message in decoding order at the current picture level.
[1418] - Otherwise, the applicable TranslatorNN is defined as the GFV SEI message existing in the picture unit puB prior to the last in output order in the current CLVS, which has the same gfv_id value as the current GFV SEI message and gfv_nn_present_flag as 1. If there are multiple such GFV SEI messages in the picture unit puB that have the same gfv_id value as the current GFV SEI message and gfv_nn_present_flag as 1, the TranslatorNN is defined as the last of these GFV SEI messages in the decoding order.
[1419] If gfv_base_pic_flag is 1, it indicates that the currently decoded output image corresponds to the base image. If gfv_base_pic_flag is 0, it indicates that the currently decoded output image does not correspond to the base image or that this SEI message does not specify a syntax element for the base image. If gfv_base_pic_flag is absent, it is inferred to be 0.
[1420] If the GFV SEI message is the first GFV SEI message with a specific gfv_id value in the current CLVS in the decoding order, the gfv_base_pic_flag value must be equal to 1.
[1421] If the gfv_base_pic_flag of a GFV SEI message with a specific gfv_id value is 1, the base picture for that specific gfv_id value (the currently cropped decoded picture) comes after the currently decoded picture in output order for the currently decoded picture and all subsequent decoded pictures of the curre...
Claims
1. A step of acquiring GFV (Generative face video) SEI messages; and A step of generating an output picture based on the above GFV (Generative face video) SEI messages; comprising Decryption method.
2. In Paragraph 1, The first GFV SEI message includes a first identifier for the first GFV SEI message within the picture unit and a first count representing a GFV SEI message count value for the first identifier, Decryption method.
3. In Paragraph 1, The second GFV SEI message includes a second identifier for the second GFV SEI message within the picture unit and a second count representing a GFV SEI message count value for the second identifier, Decryption method.
4. In paragraph 3, the above method is: The method further comprises the step of determining that the second GFV SIE message is a repetition of the first GFV SEI message based on the fact that the second identifier is equal to the first identifier and the second count is equal to the first count. Decryption method.
5. In Paragraph 4, Based on the fact that the second identifier is equal to the first identifier and the second count is equal to the first count, the first GFV SEI message and the second GFV SEI message have the same SEI payload content, Decryption method.
6. Memory; and At least one processor connected to the memory; comprising, wherein the at least one processor: Acquire GFV (Generative face video) SEI messages; and Configured to generate an output picture based on the above GFV (Generative face video) SEI messages, Decoding device.
7. In Paragraph 6, The first GFV SEI message includes a first identifier for the first GFV SEI message within the picture unit and a first count representing a GFV SEI message count value for the first identifier, Decoding device.
8. In Paragraph 6, The second GFV SEI message includes a second identifier for the second GFV SEI message within the picture unit and a second count representing a GFV SEI message count value for the second identifier, Decoding device.
9. In paragraph 8, the above at least one processor is: Further configured to determine that the second GFV SEI message is a repetition of the first GFV SEI message based on the fact that the second identifier is equal to the first identifier and the second count is equal to the first count. Decoding device.
10. In Paragraph 9, Based on the fact that the second identifier is equal to the first identifier and the second count is equal to the first count, the first GFV SEI message and the second GFV SEI message have the same SEI payload content, Decoding device.
11. A step of encoding a picture unit for a picture; and The method comprising the step of generating GFV (Generative face video) SEI messages for the picture unit; Encoding method.
12. In Paragraph 11, The first GFV SEI message includes a first identifier for the first GFV SEI message within the picture unit and a first count representing a GFV SEI message count value for the first identifier, Encoding method.
13. Memory; and At least one processor connected to the memory; comprising, wherein the at least one processor: Encoding a picture unit for a picture; and Configured to generate GFV (Generative face video) SEI messages for the above picture unit, Encoding device.
14. A computer-readable storage medium for storing a bitstream generated by the method according to paragraph 11.
15. Step of acquiring the bitstream for the picture, The bitstream is generated based on the steps of: encoding a picture unit for the picture; and generating Generative face video (GFV) SEI messages for the picture unit; and A method comprising the step of transmitting data including the bitstream above.