Image encoding method, image encoding device, image decoding method, image decoding device, method for transmitting bitstream, and recording medium storing bitstream
The proposed image encoding and decoding method addresses the challenge of high data volume in high-resolution images by using SEI messages and advanced encoding techniques, achieving efficient compression and reduced costs.
Patent Information
- Application Number
- PCT/KR2025/009501
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-09
- Filing Date
- 2025-07-03
- Publication Date
- 2026-01-15
AI Technical Summary
The increasing demand for high-resolution, high-quality images leads to a significant increase in the amount of information transmitted, resulting in higher transmission and storage costs, necessitating a more efficient image compression technology.
Implementing an image encoding and decoding method that includes Supplemental Enhancement Information (SEI) messages for improved encoding and decoding efficiency, utilizing techniques such as intra prediction, inter prediction, transformation, quantization, and entropy encoding to optimize image compression.
Enhances encoding and decoding efficiency, reducing the amount of data required for high-resolution images, thereby lowering transmission and storage costs while maintaining image quality.
Smart Images

Figure KR2025009501_15012026_PF_FP_ABST
Abstract
Description
Image encoding method, image encoding device, image decoding method, image decoding device, method for transmitting bitstream, and recording medium storing bitstream
[0001] The embodiments relate to a video encoding method, a video encoding device, a video decoding method, a video decoding device, a method for transmitting a bitstream, and a recording medium storing a bitstream.
[0002] Recently, demand for high-resolution, high-quality images, such as HD (High Definition) and UHD (Ultra High Definition) images, has been increasing across various fields. As image data becomes higher resolution and higher quality, the amount of information transmitted, or bits, increases relative to conventional image data. This increase in information or bits transmitted leads to increased transmission and storage costs.
[0003] Accordingly, a highly efficient image compression technology is required to effectively transmit, store, and play high-resolution, high-quality image information.
[0004] The embodiments provide an image encoding method, an image encoding device, an image decoding method, an image decoding device, a method for transmitting a bitstream, and a recording medium storing a bitstream.
[0005] Embodiments provide an image encoding method with improved encoding and decoding efficiency, an image encoding device, an image decoding method, an image decoding device, a method for transmitting a bitstream, and a recording medium storing a bitstream.
[0006] However, the scope of the embodiments is not limited to the technical tasks described above, and the scope of the embodiments may be expanded to other technical tasks that can be inferred by a person skilled in the art based on the entire described content.
[0007] A method according to embodiments may include the steps of: obtaining a Supplemental Enhancement Information (SEI) message for pictures; and decoding the pictures. A method according to embodiments may include the steps of: deriving a Supplemental Enhancement Information (SEI) message for pictures; and encoding the pictures.
[0008] Embodiments provide a video encoding / decoding method and device with improved encoding / decoding efficiency.
[0009] Embodiments provide a non-transitory computer-readable recording medium for storing a bitstream generated by an image encoding method.
[0010] Embodiments provide a non-transitory computer-readable recording medium that stores a bitstream received and decoded by an image decoding device and used to restore an image.
[0011] Embodiments provide a method for transmitting a bitstream generated by a video encoding method.
[0012] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the description below.
[0013] The drawings are included to further understand the embodiments, and the drawings illustrate the embodiments together with the description related to the embodiments. For a better understanding of the various embodiments described below, reference should be made to the following description of the embodiments in conjunction with the following drawings, in which like reference numerals correspond to corresponding parts throughout the drawings.
[0014] Figure 1 illustrates a video and / or image coding system according to embodiments.
[0015] Figure 2 illustrates an encoding device according to embodiments.
[0016] Figure 3 shows a decoding device according to embodiments.
[0017] Figure 4 illustrates a content streaming system structure according to embodiments.
[0018] Figure 5 shows an example of a picture divided into Coding Tree Units (CTUs) according to embodiments.
[0019] FIG. 6 illustrates an example of a picture partitioned into tiles and raster-scan slices according to embodiments.
[0020] FIG. 7 illustrates an example of a picture partitioned into tiles and raster-scan slices according to embodiments.
[0021] FIG. 8 illustrates an example of a picture partitioned into tiles, bricks, and rectangular slices according to embodiments.
[0022] FIG. 9 illustrates an example of a picture including subpictures according to embodiments.
[0023] Figure 10 illustrates an example picture including tiles and CTUs according to embodiments.
[0024] Figure 11 illustrates a multi-type tree splitting mode according to embodiments.
[0025] Figure 12 shows splitting flags within a quad tree of a multi-type tree coding structure according to embodiments.
[0026] Figure 13 shows an example of a quad tree of a multi-type tree coding block structure according to embodiments.
[0027] Figure 14 illustrates the TT (Ternary Tree) split prohibition for coding blocks according to embodiments.
[0028] Figure 15 shows transforms and inverse transforms according to embodiments.
[0029] Figure 16 shows a low-frequency non-separable transform (LFNST) according to embodiments.
[0030] Figure 17 illustrates CABAC (Context Adaptive Binary Arithmetic Coding) encoding according to embodiments.
[0031] Figure 18 shows an entropy encoding method according to embodiments.
[0032] Figure 19 illustrates an entropy decoding method according to embodiments.
[0033] Figure 20 illustrates a picture decoding method according to embodiments.
[0034] Figure 21 illustrates a picture encoding method according to embodiments.
[0035] Figure 22 shows a hierarchical structure for coded images according to embodiments.
[0036] Figures 23a, 23b, 23c, 23d, and 23e illustrate picture header structures (picture_header_structure) according to embodiments.
[0037] Figures 24a, 24b, and 24c illustrate the syntax of a neural network post-filter characteristics SEI message according to embodiments.
[0038] Figure 25 illustrates a process of deriving a luma channel from a luma component according to embodiments.
[0039] Figure 26 illustrates the syntax of a neural network post-filter activation SEI (Supplemental enhancement information) message according to embodiments.
[0040] Figure 27 illustrates the syntax of a neural network post-filler group characteristic SEI message according to embodiments.
[0041] Figure 28 illustrates the syntax of a neural network post-filter group activation SEI message according to embodiments.
[0042] Figure 29 shows source picture timing information according to embodiments.
[0043] Figure 30 shows an object mask information SEI message according to embodiments.
[0044] Figure 31 shows an SEI processing order SEI message according to embodiments.
[0045] Figure 32 illustrates a processing order nesting SEI message according to embodiments.
[0046] Figure 33 shows the syntax of an encoder optimization information SEI message according to embodiments.
[0047] Figure 34 shows the syntax of a text description information SEI message according to embodiments.
[0048] Figure 35 shows the syntax of photosensitive content information according to embodiments.
[0049] Figure 36 shows a content event information SEI according to embodiments.
[0050] Figure 37 shows an encoding method according to embodiments.
[0051] Figure 38 shows a decryption method according to embodiments.
[0052] Preferred embodiments of the embodiments are described in detail, examples of which are illustrated in the accompanying drawings. The following detailed description, with reference to the accompanying drawings, is intended to illustrate preferred embodiments of the embodiments, rather than merely show embodiments that can be implemented according to the embodiments. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it will be apparent to those skilled in the art that the embodiments may be practiced without these details.
[0053] While most of the terms used in the examples are commonly used in the field, some terms were arbitrarily selected by the applicant, and their meanings are described in detail in the following descriptions as needed. Therefore, the examples should be understood based on the intended meaning of the terms, not simply their names or meanings.
[0054] Embodiments include a method and apparatus for signaling a content event in an SEI message for a coded video bitstream. Embodiments describe a method for signaling a content event in an SEI for a coded video bitstream. The described method is based on Versatile Video Coding (VVC) and Versatile Supplemental Enhancement Information Messages for Coded Video Bitstreams (VSEI), but may also be applied to other video coding technologies.
[0055] Related technical fields: Versatile Video Coding (VVC), Versatile supplemental enhancement information messages for coded video bitstreams (VSEI), Additional SEI messages for VSEI (Draft 3), SEI processing order and processing order nesting SEI messages in VVC (draft 7), Technologies under consideration for future extensions of VSEI (draft 4), SEI messages for VSEI version 4 (Draft 2).
[0056] Figure 1 illustrates a video and / or image coding system according to embodiments.
[0057] As illustrated in FIG. 1, a video / image coding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image information or data to the receiving device via a digital storage medium or network in the form of a file or streaming.
[0058] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / video encoding device, and the decoding device may be referred to as a video / video decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be configured as a separate device or an external component.
[0059] A video source may obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source may include a video / image capture device and / or a video / image generation device. A video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device may include, for example, a computer, a tablet, a smartphone, etc., and may (electronically) generate video / images. For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process of generating related data.
[0060] An encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, to improve compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0061] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit can include an element for generating a media file using a predetermined file format and an element for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0062] The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device.
[0063] The renderer can render decoded video / images. The rendered video / images can be displayed through the display unit.
[0064] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be applied to methods disclosed in the VVC (versatile video coding) standard, the EVC (essential video coding) standard, the AV1 (AOMedia Video 1) standard, the AVS2 (2nd generation of audio video coding standard), or the next generation video / image coding standard (e.g., H.267 or H.268).
[0065] This document presents various embodiments of video / image coding, and unless otherwise stated, the embodiments may be performed in combination with each other.
[0066] In this document, video can mean a collection of images over time. A picture generally refers to a unit representing one image at a specific time point, and a slice / tile is a unit that constitutes part of a picture in coding. A slice / tile can include one or more CTUs (coding tree units). A picture can be composed of one or more slices / tiles. A picture can be composed of one or more tile groups. A tile group can include one or more tiles. A brick can represent a rectangular area of CTU rows within a tile within a picture.
[0067] A brick can represent a rectangular area of a CTU row within a tile in a picture. A tile can be divided into multiple bricks, each of which consists of one or more CTU rows within the tile. A tile that is not divided into multiple bricks is also called a brick. A brick scan is a specific sequential order of CTUs that divide a picture. In a brick, CTUs are arranged consecutively by the CTU raster scan, bricks within a tile are arranged consecutively by the tile's brick raster scan, and tiles within a picture are arranged consecutively by the picture's tile raster scan. A tile is a rectangular area of a CTU within a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of a CTU with a height equal to the picture and a width specified by a syntax element of the picture parameter set. A tile row is a rectangular area of a CTU with a height specified by a syntax element of the picture parameter set and a width equal to the picture's width. A tile scan is a specific sequential order of CTUs that divide a picture. In tiles, CTUs are arranged consecutively by the CTU raster scan, and in pictures, tiles are arranged consecutively by the tile raster scan. A slice contains an integer number of bricks of a picture, which can be exclusively contained in a single NAL unit. A slice can consist of multiple complete tiles or a sequence of complete bricks of a tile arranged consecutively.
[0068] In this document, the terms tile group and slice may be used interchangeably. For example, in this document, the terms tile group / tile group header may be referred to as slice / slice header.
[0069] A pixel or pel can refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component.
[0070] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.
[0071] Figure 2 illustrates an encoding device according to embodiments.
[0072] FIG. 2 shows a schematic block diagram of an encoding device to which the embodiment(s) of the present document can be applied and in which encoding of a video / image signal is performed.
[0073] As shown in FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a prediction unit (predictor) 220, a residual processor (residual processor) 230, an entropy encoder (entropy encoder) 240, an adder (adder) 250, a filter (filter) 260, and a memory (memory) 270. The prediction unit (220) may include an inter prediction unit (221) and an intra prediction unit (222). The residual processor (230) may include a transformer (transformer) 232, a quantizer (quantizer) 233, a dequantizer (dequantizer) 234, and an inverse transformer (inverse transformer) 235. The residual processing unit (230) may further include a subtractor (231). The addition unit (250) may be called a reconstructor or a recontructed block generator. The image segmentation unit (210), the prediction unit (220), the residual processing unit (230), the entropy encoding unit (240), the addition unit (250), and the filtering unit (260) described above may be configured by one or more hardware components (e.g., an encoder chipset or a processor) depending on the embodiment. In addition, the memory (270) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.
[0074] The image segmentation unit (210) can segment an input image (or picture, frame) input to the encoding device (200) into one or more processing units. For example, the processing units may be referred to as coding units (CUs). In this case, the coding units may be recursively segmented from a coding tree unit (CTU) or a largest coding unit (LCU) according to a Quad-tree binary-tree ternary-tree (QTBTTT) structure. For example, one coding unit may be segmented into multiple coding units of deeper depth based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied first. The coding procedure according to the present document may be performed based on the final coding unit that is no longer segmented. In this case, based on coding efficiency according to image characteristics, etc., the maximum coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units of lower depths, and the coding unit of the optimal size can be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration described below. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may each be divided or partitioned from the final coding unit described above. The prediction unit may be a unit of sample prediction, and the transformation unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from a transform coefficient.
[0075] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the case. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or pel in a picture (or image).
[0076] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (predicted block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input video signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, as illustrated, a unit that subtracts a prediction signal (prediction block, prediction sample array) from an input video signal (original block, original sample array) within the encoder (200) may be called a subtraction unit (231). The prediction unit can perform prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied on a current block or CU basis. The prediction unit can generate various information regarding prediction, such as prediction mode information, as described later in the description of each prediction mode, and transmit the information to the entropy encoding unit (240). The information regarding prediction can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.
[0077] The intra prediction unit (222) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from it depending on the prediction mode. In intra prediction, the prediction modes may include multiple non-directional modes and multiple directional modes. The non-directional modes may include, for example, a DC mode and a planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the degree of detail in the prediction direction. However, this is merely an example, and a greater or lesser number of directional prediction modes may be used depending on the settings. The intra prediction unit (222) may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0078] The inter prediction unit (221) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. Temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and reference pictures including temporal neighboring blocks may be called collocated pictures (colPic). For example, the inter prediction unit (221) may construct a motion information candidate list based on neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of neighboring blocks as the motion information of the current block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0079] The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, sample values within a picture can be signaled based on information about the palette table and palette index.
[0080] The prediction signal generated through the prediction unit (including the inter-prediction unit (221) and / or the intra-prediction unit (222)) can be used to generate a restored signal or a residual signal. The transform unit (232) can apply a transform technique to the residual signal to generate transform coefficients. For example, the transform technique can include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is expressed as a graph. CNT refers to a transform obtained based on a prediction signal generated by using all previously reconstructed pixels. Additionally, the transformation process can be applied to blocks of pixels of equal size, either square or variable size, non-square.
[0081] The quantization unit (233) quantizes the transform coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients may be called residual information. The quantization unit (233) can rearrange the quantized transform coefficients in the form of a block into the form of a one-dimensional vector based on the coefficient scan order, and can also generate information about the quantized transform coefficients based on the quantized transform coefficients in the form of a one-dimensional vector. The entropy encoding unit (240) can perform various encoding methods, such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit (240) may encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in the form of a network abstraction layer (NAL) unit. The video / image information may further include information on various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. In this document, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the above-described encoding procedure and included in the bitstream.The bitstream may be transmitted over a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit (240) may be configured as an internal / external element of the encoding device (200) by a transmitting unit (not shown) and / or a storing unit (not shown), or the transmitting unit may be included in the entropy encoding unit (240).
[0082] The quantized transform coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (234) and the inverse transform unit (235), a residual signal (residual block or residual samples) can be reconstructed. The addition unit (155) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit (221) or the intra prediction unit (222). When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as a reconstructed block. The addition unit (250) may be called a reconstructor or a reconstructed block generation unit. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, and can also be used for inter prediction of the next picture after filtering as described below.
[0083] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.
[0084] The filtering unit (260) can improve subjective / objective picture quality by applying filtering to the restoration signal. For example, the filtering unit (260) can apply various filtering methods to the restoration picture to generate a modified restoration picture and store the modified restoration picture in the memory (270), specifically, in the DPB of the memory (270). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit the information to the entropy encoding unit (240), as described later in the description of each filtering method. The information regarding filtering may be encoded by the entropy encoding unit (240) and output in the form of a bitstream.
[0085] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter prediction unit (221). Through this, when inter prediction is applied, the encoding device can avoid prediction mismatch between the encoding device (100) and the decoding device, and can also improve encoding efficiency.
[0086] The memory (270) DPB can store the modified restored picture to be used as a reference picture in the inter prediction unit (221). The memory (270) can store motion information of a block from which motion information is derived (or encoded) within the current picture and / or motion information of blocks within a picture that has already been restored. The stored motion information can be transferred to the inter prediction unit (221) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (270) can store restored samples of restored blocks within the current picture and transfer them to the intra prediction unit (222).
[0087] Figure 3 shows a decoding device according to embodiments.
[0088] FIG. 3 shows a schematic block diagram of a decoding device to which the embodiment(s) of the present document can be applied and in which decoding of a video / image signal is performed.
[0089] As shown in FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-prediction unit (331) and an intra-prediction unit (332). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321). The entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) described above may be configured by a single hardware component (e.g., decoder chipset or processor) depending on the embodiment. In addition, the memory (360) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.
[0090] When a bitstream including video / image information is input, the decoding device (300) can restore the image corresponding to the process in which the video / image information is processed in the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division related information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied in the encoding device. Therefore, the processing unit of decoding may be, for example, a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. Then, the restored image signal decoded and output through the decoding device (300) can be reproduced through a reproduction device.
[0091] The decoding device (300) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through the entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The decoding device can decode the picture further based on the information on the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded and obtained from the bitstream through a decoding procedure. For example, the entropy decoding unit (310) can decode information in a bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements required for image restoration and the quantized values of transform coefficients for residuals. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element in the bitstream, determines a context model using information of the syntax element to be decoded and decoding information of the surrounding and decoding target blocks or information of symbols / bins decoded in the previous step, and predicts the occurrence probability of the bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model.Among the information decoded by the entropy decoding unit (310), information regarding prediction is provided to the prediction unit (inter-prediction unit (332) and intra-prediction unit (331)), and residual values on which entropy decoding is performed by the entropy decoding unit (310), for example, quantized transform coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive a residual signal (residual block, residual samples, residual sample array). In addition, information regarding filtering among the information decoded by the entropy decoding unit (310) can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of a decoding device (300), or the receiving unit may be a component of an entropy decoding unit (310). Meanwhile, a decoding device according to the present document may be called a video / video / picture decoding device, and the decoding device may be divided into an information decoder (video / video / picture information decoder) and a sample decoder (video / video / picture sample decoder). The information decoder may include an entropy decoding unit (310), and the sample decoder may include at least one of an inverse quantization unit (321), an inverse transformation unit (322), an addition unit (340), a filtering unit (350), a memory (360), an inter prediction unit (332), and an intra prediction unit (331).
[0092] The inverse quantization unit (321) can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit (321) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.
[0093] In the inverse transform unit (322), the transform coefficients are inversely transformed to obtain a residual signal (residual block, residual sample array).
[0094] The prediction unit can perform a prediction on the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit (310), the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block, and can determine a specific intra- / inter-prediction mode.
[0095] The prediction unit (320) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. IBC can utilize at least one of the inter prediction techniques described in this document. Palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, information about the palette table and palette index can be signaled and included in the video / image information.
[0096] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from it, depending on the prediction mode. In intra prediction, the prediction modes may include multiple non-directional modes and multiple directional modes. The intra prediction unit (331) can also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0097] The inter prediction unit (332) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) can construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and the information about the prediction can include information indicating the mode of inter prediction for the current block.
[0098] The addition unit (340) can generate a restoration signal (restored picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (predicted block, prediction sample array) output from the prediction unit (including the inter-prediction unit (332) and / or intra-prediction unit (331)). When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the restoration block.
[0099] The addition unit (340) may be referred to as a restoration unit or restoration block generation unit. The generated restoration signal may be used for intra prediction of the next processing target block within the current picture, may be output after filtering as described below, or may be used for inter prediction of the next picture.
[0100] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0101] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can apply various filtering methods to the restored picture to generate a modified restored picture, and transmit the modified restored picture to the memory (360), specifically, to the DPB of the memory (360). The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0102] The (modified) reconstructed picture stored in the DPB of the memory (360) can be used as a reference picture in the inter prediction unit (332). The memory (360) can store motion information of a block from which motion information is derived (or decoded) within the current picture and / or motion information of blocks within a picture that has already been reconstructed. The stored motion information can be transferred to the inter prediction unit (260) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (360) can store reconstructed samples of reconstructed blocks within the current picture and transfer them to the intra prediction unit (331).
[0103] In this specification, the embodiments described in the filtering unit (260), the inter prediction unit (221), and the intra prediction unit (222) of the encoding device (100) can be applied to the filtering unit (350), the inter prediction unit (332), and the intra prediction unit (331) of the decoding device (300) in the same or corresponding manner, respectively.
[0104] Implementation and application examples:
[0105] The embodiments described in this document may be implemented and performed on a processor, microprocessor, controller, or chip. For example, the functional units depicted in each drawing may be implemented and performed on a computer, processor, microprocessor, controller, or chip. In this case, information for implementation (e.g., information on instructions) or algorithms may be stored on a digital storage medium.
[0106] In addition, the decoding device and encoding device to which the embodiment(s) of the present document are applied may be included in a multimedia broadcasting transmitting and receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as a video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VoD) service providing device, an OTT (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (argumente reality) device, a video phone video device, a transportation terminal (ex. a vehicle (including an autonomous vehicle) terminal, an airplane terminal, a ship terminal, etc.), and a medical video device, and may be used to process a video signal or a data signal. For example, the OTT (Over the top video) device may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc.
[0107] In addition, the processing method to which the embodiment(s) of the present document are applied can be produced in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of the present document can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium can include, for example, a Blu-ray Disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission via the Internet). In addition, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0108] Additionally, the embodiments of the present document may be implemented as a computer program product by program code, and the program code may be executed on a computer by the embodiments of the present document. The program code may be stored on a computer-readable carrier.
[0109] Figure 4 illustrates a content streaming system structure according to embodiments.
[0110] A content streaming system to which the embodiments of this document are applied may broadly include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0111] The encoding server compresses content input from multimedia input devices, such as smartphones, cameras, and camcorders, into digital data, generates a bitstream, and transmits it to a streaming server. Alternatively, if multimedia input devices, such as smartphones, cameras, and camcorders, generate the bitstream directly, the encoding server may be omitted.
[0112] A bitstream can be generated by an encoding method or a bitstream generation method to which the embodiment(s) of this document are applied, and a streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0113] A streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary, informing the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. The content streaming system may include a separate control server, in which case the control server controls the command / response flow between each device within the content streaming system.
[0114] A streaming server can receive content from a media repository and / or encoding server. For example, if content is received from an encoding server, it can be received in real time. In this case, the streaming server can store the bitstream for a certain period of time to provide a smooth streaming service.
[0115] Examples of user devices include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs), digital TVs, desktop computers, digital signage, etc.
[0116] Each server within a content streaming system can be operated as a distributed server, in which case the data received from each server can be processed in a distributed manner.
[0117] Partitioning structure:
[0118] The video / image coding method according to this document can be performed based on the following partitioning structure. Specifically, the procedures such as prediction, residual processing ((inverse) transformation, (inverse) quantization, etc.), syntax element coding, and filtering described below can be performed based on CTU, CU (and / or TU, PU) derived based on the partitioning structure. The block partitioning procedure is performed in the image partitioning unit (210) of the encoding device described above, and partitioning-related information can be (encoded) processed in the entropy encoding unit (240) and transmitted to the decoding device in the form of a bitstream. The entropy decoding unit (310) of the decoding device can derive the block partitioning structure of the current picture based on the partitioning-related information obtained from the bitstream, and perform a series of procedures for image decoding (e.g., prediction, residual processing, block / picture restoration, in-loop filtering, etc.) based on the same. The CU size and the TU size may be the same, or multiple TUs may exist within the CU area. Meanwhile, the CU size may generally represent the luma component (sample) CB size. The TU size may generally represent the luma component (sample) TB size. The chroma component (sample) CB or TB size may be derived based on the luma component (sample) CB or TB size according to the component ratio according to the color format of the picture / video (chroma format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.). The TU size may be derived based on maxTbSize. For example, when the CU size is larger than maxTbSize, multiple TUs (TBs) of maxTbSize may be derived from C, and transformation / inverse transformation may be performed in units of TU (TB). Additionally, for example, when intra prediction is applied, the intra prediction mode / type can be derived in units of CU (or CB), and the procedures for deriving surrounding reference samples and generating prediction samples can be performed in units of TU (or TB).In this case, one or more TUs (or TBs) may exist within one CU (or CB) region, and in this case, multiple TUs (or TBs) may share the same intra prediction mode / type.
[0119] In addition, in the coding of video / images according to this document, the image processing unit may have a hierarchical structure. One picture may be divided into one or more tiles, bricks, slices, and / or tile groups. One slice may include one or more bricks. One brick may include one or more CTU rows within the tile. A slice may include an integer number of bricks of the picture. One tile group may include one or more tiles. One tile may include one or more CTUs. A CTU may be divided into one or more CUs. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. A tile group may include an integer number of tiles according to a tile raster scan within a picture. A slice header may carry information / parameters applicable to the corresponding slice (blocks within the slice). When the encoding / decoding device has a multi-core processor, the encoding / decoding procedure for tiles, slices, bricks, and / or tile groups can be processed in parallel. In this document, the terms slice and tile group can be used interchangeably. The tile group header may be called a slice header. Here, the slice may have one of the following slice types: intra (I) slice, predictive (P) slice, and bi-predictive (B) slice. For blocks within an I slice, inter prediction is not used for prediction, and only intra prediction can be used. Of course, in this case, the original sample values can be coded and signaled without prediction.For blocks within a P slice, either intra prediction or inter prediction can be used, and when inter prediction is used, only uni prediction can be used. On the other hand, for blocks within a B slice, either intra prediction or inter prediction can be used, and when inter prediction is used, up to bi prediction can be used.
[0120] The encoder may determine the tile / tile group, brick, slice, maximum and minimum coding unit sizes based on the characteristics of the video image (e.g., resolution) or considering coding efficiency or parallel processing, and information about this or information that can derive this may be included in the bitstream.
[0121] The decoder can obtain information indicating whether the tile / tile group of the current picture, bricks, slias, and CTUs within a tile are divided into multiple coding units. Efficiency can be improved by ensuring that this information is obtained (transmitted) only under certain conditions.
[0122] A slice header (slice header syntax) may include information / parameters that are commonly applicable to slices. An APS (APS syntax) or a PPS (PPS syntax) may include information / parameters that are commonly applicable to one or more pictures. An SPS (SPS syntax) may include information / parameters that are commonly applicable to one or more sequences. A VPS (VPS syntax) may include information / parameters that are commonly applicable to multiple layers. A DPS (DPS syntax) may include information / parameters that are commonly applicable to the entire video. A DPS may include information / parameters related to the concatenation of CVS (coded video sequence).
[0123] In this document, high-level syntax may include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, and slice header syntax.
[0124] Additionally, for example, information regarding the division and composition of tiles / tile groups / bricks / slices, etc., can be configured at the encoding stage through higher-level syntax and transmitted to the decoding device in the form of a bitstream.
[0125] Figure 5 shows an example of a picture divided into Coding Tree Units (CTUs) according to embodiments.
[0126] Partitioning of picture into CTUs:
[0127] Pictures can be divided into a sequence of coding tree units (CTUs). A CTU may correspond to a coding tree block (CTB). Alternatively, a CTU may include a coding tree block of luma samples and two corresponding coding tree blocks of chroma samples. In other words, for a picture containing three sample arrays, a CTU may include an NxN block of luma samples and two corresponding blocks of chroma samples. Figure 5 illustrates an example of a picture being divided into CTUs.
[0128] The maximum allowable size of a CTU for coding and prediction, etc., may differ from the maximum allowable size of a CTU for transformation. For example, the maximum allowable size of a luma block within a CTU may be 128x128 (even though the luma ring blocks have a maximum size of 64x64).
[0129] FIG. 6 illustrates an example of a picture partitioned into tiles and raster-scan slices according to embodiments.
[0130] Partitioning of pictures into subpictures, slices, and tiles:
[0131] A picture is divided into one or more rows of tiles and one or more columns of tiles. A tile is a sequence of CTUs that cover a rectangular region of the picture. The CTUs within a tile are scanned in raster scan order within that tile.
[0132] A slice consists of an integer number of complete tiles or an integer number of contiguous rows of complete CTUs within a tile of a picture.
[0133] Slices support two modes: raster scan slice mode and rectangular slice mode. In raster scan slice mode, a slice contains a complete sequence of tiles from a tile raster scan of a picture. In rectangular slice mode, a slice contains multiple complete tiles that make up a rectangular region of a picture, or multiple contiguous complete CTU rows of a tile that makes up a rectangular region of a picture. Tiles within a rectangular slice are scanned in tile raster scan order within the rectangular region corresponding to that slice.
[0134] A subpicture consists of one or more slices that entirely cover a rectangular area of the picture.
[0135] Figure 6 shows an example of partitioning a picture into raster scan slices. Here, the picture is divided into 12 tiles and 3 raster scan slices.
[0136] FIG. 7 illustrates an example of a picture partitioned into tiles and raster-scan slices according to embodiments.
[0137] Figure 7 shows an example of partitioning a picture into rectangular slices. Here, the picture is divided into 24 tiles (6 tile columns and 4 tile rows) and 9 rectangular slices.
[0138] FIG. 8 illustrates an example of a picture partitioned into tiles, bricks, and rectangular slices according to embodiments.
[0139] Figure 8 shows an example of a picture divided into tiles and rectangular slices. Here, the picture is divided into four tiles (two tile columns and two tile rows) and four rectangular slices.
[0140] FIG. 9 illustrates an example of a picture including subpictures according to embodiments.
[0141] Figure 9 shows an example of subpicture segmentation of a picture. Here, the picture is segmented into 28 subpictures of various dimensions.
[0142] Figure 10 illustrates an example picture including tiles and CTUs according to embodiments.
[0143] When a picture is coded using three separate color planes (separate_colour_plane_flag is 1), a slice contains only a CTU of one color component identified by its color_plane_id value, and each color component array of the picture consists of slices with the same color_plane_id value. Within a picture, coded slices with different color_plane_id values may be interleaved with each other under the constraint that for each color_plane_id value, the coded slice NAL units with that color_plane_id value must be in increasing order of CTU address in the tile scan order with respect to the first CTU of each coded slice NAL unit.
[0144] Note - If separate_colour_plane_flag is 0, each CTU of the picture is contained in exactly one slice. If separate_colour_plane_flag is 1, each CTU of the colour component is contained in exactly one slice (information about each CTU of the picture is contained in exactly three slices, and these three slices have different colour_plane_id values).
[0145] Tile changes the order of CTUs in a picture. If a picture is divided into two or more tiles, the CTU order is the raster scan order within each tile, as shown in Figure 10. In Figure 10, the picture is divided into two tiles, and each tile contains eight CTUs. The order of CTUs within a tile is the raster scan order.
[0146] Figure 11 illustrates a multi-type tree splitting mode according to embodiments.
[0147] Partitioning of the CTUs using a tree structure:
[0148] A CTU can be partitioned into CUs based on a quad-tree (QT) structure. The quad-tree structure can be called a quaternary tree structure to reflect various local characteristics. Meanwhile, in this document, a CTU can be partitioned based on a multi-type tree structure partitioning that includes not only a quad-tree but also a binary-tree (BT) and a ternary-tree (TT). Hereinafter, the QTBT structure may include partitioning structures based on quad-tree and binary trees, and the QTBTTT structure may include partitioning structures based on quad-tree, binary trees, and ternary trees. Alternatively, the QTBT structure may include partitioning structures based on quad-tree, binary trees, and ternary trees. In the coding tree structure, a CU may have a square or rectangular shape. A CTU can first be partitioned into a quad-tree structure. Then, the leaf nodes of the quad-tree structure can be further partitioned by a multi-type tree structure. For example, as shown in Fig. 11, a multi-type tree structure can roughly include four partition types.
[0149] The four split types can include vertical binary splitting (SPLIT_BT_VER), horizontal binary splitting (SPLIT_BT_HOR), vertical ternary splitting (SPLIT_TT_VER), and horizontal ternary splitting (SPLIT_TT_HOR). The leaf nodes of the multi-type tree structure can be called CUs. These CUs can be used for prediction and transformation procedures. In this document, CUs, PUs, and TUs can generally have the same block size. However, CUs and TUs can have different block sizes if the maximum supported transform length is less than the width or height of the color component of the CU.
[0150] Figure 12 shows splitting flags within a quad tree of a multi-type tree coding structure according to embodiments.
[0151] Figure 12 illustrates an example of a signaling mechanism for partitioning information of a quadtree structure with nested multi-type trees.
[0152] Here, the CTU is treated as the root of the quadtree and is first partitioned into a quadtree structure. Each quadtree leaf node can be further partitioned into a multitype tree structure. In the multitype tree structure, a first flag (e.g., mtt_split_cu_flag) is signaled to indicate whether the node is further partitioned. If the node is further partitioned, a second flag (e.g., mtt_split_cu_vertical_flag) can be signaled to indicate the splitting direction. A third flag (e.g., mtt_split_cu_binary_flag) can then be signaled to indicate whether the splitting type is binary partitioning or ternary partitioning. For example, based on mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the multi-type tree splitting mode (MttSplitMode) of CU can be derived as shown in Table 1 (MttSplitMode derviation based on multi-type tree syntax elements).
[0153] [Table 1]
[0154]
[0155] Figure 13 shows an example of a quad tree of a multi-type tree coding block structure according to embodiments.
[0156] Figure 13 illustrates an example of a CTU being divided into multiple CUs based on a quadtree and nested multi-type tree structure.
[0157] Here, bold block edges represent quadtree partitioning, and the remaining edges represent multi-type tree partitioning. A quadtree partition involving a multi-type tree can provide a content-adaptive coding tree structure. A CU may correspond to a coding block (CB). Alternatively, a CU may include a coding block of luma samples and two coding blocks of corresponding chroma samples. The size of a CU may be as large as a CTU, or may be cut into units of 4x4 in luma samples. For example, in the case of a 4:2:0 color format (or chroma format), the maximum chroma CB size may be 64x64, and the minimum chroma CB size may be 2x2.
[0158] For example, in this document, the maximum allowed luma TB size may be 64x64, and the maximum allowed chroma TB size may be 32x32. If the width or height of a CB split according to the tree structure is larger than the maximum transform width or height, the CB may be automatically (or implicitly) split until it satisfies the horizontal and vertical TB size constraints.
[0159] Meanwhile, for a quadtree coding tree scheme involving a multi-type tree, the following parameters can be defined and identified as SPS syntax elements.
[0160] CTU size: The size of the root node of a 4th-order tree
[0161] MinQTSize: Minimum allowed size of a 4th order tree leaf node.
[0162] MaxBtSize: Maximum allowed binary tree root node size.
[0163] MaxTtSize: Maximum allowed size of a tertiary tree root node.
[0164] MaxMttDepth: The maximum allowed layer depth of a multi-type tree split at a 4th-order tree leaf.
[0165] MinBtSize: Minimum allowed binary tree leaf node size.
[0166] MinTtSize: Minimum allowed tertiary tree leaf node size
[0167] As an example of a quadtree coding tree structure involving a multi-type tree, the CTU size can be set to 64x64 blocks of 128x128 luma samples and two corresponding chroma samples (in a 4:2:0 chroma format). In this case, MinOTSize can be set to 16x16, MaxBtSize can be set to 128x128, MaxTtSzie can be set to 64x64, MinBtSize and MinTtSize (for both width and height) can be set to 4x4, and MaxMttDepth can be set to 4. Quadtree partitioning can be applied to the CTU to generate quadtree leaf nodes. The quadtree leaf nodes can be called leaf QT nodes. The quadtree leaf nodes can have sizes from 16x16 (i.e. the MinOTSize) to 128x128 (i.e. the CTU size). If a leaf QT node is 128x128, it may not be further split into binary trees / ternary trees. This is because in this case, even if it is split, it exceeds MaxBtsize and MaxTtszie (i.e. 64x64). In other cases, a leaf QT node may be further split into a multitype tree. Therefore, a leaf QT node is a root node for a multitype tree, and a leaf QT node may have a multitype tree depth (mttDepth) value of 0. If the multitype tree depth reaches MaxMttdepth (e.g. 4), no further splits may be considered. If the width of a multitype tree node is equal to MinBtSize and less than or equal to 2xMinTtSize, no further horizontal splits may be considered. If the height of a multitype tree node is equal to MinBtSize and less than or equal to 2xMinTtSize, no further vertical splits may be considered.
[0168] Figure 14 illustrates the TT (Ternary Tree) split prohibition for coding blocks according to embodiments.
[0169] In hardware decoders, to allow for 64x64 luma block and 32x32 chroma pipeline design, TT splitting may be prohibited in certain cases. For example, if the width or height of a luma coding block is greater than 64, TT splitting may be prohibited, as illustrated in FIG. 14. Also, for example, if the width or height of a chroma coding block is greater than 32, TT splitting may be prohibited.
[0170] In this document, the coding tree scheme may support luma and chroma (component) blocks having separate block tree structures. When luma and chroma blocks within a CTU have the same block tree structure, it may be denoted as SINGLE_TREE. When luma and chroma blocks within a CTU have separate block tree structures, it may be denoted as DUAL_TREE. In this case, the block tree type for the luma component may be called DUAL_TREE_LUMA, and the block tree type for the chroma component may be called DUAL_TREE_CHROMA. For P and B slice / tile groups, luma and chroma CTBs within a CTU may be restricted to have the same coding tree structure. However, for I slice / tile groups, luma and chroma blocks may have separate block tree structures. When the individual block tree mode is applied, the luma CTB may be partitioned into CUs based on a specific coding tree structure, and the chroma CTB may be partitioned into chroma CUs based on a different coding tree structure. This may mean that a CU in an I slice / tile group may be composed of coding blocks of a luma component or coding blocks of two chroma components, and a CU in a P or B slice / tile group may be composed of blocks of three color components. In this document, a slice may be referred to as a tile / tile group, and a tile / tile group may be referred to as a slice.
[0171] Although the quadtree coding tree structure including a multi-type tree was described in the above-mentioned "Partitioning of the CTUs using a tree structure", the structure into which the CU is partitioned is not limited thereto. For example, the BT structure and the TT structure can be interpreted as concepts included in the Multiple Partitioning Tree (MPT) structure, and the CU can be interpreted as being partitioned through the QT structure and the MPT structure. In an example in which the CU is partitioned through the QT structure and the MPT structure, the partitioning structure can be determined by signaling a syntax element (e.g., MPT_split_type) including information on how many blocks the leaf node of the QT structure is partitioned into and a syntax element (e.g., MPT_split_mode) including information on whether the leaf node of the QT structure is partitioned vertically or horizontally.
[0172] In another example, a CU may be partitioned in a different way than the QT structure, the BT structure, or the TT structure. That is, unlike the QT structure where the CU of the lower depth is partitioned to 1 / 4 the size of the CU of the upper depth, the BT structure where the CU of the lower depth is partitioned to 1 / 2 the size of the CU of the upper depth, or the TT structure where the CU of the lower depth is partitioned to 1 / 4 or 1 / 2 the size of the CU of the upper depth, the CU of the lower depth may be partitioned to 1 / 5, 1 / 3, 3 / 8, 3 / 5, 2 / 3, or 5 / 8 the size of the CU of the upper depth, as the case may be, and the way in which the CU is partitioned is not limited thereto.
[0173] Transform / Inverse Transform:
[0174] As described above, the encoding device can derive a residual block (residual samples) based on a predicted block (prediction samples) through intra / inter / IBC prediction, etc., and can derive quantized transform coefficients by applying transformation and quantization to the derived residual samples. Information about the quantized transform coefficients (residual information) can be included in the residual coding syntax and output in the form of a bitstream after encoding. The decoding device can obtain information about the quantized transform coefficients (residual information) from the bitstream and decode it to derive the quantized transform coefficients. The decoding device can derive residual samples by performing inverse quantization / inverse transformation based on the quantized transform coefficients. As described above, at least one of quantization / inverse quantization and / or transformation / inverse transformation can be omitted. When transform / inverse transform is omitted, the transform coefficients may be called coefficients or residual coefficients, or may still be called transform coefficients for consistency of representation. Whether transform / inverse transform is omitted can be signaled based on transform_skip_flag.
[0175] Transformation / inverse transformation can be performed based on transformation kernel(s). For example, according to this document, the multiple transform selection (MTS) scheme can be applied. In this case, some of a set of multiple transformation kernels can be selected and applied to the current block. The transformation kernel can be referred to by various terms, such as transformation matrix, transformation type, etc. For example, the set of transformation kernels can represent a combination of vertical transformation kernels (vertical transformation kernels) and horizontal transformation kernels (horizontal transformation kernels).
[0176] For example, MTS index information (or tu_mts_idx syntax element) can be generated / encoded by an encoding device and signaled to a decoding device to indicate one of a set of transformation kernels. For example, a set of transformation kernels depending on the value of the MTS index information can be derived as in Table 2 (Specification of trTypeHor and trTypeVer depending on tu_mts_idx[ x ][ y ]), Table 3 (Specification of trTypeHor and trTypeVer depending on cu_sbt_horizontal_flag and cu_sbt_pos_flag), and / or Table 4 (Specification of trTypeHor and trTypeVer depending on predModeIntra).
[0177] [Table 2]
[0178]
[0179] The set of transformation kernels may also be determined based on, for example, cu_sbt_horizontal_flag and cu__sbt_pos_flag.
[0180] If cu_sbt_horizontal_flag is 1, it indicates that the current coding unit is split horizontally into two transform units. If cu_sbt_horizontal_flag[ x0 ][ y0 ] is 0, it indicates that the current coding unit is split vertically into two transform units. If cu_sbt_pos_flag is 1, it indicates that tu_cbf_luma, tu_cbf_cb, and tu_cbf_cr of the first transform unit of the current coding unit are not in the bitstream. If cu_sbt_pos_flag is 0, it indicates that tu_cbf_luma, tu_cbf_cb, and tu_cbf_cr of the second transform unit of the current coding unit are not in the bitstream.
[0181] [Table 3]
[0182]
[0183] The set of transformation kernels may be determined based on, for example, the intra prediction mode for the current block.
[0184] [Table 4]
[0185]
[0186] In the above tables, trTypeHor can represent a horizontal transform kernel, and trTypeVer can represent a vertical transform kernel. Here, a trTypeHor / trTypeVer value of 0 can represent DCT2, a trTypeHor / trTypeVer value of 1 can represent DST7, and a trTypeHor / trTypeVer value of 2 can represent DCT8. However, this is just an example, and other values may be mapped to other DCTs / DSTs by convention.
[0187] The following Table 5 (Transform basis functions of DCT-II / VIII and DSTVII for N-point input) shows examples of basis functions for the above-described DCT2, DCT8, and DST7.
[0188] [Table 5]
[0189]
[0190] Figure 15 shows transforms and inverse transforms according to embodiments.
[0191] In this document, the MTS-based transform is applied as a primary transform, and a secondary transform can be further applied. The secondary transform may be applied only to the coefficients in the upper left wxh region of the coefficient block to which the primary transform is applied, and may be called a reduced secondary transform (RST). For example, w and / or h may be 4 or 8. In the transform, the primary and secondary transforms may be sequentially applied to the residual block, and in the inverse transform, the inverse secondary transform and the inverse primary transform may be sequentially applied to the transform coefficients. The secondary transform (RST transform) may be called a low frequency coefficients transform (LFCT) or a low frequency non-seperable transform (LFNST). The inverse secondary transform may be called an inverse LFCT or an inverse LFNST.
[0192] Figure 16 shows a low-frequency non-separable transform (LFNST) according to embodiments.
[0193] The Low Frequency Non-Separable Transform (LFNST), also known as the Reduced Secondary Transform, is applied between the forward primary transform and quantization (encoder side), and between the inverse quantization and the inverse primary transform (decoder side), as shown in Fig. 16. In LFNST, either a 4x4 non-separable transform or an 8x8 non-separable transform is applied depending on the block size. For example, a 4x4 LFNST is applied to a small block (i.e., minimum (width, height) < 8), and an 8x8 LFNST is applied to a large block (i.e., minimum (width, height) > 4).
[0194] The application of the non-separable transform used in LFNST is explained as follows with an example input. To apply 4x4 LFNST, a 4x4 input block X is represented as a vector as follows.
[0195]
[0196]
[0197] Inseparable transformations are as follows: is calculated as . Here represents the transformation coefficient vector, and T is a 16x16 transformation matrix. 16x1 coefficient The vector is then reconstructed into 4x4 blocks using the scan order of the corresponding block (horizontal, vertical, or diagonal). Coefficients with smaller indices are placed at smaller scan indices in the 4x4 coefficient blocks.
[0198] Transformation / inverse transformation can be performed on a CU or TU basis. That is, transformation / inverse transformation can be applied to residual samples within a CU or residual samples within a TU. The CU size and the TU size can be the same, or multiple TUs can exist within a CU area. Meanwhile, the CU size can generally represent the luma component (sample) CB size. The TU size can generally represent the luma component (sample) TB size. The chroma component (sample) CB or TB size can be derived based on the luma component (sample) CB or TB size according to the component ratio according to the color format (chroma format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.). The TU size can be derived based on maxTbSize. For example, if the CU size is larger than maxTbSize, multiple TUs (TB) of maxTbSize can be derived from the CU, and transformation / inverse transformation can be performed in units of TU (TB). maxTbSize can be considered when determining whether to apply various intra prediction types such as ISP. Information about maxTbSize can be determined in advance, or can be generated and encoded by the encoding device and signaled to the decoding device.
[0199] Quantization / Dequantization:
[0200] As described above, the quantization unit of the encoding device can apply quantization to the transform coefficients to derive quantized transform coefficients, and the inverse quantization unit of the encoding device or the inverse quantization unit of the decoding device can apply inverse quantization to the quantized transform coefficients to derive transform coefficients.
[0201] In general, in video / image coding, the quantization rate can be changed, and the compression level can be adjusted using the changed quantization rate. From an implementation perspective, considering the complexity, instead of using the quantization rate directly, the quantization parameter (QP) can be used. For example, the quantization parameter with an integer value from 0 to 63 can be used, and each quantization parameter value can correspond to the actual quantization rate. The quantization parameter (QPY) for the luma component (luma sample) and the quantization parameter (QPC) for the chroma component (chroma sample) can be set differently.
[0202] The quantization process takes a transform coefficient (C) as input, divides it by a quantization rate (Qstep), and obtains a quantized transform coefficient (C`) based on this. In this case, considering the computational complexity, the quantization rate can be multiplied by a scale to make it into an integer, and a shift operation can be performed by a value corresponding to the scale value. The quantization scale can be derived based on the product of the quantization rate and the scale value. In other words, the quantization scale can be derived according to the QP. The quantized transform coefficient (C`) can also be derived based on this by applying the quantization scale to the transform coefficient (C).
[0203] The inverse quantization process is the reverse process of the quantization process. By multiplying the quantized transform coefficient (C`) by the quantization rate (Qstep), a restored transform coefficient (C``) can be obtained based on this. In this case, a level scale can be derived according to the quantization parameter, and by applying the level scale to the quantized transform coefficient (C`), a restored transform coefficient (C``) can be derived based on this. The restored transform coefficient (C``) may be somewhat different from the original transform coefficient (C) due to loss in the transformation and / or quantization process. Therefore, in the encoding device, in the same manner as in the decoding device, inverse quantization is performed.
[0204] Meanwhile, an adaptive frequency-based weighted quantization technique that adjusts the quantization strength according to frequency may be applied. The adaptive frequency-based weighted quantization technique is a method of applying different quantization strengths to different frequencies. The adaptive frequency-based weighted quantization can apply different quantization strengths to each frequency using a predefined quantization scaling matrix. That is, the quantization / dequantization process described above may be performed further based on the quantization scaling matrix. For example, different quantization scaling matrices may be used depending on the size of the current block and / or whether the prediction mode applied to the current block to generate the residual signal of the current block is inter-prediction or intra-prediction. The quantization scaling matrix may be referred to as a quantization matrix or a scaling matrix. The quantization scaling matrix may be predefined. In addition, for frequency-adaptive scaling, frequency-based quantization scale information for the quantization scaling matrix may be configured / encoded in an encoding device and signaled to a decoding device. The frequency-specific quantization scale information may be referred to as quantization scaling information. The frequency-specific quantization scale information may include scaling list data (scaling_list_data). Based on the scaling list data, a (modified) quantization scaling matrix may be derived. In addition, the frequency-specific quantization scale information may include a presence flag (present flag) information indicating whether the scaling list data is present. Alternatively, if the scaling list data is signaled at a higher level (e.g., SPS), information indicating whether the scaling list data is modified at a lower level (e.g., PPS or tile group header, etc.) may be further included.
[0205] Entropy coding:
[0206] As described above in the description of FIG. 2, some or all of the video / image information may be entropy encoded by the entropy encoding unit (240), and as described above in the description of FIG. 3, some or all of the video / image information may be entropy decoded by the entropy decoding unit (310). In this case, the video / image information may be encoded / decoded in units of syntax elements. In this document, encoding / decoding information may include encoding / decoding by the method described in this paragraph.
[0207] Figure 17 illustrates CABAC (Context Adaptive Binary Arithmetic Coding) encoding according to embodiments.
[0208] Figure 17 shows a block diagram of CABAC for encoding a single syntax element. The CABAC encoding process first converts the input signal into a binary value through binarization if the input signal is a syntax element rather than a binary value. If the input signal is already a binary value, binarization is bypassed. Here, each binary digit 0 or 1 that constitutes the binary value is called a bin. For example, if the binary string (bin string) after binarization is 110, 1, 1, and 0 are each called a bin. The bin(s) for a single syntax element can represent the value of the corresponding syntax element.
[0209] Binarized bins are input to a regular coding engine or a bypass coding engine. The regular coding engine assigns a context model that reflects the probability value for each bin and encodes the bin based on the assigned context model. The regular coding engine can update the probability model for each bin after coding. Bins coded in this way are called context-coded bins. The bypass coding engine omits the process of estimating the probability of each input bin and updating the probability model applied to the bin after coding. Instead of assigning a context, it applies a uniform probability distribution (e.g., 50:50) to the input bins, thereby improving coding speed. Bins coded in this way are called bypass bins. A context model can be assigned and updated for each bin subject to context coding (regular coding), and the context model can be indicated based on ctxidx or ctxInc. ctxidx can be derived based on ctxInc. Specifically, for example, the context index (ctxidx) indicating the context model for each of the regular coded bins can be derived as the sum of the context index increment (ctxInc) and the context index offset (ctxIdxOffset). Here, ctxInc can be derived differently for each bin. ctxIdxOffset can be represented by the lowest value of ctxIdx. The lowest value of ctxIdx can be called the initial value (initValue) of ctxIdx. ctxIdxOffset is a value generally used to distinguish from context models for other syntax elements, and a context model for a syntax element can be distinguished / derived based on ctxinc.
[0210] In the entropy encoding process, encoding can be performed via the regular coding engine or the bypass coding engine, and the coding path can be switched. Entropy decoding performs the same process as entropy encoding in reverse order.
[0211] Figure 18 shows an entropy encoding method according to embodiments.
[0212] The entropy coding described in Fig. 17 can be performed, for example, as in Fig. 18.
[0213] Referring to FIG. 18, an encoding device (entropy encoding unit) performs an entropy coding procedure on image / video information. The image / video information may include partitioning-related information, prediction-related information (e.g., inter / intra prediction distinction information, intra-prediction mode information, inter-prediction mode information, etc.), residual information, in-loop filtering-related information, etc., or may include various syntax elements related thereto. Entropy coding may be performed on a syntax element basis. S600 to S610 may be performed by the entropy encoding unit (240) of the encoding device of FIG. 2 described above.
[0214] The encoding device performs binarization on the target syntax element (S600). Here, binarization may be based on various binarization methods, such as the Truncated Rice binarization process and the Fixed-length binarization process, and the binarization method for the target syntax element may be predefined. The binarization procedure may be performed by the binarization unit (242) within the entropy encoding unit (240).
[0215] The encoding device performs entropy encoding on the target syntax element (S610). The encoding device can encode an empty string of the target syntax element based on a regular coding-based (context-based) or bypass coding-based entropy coding technique such as CABAC (context-adaptive arithmetic coding) or CAVLC (context-adaptive variable length coding), and the output thereof can be included in the bitstream. The entropy encoding procedure can be performed by an entropy encoding processing unit (243) within the entropy encoding unit (240). As described above, the bitstream can be transmitted to the decoding device via a (digital) storage medium or a network.
[0216] Figure 19 illustrates an entropy decoding method according to embodiments.
[0217] As shown in FIG. 19, a decoding device (entropy decoding unit) can decode encoded image / video information. The image / video information may include partitioning-related information, prediction-related information (e.g., inter / intra prediction distinction information, intra-prediction mode information, inter-prediction mode information, etc.), residual information, in-loop filtering-related information, etc., or may include various syntax elements related thereto. Entropy coding may be performed in units of syntax elements. S700 to S710 may be performed by the entropy decoding unit (310) of the decoding device of FIG. 3 described above.
[0218] The decoding device performs binarization on the target syntax element (S700). Here, the binarization may be based on various binarization methods, such as the Truncated Rice binarization process and the Fixed-length binarization process, and the binarization method for the target syntax element may be predefined. The decoding device may derive available empty strings (empty string candidates) for available values of the target syntax element through the binarization process. The binarization process may be performed by the binarization unit (312) within the entropy decoding unit (310).
[0219] The decoding device performs entropy decoding on the target syntax element (S710). The decoding device sequentially decodes and parses each bin for the target syntax element from the input bit(s) in the bitstream, and compares the derived bin string with the available bin strings for the corresponding syntax element. If the derived bin string is equal to one of the available bin strings, the value corresponding to the bin string is derived as the value of the corresponding syntax element. If not, the next bit in the bitstream is further parsed and the above-described procedure is performed again. Through this process, it is possible to signal specific information (specific syntax element) using variable-length bits without using start bits or end bits in the bitstream. This allows relatively fewer bits to be allocated to low values, thereby improving overall coding efficiency.
[0220] The decoding device can decode each bin within the bin string from the bitstream based on a context-based or bypass-based entropy coding technique such as CABAC or CAVLC. The entropy decoding procedure can be performed by the entropy decoding processing unit (313) within the entropy decoding unit (310). As described above, the bitstream can include various information for image / video decoding. As described above, the bitstream can be transmitted to the decoding device via a (digital) storage medium or a network.
[0221] In this document, a table including syntax elements (syntax table) may be used to indicate signaling of information from an encoding device to a decoding device. The order of the syntax elements in the table including the syntax elements used in this document may indicate a parsing order of the syntax elements from a bitstream. The encoding device may configure and encode the syntax table so that the syntax elements can be parsed by a decoding device in the parsing order, and the decoding device may parse and decode the syntax elements of the corresponding syntax table from the bitstream in the parsing order to obtain values of the syntax elements.
[0222] Figure 20 illustrates a picture decoding method according to embodiments.
[0223] General Video Coding Procedure:
[0224] In image / video coding, the pictures that make up an image / video can be encoded / decoded according to a series of decoding orders. The picture order corresponding to the output order of the decoded pictures can be set differently from the decoding order, and based on this, not only forward prediction but also backward prediction can be performed during inter prediction.
[0225] Fig. 20 illustrates an example of a schematic picture decoding procedure to which the embodiments of the present document are applicable. In Fig. 20, S900 may be performed in the entropy decoding unit (310) of the decoding device described above in Fig. 3, S910 may be performed in the prediction unit (330), S920 may be performed in the residual processing unit (320), S930 may be performed in the addition unit (340), and S940 may be performed in the filtering unit (350). S900 may include the information decoding procedure described in this document, S910 may include the inter / intra prediction procedure described in this document, S920 may include the residual processing procedure described in this document, S930 may include the block / picture restoration procedure described in this document, and S940 may include the in-loop filtering procedure described in this document.
[0226] As shown in FIG. 20, the picture decoding procedure may roughly include a procedure for obtaining image / video information (through decoding) from a bitstream (S900), a procedure for restoring a picture (S910 to S930), and an in-loop filtering procedure for a restored picture (S940), as shown in the description of FIG. 3. The picture restoration procedure may be performed based on prediction samples and residual samples obtained through the inter / intra prediction (S910) and residual processing (S920, inverse quantization and inverse transformation for quantized transform coefficients) processes described in this document. A modified restored picture can be generated through an in-loop filtering procedure for a restored picture generated through a picture restoration procedure, and the modified restored picture can be output as a decoded picture and stored in a decoded picture buffer or memory (360) of a decoding device so that it can be used as a reference picture in an inter prediction procedure when decoding a subsequent picture. In some cases, the in-loop filtering procedure can be omitted, in which case the restored picture can be output as a decoded picture and stored in a decoded picture buffer or memory (360) of a decoding device so that it can be used as a reference picture in an inter prediction procedure when decoding a subsequent picture. The in-loop filtering procedure (S940) can include a deblocking filtering procedure, a sample adaptive offset (SAO) procedure, an adaptive loop filter (ALF) procedure, and / or a bi-lateral filter procedure as described above, and some or all of them can be omitted. Additionally, one or more of the deblocking filtering procedure, the sample adaptive offset (SAO) procedure, the adaptive loop filter (ALF) procedure, and the bilateral filter procedure may be applied sequentially, or all of them may be applied sequentially. For example, the SAO procedure may be performed after the deblocking filtering procedure is applied to the restored picture.Alternatively, for example, the ALF procedure may be performed after a deblocking filtering procedure has been applied to the restored picture. This may also be performed in the encoding device.
[0227] Figure 21 illustrates a picture encoding method according to embodiments.
[0228] Fig. 21 illustrates an example of a schematic picture encoding procedure to which the embodiments of the present document are applicable. In Fig. 21, S800 may be performed in the prediction unit (220) of the encoding device described above in Fig. 2, S810 may be performed in the residual processing unit (230), and S820 may be performed in the entropy encoding unit (240). S800 may include the inter / intra prediction procedure described in the present document, S810 may include the residual processing procedure described in the present document, and S820 may include the information encoding procedure described in the present document.
[0229] As shown in Fig. 21, the picture encoding procedure may include not only a procedure for encoding information for picture restoration (e.g., prediction information, residual information, partitioning information, etc.) and outputting it in the form of a bitstream as shown in the description for Fig. 02, but also a procedure for generating a restored picture for the current picture and a procedure (optional) for applying in-loop filtering to the restored picture. The encoding device can derive (corrected) residual samples from the quantized transform coefficients through the inverse quantization unit (234) and the inverse transformation unit (235), and can generate a restored picture based on the prediction samples and (corrected) residual samples, which are outputs of S800. The restored picture generated in this way may be identical to the restored picture generated by the decoding device described above. Through the in-loop filtering procedure for the restored picture, a modified restored picture can be generated, which can be stored in the decoded picture buffer or memory (270), and can be used as a reference picture in the inter prediction procedure when encoding a subsequent picture, as in the case of the decoding device. As described above, some or all of the in-loop filtering procedure may be omitted depending on the case. When the in-loop filtering procedure is performed, (in-loop) filtering-related information (parameters) can be encoded in the entropy encoding unit (240) and output in the form of a bitstream, and the decoding device can perform the in-loop filtering procedure in the same manner as the encoding device based on the filtering-related information.
[0230] This in-loop filtering procedure can reduce noise that occurs during image / video coding, such as blocking artifacts and ringing artifacts, and improve subjective / objective visual quality. Furthermore, by performing the in-loop filtering procedure on both the encoding device and the decoding device, the encoding device and the decoding device can produce identical prediction results, thereby increasing the reliability of picture coding and reducing the amount of data that must be transmitted for picture coding.
[0231] As described above, the picture restoration procedure can be performed not only in the decoding device but also in the encoding device. A restoration block can be generated based on intra-prediction / inter-prediction for each block, and a restoration picture including the restoration blocks can be generated. If the current picture / slice / tile group is an I picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based only on intra-prediction. On the other hand, if the current picture / slice / tile group is a P or B picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based on intra-prediction or inter-prediction. In this case, inter-prediction may be applied to some blocks in the current picture / slice / tile group, and intra-prediction may be applied to some blocks in the remaining blocks. The color component of a picture may include a luma component and a chroma component, and unless explicitly limited in this document, the methods and embodiments proposed in this document may be applied to the luma component and the chroma component.
[0232] Examples of coding hierarchies and structures:
[0233] The coded video / image according to this document can be processed according to the coding layers and structures described below, for example.
[0234] Figure 22 shows a hierarchical structure for coded images according to embodiments.
[0235] Figure 22 is a diagram showing the hierarchical structure for a coded image.
[0236] The coded video is divided into the video coding layer (VCL), which handles the decoding process of the video and the video itself, the subsystem that transmits and stores the coded information, and the network abstraction layer (NAL), which exists between the VCL and the subsystem and is responsible for network adaptation functions.
[0237] In VCL, VCL data containing compressed image data (slice data) can be generated, or a parameter set containing information such as a picture parameter set (PPS), a sequence parameter set (SPS), a video parameter set (VPS), etc., or an SEI (Supplemental Enhancement Information) message additionally required for the image decoding process can be generated.
[0238] In NAL, a NAL unit can be created by adding header information (NAL unit header) to an RBSP (Raw Byte Sequence Payload) generated from a VCL. At this time, RBSP refers to slice data, parameter sets, SEI messages, etc. generated from a VCL. The NAL unit header can include NAL unit type information that is specific to the RBSP data included in the NAL unit.
[0239] NAL units can be divided into VCL NAL units and non-VCL NAL units according to the RBSP generated from VCL. A VCL NAL unit can refer to a NAL unit that contains information about a video (slice data), and a non-VCL NAL unit can refer to a NAL unit that contains information necessary for decoding a video (parameter set or SEI message).
[0240] The above-described VCL NAL units and non-VCL NAL units can be transmitted over a network by attaching header information according to the data specifications of the lower system. For example, NAL units can be transformed into data formats of a certain standard, such as the H.266 / VVC file format, RTP (Real-time Transport Protocol), TS (Transport Stream), etc., and transmitted over various networks.
[0241] A NAL unit can be specified as a NAL unit type according to the RBSP data structure included in the NAL unit, and information about the NAL unit type can be stored and signaled in the NAL unit header.
[0242] For example, depending on whether a NAL unit contains information about a picture (slice data), it can be broadly classified into VCL NAL unit types and Non-VCL NAL unit types. The VCL NAL unit type can be classified according to the nature and type of the picture contained in the VCL NAL unit, and the Non-VCL NAL unit type can be classified according to the type of parameter set, etc.
[0243] Below are examples of NAL unit types specified by the type of parameter set included in the Non-VCL NAL unit type: APS (Adaptation Parameter Set) NAL unit: Type for a NAL unit that includes an APS. DPS (Decoding Parameter Set) NAL unit: Type for a NAL unit that includes a DPS. VPS (Video Parameter Set) NAL unit: Type for a NAL unit that includes a VPS. SPS (Sequence Parameter Set) NAL unit: Type for a NAL unit that includes an SPS. PPS (Picture Parameter Set) NAL unit: Type for a NAL unit that includes a PPS.
[0244] The above-described NAL unit types have syntax information for the NAL unit type, and the syntax information can be stored and signaled in the NAL unit header. For example, the syntax information can be nal_unit_type, and NAL unit types can be specified by the nal_unit_type value.
[0245] A slice header (slice header syntax) may include information / parameters that are commonly applicable to slices. An APS (APS syntax) or a PPS (PPS syntax) may include information / parameters that are commonly applicable to one or more slices or pictures. An SPS (SPS syntax) may include information / parameters that are commonly applicable to one or more sequences. A VPS (VPS syntax) may include information / parameters that are commonly applicable to multiple layers. A DPS (DPS syntax) may include information / parameters that are commonly applicable to the entire video. A DPS may include information / parameters related to the concatenation of CVSs (coded video sequences). In this document, a high level syntax (HLS) may include at least one of an APS syntax, a PPS syntax, an SPS syntax, a VPS syntax, a DPS syntax, and a slice header syntax.
[0246] In this document, the image / video information encoded from an encoding device to a decoding device and signaled in the form of a bitstream may include information related to partitioning within a picture, intra / inter prediction information, residual information, in-loop filtering information, etc., as well as information included in a slice header, information included in an APS, information included in the PPS, information included in an SPS, and / or information included in a VPS.
[0247] Coding descriptors:
[0248] The following descriptors represent the parsing process for each syntax element: ae(v): A context-adaptive arithmetic entropy coded syntax element. b(8): A byte containing a bit string (8 bits) of arbitrary pattern. The parsing process for this descriptor is specified by the return value of the read_bits(8) function. f(n): A bit string with a fixed pattern, written left-to-right, with the left bit first. The parsing process for this descriptor is specified by the return value of the read_bits(n) function. i(n): A signed integer with n bits. If n is "v" in the syntax table, the number of bits depends on the values of other syntax elements. The parsing process for this descriptor is specified by the return value of the read_bits(n) function, interpreted as a two's complement integer representation, most significant bit first. se(v): A signed integer, zeroth-order, Exp-Golomb coded syntax element, left bit first. The parsing process for this descriptor is specified by the order of k = 0. st(v): A null-terminated string encoded as Universal Coded Character Set (UCS) Transmission Format-8 (UTF-8) characters as specified in ISO / IEC 10646. The parsing process is as follows: st(v) moves the bitstream pointer (stringLength + 1) * 8 bit positions starting at a byte-aligned position in the bitstream, up to but not including the next byte-aligned byte equal to 0x00, starting at the current position, where stringLength is equal to the number of bytes returned. The st(v) syntax descriptor is only used in this specification when the current position in the bitstream is byte-aligned. tu(v): A truncated unary operator that uses up to maxVal bits. maxVal is defined in the semantics of the symtax element. u(n): An unsigned integer that uses n bits.In the syntax table, if n is "v", the number of bits depends on the values of other syntax elements. The parsing process for this descriptor is specified by the return value of the function read_bits(n), which is interpreted as the binary representation of an unsigned integer written most-significant bit first. ue(v): An unsigned integer of a 0th-order Golomb-encoded syntax element written left bit first. The parsing process for this descriptor is specified by setting the kth order to 0.
[0249] Below, high level syntax signaling and semantics are described with reference to each drawing.
[0250] Figures 23a, 23b, 23c, 23d, and 23e illustrate picture header structures (picture_header_structure) according to embodiments.
[0251] Picture header and slice header:
[0252] A coded picture may consist of one or more slices. Parameters describing a coded picture are conveyed within a picture header (PH), and parameters describing a slice are conveyed within a slice header. PH is conveyed as its own NAL unit type. SH is located at the beginning of a NAL unit containing the slice's payload (i.e., slice data). For detailed information on the syntax and semantics of PH and SH, see Section 7 of the VVC specification.
[0253] SEI Messages:
[0254] Neural-network post-filter SEI messages
[0255] General post-processing filtering process using NNPFs
[0256] The input to this process is the bitstream BitstreamToFilter. The output of this process is a list of NNPF output pictures ListNnpfOutputPics.
[0257] First, BitstreamToFilter is decoded, and the CroppedDecodedPictures list is set to a list of cropped decoded pictures in the order of the BitstreamToFilter decoding result output.
[0258] Second, the filtering process for a picture is in CroppedDecodedPictures and is called repeatedly in output order for each cropped decoded picture with one or more NNPFs enabled.
[0259] The order of the pictures in ListNnpfOutputPics is the output order.
[0260] Within ListNnpfOutputPics , there must be exactly one picture associated with a particular output time instance. If there are multiple NNPFs enabled for a particular picture in CroppedDecodedPictures and only one NNPF can be selected to apply (other NNPFs can also be selected), the above constraints apply regardless of which NNPF is applied to the particular picture.
[0261] Single picture filtering process using NNPF:
[0262] The filtering process is applied to each cropped decoded picture (called the current picture) belonging to CroppedDecodedPictures and having one or more NNPFs enabled.
[0263] When applying NNPF to the current picture, the filtered and / or interpolated picture is generated by NNPF by applying the NNPF process specified in the semantics of the NNPFFC SEI message to the current picture in a patch manner.
[0264] When applying NNPF to the current picture, the order of the pictures generated by NNPF by applying the NNPF process is the same as the output order stored in the output tensor of NNPF.
[0265] If the applied NNPF is the last NNPF applied to the current photo, the photos generated by the NNPF and the photos output from the NNPF process are included in ListNnpfOutputPics in the same order in which the photos were stored in the NNPF's output tensor.
[0266] Figures 24a, 24b, and 24c illustrate the syntax of a neural network post-filter characteristics SEI message according to embodiments.
[0267] The syntax of the NNPFC SEI message related to the neural network post-filter characteristics SEI message (NNPFC) is as shown in Figs. 24a, 24b, and 24c.
[0268] The NNPFC SEI message indicates a neural network that can be used as a post-processing filter. The use of a specified neural network post-processing filter (NNPF) for a given picture is indicated by the Neural Network Post-processing Filter Activation (NNPFA) SEI message.
[0269] Using this SEI message requires defining the following variables:
[0270] The input picture width and height in luma sample units (denoted as CroppedWidth and CroppedHeight, respectively).
[0271] An array of luma samples CroppedYPic[idx] and chroma sample arrays CroppedCbPic[idx] and CroppedCrPic[idx] (if present) of input pictures with indices idx in the range 0 to numInputPics - 1 (inclusive) used as input to NNPF.
[0272] BitDepthY, the bit depth of the luma sample array of the input picture.
[0273] BitDepthC, the bit depth of the chroma sample array (if any) of the input picture.
[0274] Chroma format indicator, represented by ChromaFormatIdc.
[0275] If nnpfc_auxiliary_inp_idc is 1, the filtering strength control value array StrengthControlVal[idx] contains real numbers in the range 0 to 1 (inclusive) for the input pictures whose index idx is in the range 0 to numInputPics - 1 (inclusive).
[0276] An input picture with index 0 corresponds to a picture for which the NNPF defined in this NNPFC SEI message is activated by the NNPFA SEI message. Input pictures with index i in the range of 1 to numInputPics-1 have precedence over the input picture with index i-1 in the output order.
[0277] The SubWidthC and SubHeightC variables are derived from ChromaFormatIdc.
[0278] There can be more than one NNPFC SEI message for the same picture. If two or more NNPFC SEI messages with different nnpfc_id values exist or are enabled for the same picture, the nnpfc_purpose and nnpfc_mode_idc values of those messages can be the same or different.
[0279] nnpfc_purpose represents the purpose of NNPF as specified in Table 6 (Definition of nnpfc_purpose). If (nnpfc_purpose & bitMask) is not 0, it indicates that NNPF has the purpose associated with the bitMask value in Table 6. If nnpfc_purpose is greater than 0 and (nnpfc_purpose & bitMask) is 0, the purpose associated with the bitMask value cannot be applied to NNPF. If nnpfc_purpose is 0, NNPF can be used as determined by the application.
[0280] The value of nnpfc_purpose is in the range 0 to 63 in bitstreams conforming to this version of this document. The values 64 to 65,535 (inclusive) for nnpfc_purpose are reserved for future use in ITU-T | ISO / IEC and are not present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_purpose in the range 64 to 65,535.
[0281] [Table 6]
[0282]
[0283] The variables chromaUpsamplingFlag, resolutionResamplingFlag, pictureRateUpsamplingFlag, bitDepthUpsamplingFlag, and colorizationFlag, which specify whether the purpose of NNPF includes chroma upsampling, resolution resampling, picture rate upsampling, bit depth upsampling, and colorization, respectively, are derived as follows.
[0284] chromaUpsamplingFlag = ((nnpfc_purpose & 0x02) > 0)? 1:0
[0285] resolutionResamplingFlag = ( ( nnpfc_purpose & 0x04 ) > 0 ) ? 1:0
[0286] pictureRateUpsamplingFlag = ((nnpfc_purpose & 0x08) > 0)? 1:0 (76)
[0287] bitDepthUpsamplingFlag = ( ( nnpfc_purpose & 0x10 ) > 0 ) ? 1:0
[0288] colourizationFlag = ( ( nnpfc_purpose & 0x20 ) > 0 ) ? 1:0
[0289] If the reserved value of nnpfc_purpose is used in the future in ITU-T | ISO / IEC, the syntax of this SEI message may be extended by syntax elements depending on whether nnpfc_purpose is equal to that value.
[0290] If ChromaFormatIdc is 3, chromaUpsamplingFlag becomes 0.
[0291] If ChromaFormatIdc or chromaUpsamplingFlag is non-zero, colorizationFlag will be 0.
[0292] If the input picture with pictureRateUpsamplingFlag equal to 1 and index 0 is associated with a frame packing array SEI message with fp_arrangement_type equal to 5, then all input pictures are associated with frame packing array SEI messages with fp_arrangement_type equal to 5 and the same fp_current_frame_is_frame0_flag value.
[0293] nnpfc_id contains an identification number that can be used to identify the NNPF. nnpfc_id values range from 0 to 2-2 (inclusive). nnpfc_id values 256 to 511 (inclusive) and 2-2 (inclusive) are reserved for future use by ITU-T | ISO / IEC. Decoders conforming to this version of this document will ignore NNPFC SEI messages with nnpfc_id values in the ranges 256 to 511 (inclusive) or 2-2 (inclusive).
[0294] If the NNPFC SEI message is the first NNPFC SEI message with a particular nnpfc_id value within the current CLVS in decoding order, the following applies:
[0295] This SEI message represents the basic NNPF.
[0296] This SEI message applies to the currently decoded picture and all subsequent decoded pictures in the current layer in output order up to the end of the current CLVS.
[0297] If nnpfc_base_flag is 1, it indicates that the SEI message specifies the base NNPF. If nnpf_base_flag is 0, it indicates that the SEI message specifies updates based on the base NNPF.
[0298] The following constraints apply to the nnpfc_base_flag value:
[0299] If the NNPFC SEI message is the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS, the nnpfc_base_flag value is equal to 1.
[0300] If NNPFC SEI message nnpfcB is not the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS and the nnpfc_base_flag value is 1, the NNPFC SEI message is a repetition of the first NNPFC SEI message nnpfcA with the same nnpfc_id value in decoding order. That is, the payload content of nnpfcB is identical to the payload content of nnpfcA.
[0301] If nnpfc_base_flag is 0, the following applies:
[0302] This SEI message defines an update relative to a previous default NNPF with the same nnpfc_id value in decoding order. Updates are not cumulative; each update is applied to the default NNPF. The default NNPF is the NNPF specified in the first NNPFC SEI message in decoding order and has a specific nnpfc_id value within the current CLVS. The NNPF defined in this SEI message is obtained by applying the updates defined in this SEI message relative to the default NNPF with the same nnpfc_id value.
[0303] This SEI message is for the currently decoded picture and all subsequent decoded pictures of the current layer (in output order), up to the end of the current CLVS or the decoded picture that follows the current decoded picture in output order within the current CLVS, and is associated with subsequent NNPFC SEI messages in decoding order, provided that nnpfc_base_flag is 0 and that a particular nnpfc_id value exists within the current CLVS (whichever is earlier).
[0304] If nnpfc_mode_idc is 0, this SEI message either contains an ISO / IEC 15938-17 bitstream specifying the default NNPF (if nnpfc_base_flag is 1) or is an updated version of the default NNPF with the same nnpfc_id value (if nnpfc_base_flag is 0).
[0305] If nnpfc_base_flag is 1, then nnpfc_mode_idc is 1, indicating that the base NNPF associated with the nnpfc_id value is a neural network identified by a URI denoted by nnpfc_uri and in a format identified by a tag URI nnpfc_tag_uri.
[0306] When nnpfc_base_flag is 0 and nnpfc_mode_idc is 1, it indicates that updates to the base NNPF with the same nnpfc_id value are defined by the URI indicated by nnpfc_uri and have a format identified by the tag URI nnpfc_tag_uri.
[0307] The value of nnpfc_mode_idc is in the range 0 to 1 inclusive in bitstreams conforming to this version of this document. nnpfc_mode_idc values from 2 to 255 (inclusive) are reserved for future use in ITU-T | ISO / IEC and are not present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_mode_idc in the range 2 to 255, inclusive. nnpfc_mode_idc values greater than 255 are not present in bitstreams conforming to this version of this document and are not reserved for future use.
[0308] nnpfc_reserved_zero_bit_a is equal to 0 in bitstreams that conform to this version of this document. The decoder ignores NNPFC SEI messages where nnpfc_reserved_zero_bit_a is not 0.
[0309] nnpfc_tag_uri contains a tag URI with the syntax and semantics specified in IETF RFC 4151 that identifies the format and associated information of the neural network used as an update to the base NNPF or to a base NNPF with the same nnpfc_id value specified in nnpfc_uri.
[0310] nnpfc_tag_uri allows you to uniquely identify the format of the neural network data specified in nnrpf_uri without the need for a central registry.
[0311] If nnpfc_tag_uri is "tag:iso.org,2023:15938-17", it indicates that the neural network data identified by nnpfc_uri complies with ISO / IEC 15938-17.
[0312] nnpfc_uri contains a URI with the syntax and semantics specified in IETF Internet Standard 66 that identifies the neural network used as the primary NNPF or as an update to a primary NNPF with the same nnpfc_id value.
[0313] When nnpfc_property_present_flag is 1, it indicates that syntax elements related to filter usage, input format, output format, and complexity are present. When nnpfc_property_present_flag is 0, it indicates that syntax elements related to filter usage, input format, output format, and complexity are not present.
[0314] If nnpfc_base_flag is 1, nnpfc_property_present_flag also becomes 1.
[0315] If nnpfc_property_present_flag is 0, the values of all syntax elements that can only be present when nnpfc_property_present_flag is 1 are inferred to be identical to the corresponding syntax elements of the NNPFC SEI message containing the underlying NNPF for which this SEI message provides updates.
[0316] If the NNPFC SEI message nnpfcCurr is not the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS, and does not duplicate the first NNPFC SEI message with that particular nnpfc_id (i.e., the nnpfc_base_flag value is 0), and the nnpfc_property_present_flag value is 1, the following constraints apply:
[0317] The nnpfc_purpose value of an NNPFC SEI message is equal to the nnpfc_purpose value of the first NNPFC SEI message in decoding order with that particular nnpfc_id value within the current CLVS.
[0318] In the NNPFC SEI message, the values of the syntax elements after nnpfc_property_present_flag and before nnpfc_complexity_info_present_flag in decoding order are identical to the values of the corresponding syntax elements of the first NNPFC SEI message with the corresponding nnpfc_id value within the current CLVS.
[0319] In the first NNPFC SEI message with the given nnpfc_id value within the current CLVS (denoted as nnpfcBase below), nnpfc_complexity_info_present_flag must be either 0 or both 1 in decoding order, and all of the following apply:
[0320] nnpfc_parameter_type_idc of nnpfcCurr is the same as nnpfc_parameter_type_idc of nnpfcBase.
[0321] If nnpfc_log2_parameter_bit_length_minus3 of nnpfcCurr exists, it is less than or equal to nnpfc_log2_parameter_bit_length_minus3 of nnpfcBase.
[0322] If nnpfc_num_parameters_idc of nnpfcBase is 0, nnpfc_num_parameters_idc of nnpfcCurr also becomes 0.
[0323] Otherwise (if nnpfc_num_parameters_idc of nnpfcBase is greater than 0), nnpfc_num_parameters_idc of nnpfcCurr is greater than 0 and less than or equal to nnpfc_num_parameters_idc of nnpfcBase.
[0324] If nnpfc_num_kmac_operations_idc of nnpfcBase is 0, nnpfc_num_kmac_operations_idc of nnpfcCurr also becomes 0.
[0325] Otherwise (if nnpfc_num_kmac_operations_idc of nnpfcBase is greater than 0), nnpfc_num_kmac_operations_idc of nnpfcCurr is greater than 0 and less than or equal to nnpfc_num_kmac_operations_idc of nnpfcBase.
[0326] If nnpfc_total_kilobyte_size of nnpfcBase is 0, nnpfc_total_kilobyte_size of nnpfcCurr also becomes 0.
[0327] Otherwise (if nnpfc_total_kilobyte_size of nnpfcBase is greater than 0), nnpfc_total_kilobyte_size of nnpfcCurr is greater than 0 and less than or equal to nnpfc_total_kilobyte_size of nnpfcBase.
[0328] nnpfc_num_input_pics_minus1 + 1 represents the number of pictures used as input to NNPF. The value of nnpfc_num_input_pics_minus1 ranges from 0 to 63. If pictureRateUpsamplingFlag is 1, the value of nnpfc_num_input_pics_minus1 is greater than 0.
[0329] The variable numInputPics, which specifies the number of pictures used as input to NNPF, is derived as follows.
[0330] numInputPics = nnpfc_num_input_pics_minus1 + 1 (77)
[0331] When nnpfc_input_pic_output_flag[ i ] is 1, it indicates that NNPF generates the corresponding output picture for the i-th input picture. When nnpfc_input_pic_output_flag[ i ] is 0, it indicates that NNPF does not generate the corresponding output picture for the i-th input picture. When nnpfc_num_input_pics_minus1 is 0, nnpfc_input_pic_output_flag
[0000] is inferred to be 1. When pictureRateUpsamplingFlag is 0 and nnpfc_num_input_pics_minus1 is greater than 0, nnpfc_input_pic_output_flag[ i ] is equal to 1 for at least one value of i in the range of 0 to nnpfc_num_input_pics_minus1, inclusive.
[0332] If nnpfc_absent_input_pic_zero_flag is 1, it indicates that NNPF should represent input pictures that are not in the bitstream as a sample array with sample values of 0. If nnpfc_absent_input_pic_flag is 0, it indicates that NNPF should represent input pictures that are not in the bitstream as the closest input picture in output order within the bitstream.
[0333] nnpfc_out_sub_c_flag indicates the values of the outSubWidthC and outSubHeightC variables when chromaUpsamplingFlag is 1. When nnpfc_out_sub_c_flag is 1, it indicates that outSubWidthC is 1 and outSubHeightC is 1. When nnpfc_out_sub_c_flag is 0, it indicates that outSubWidthC is 2 and outSubHeightC is 1. If ChromaFormatIdc is 2 and nnpfc_out_sub_c_flag is present, the value of nnpfc_out_sub_c_flag will be 1.
[0334] nnpfc_out_colour_format_idc specifies the colour format of the NNPF output when colourizationFlag is 1, and consequently the values of the outSubWidthC and outSubHeightC variables. If nnpfc_out_colour_format_idc is 1, the colour format of the NNPF output is 4:2:0, and both outSubWidthC and outSubHeightC are 2. If nnpfc_out_colour_format_idc is 2, the colour format of the NNPF output is 4:2:2, and outSubWidthC is 2 and outSubHeightC is 1. If nnpfc_out_colour_format_idc is 3, the colour format of the NNPF output is 4:4:4, and both outSubWidthC and outSubHeightC are 1. The value of nnpfc_out_colour_format_idc is not 0.
[0335] If both chromaUpsamplingFlag and colorizationFlag are 0, outSubWidthC and outSubHeightC are inferred as follows: SubWidthC and SubHeightC.
[0336] nnpfc_pic_width_num_minus1 + 1 and nnpfc_pic_width_denom_minus1 + 1 represent the numerator and denominator, respectively, of the resampling ratio of the NNPF output picture width relative to CroppedWidth. The value of (nnpfc_pic_width_num_minus1 + 1) / (nnpfc_pic_width_denom_minus1 + 1) ranges from 1 to 16, inclusive. If nnpfc_pic_width_num_minus1 and nnpfc_pic_width_denom_minus1 are absent, the values of nnpfc_pic_width_num_minus1 and nnpfc_pic_width_denom_minus1 are both inferred to be 0.
[0337] The nnpfcOutputPicWidth variable, which represents the width of the luma sample array of the picture resulting from applying the NNPF identified by nnpfc_id to the input picture, is derived as follows.
[0338] nnpfcOutputPicWidth = Ceil(CroppedWidth * (78)
[0339] ( nnpfc_pic_width_num_minus1 + 1 ) / ( nnpfc_pic_width_denom_minus1 + 1 ) )
[0340] For bitstream conformance, the value of nnpfcOutputPicWidth % outSubWidthC is equal to 0.
[0341] nnpfc_pic_height_num_minus1 + 1 and nnpfc_pic_height_denom_minus1 + 1 represent the numerator and denominator of the resampling ratio of the NNPF output picture height relative to the CroppedHeight, respectively. The value of ( nnpfc_pic_height_num_minus1 + 1 ) / ( nnpfc_pic_height_denom_minus1 + 1 ) ranges from 1 / 16 to 16. If nnpfc_pic_height_num_minus1 and nnpfc_pic_height_denom_minus1 are absent, the values of nnpfc_pic_height_num_minus1 and nnpfc_pic_height_denom_minus1 are both inferred to be 0.
[0342] The nnpfcOutputPicHeight variable, which represents the height of the luma sample array of the picture resulting from applying the NNPF identified by nnpfc_id to the input picture, is derived as follows.
[0343] nnpfcOutputPicHeight = Ceil( CroppedHeight * (79)
[0344] (nnpfc_pic_height_num_minus1 + 1) / (nnpfc_pic_height_denom_minus1 + 1) )
[0345] For bitstream conformance, the value of nnpfcOutputPicHeight % outSubHeightC is equal to 0.
[0346] If nnpfc_pic_width_num_minus1, nnpfc_pic_width_denom_minus1, nnpfc_pic_height_num_minus1, nnpfc_pic_height_denom_minus1 exist, then at least one of the following is true:
[0347] The value of nnpfcOutputPicWidth is not equal to CroppedWidth.
[0348] The value of nnpfcOutputPicHeight is not equal to CroppedHeight.
[0349] nnpfc_interpolated_pics[ i ] represents the number of interpolated pictures generated by NNPF between the i-th and ( i + 1 )-th pictures used as input to NNPF. The value of nnpfc_interpolated_pics[ i ] ranges from 0 to 63. The value of nnpfc_interpolated_pics[ i ] is greater than 0 for at least one value of i in the range of 0 to nnpfc_num_input_pics_minus1 - 1.
[0350] The variables NumInpPicsInOutputTensor, which exist in the output tensor of NNPF and specify the number of pictures that have the corresponding input picture, InpIdx[ idx ], which exist in the output tensor of NNPF and specify the input picture index of the idxth picture that has the corresponding input picture, and numOutputPics, which specify the total number of pictures that exist in the output tensor of NNPF, are derived as follows.
[0351] for( i = 0, numOutputPics = 0; i < numInputPics; i++ )
[0352] if(nnpfc_input_pic_output_flag[i]) {
[0353] InpIdx[ numOutputPics ] = i
[0354] numOutputPics++
[0355] }
[0356] NumInpPicsInOutputTensor = numOutputPics
[0357] if(pictureRateUpsamplingFlag)
[0358] for( i = 0; i <= numInputPics - 2; i++ )
[0359] numOutputPics += nnpfc_interpolated_pics[ i ]
[0360] If nnpfc_component_last_flag is 1, it indicates that the last dimension of the input tensor for NNPF, inputTensor, and the output tensor generated by NNPF, outputTensor, is used for the current channel. If nnpfc_component_last_flag is 0, it indicates that the third dimension of the input tensor for NNPF, inputTensor, and the output tensor generated by NNPF, outputTensor, is used for the current channel.
[0361] The first dimension of the input and output tensors is used for batch indices, a practice used in some neural network frameworks. The semantics of this SEI message use a batch size with a batch index of 0, but it's up to the postprocessing implementation to determine the batch size used as input for neural network inference.
[0362] For example, if nnpfc_inp_order_idc is 3 and nnpfc_auxiliary_inp_idc is 1, the input tensor has a total of 7 channels, including 4 luma matrices, 2 chroma matrices, and 1 auxiliary input matrix. In this case, the DeriveInputTensors() process derives the 7 channels of the input tensor one by one, and when a particular channel among these channels is processed, that channel is called the current channel during the process.
[0363] nnpfc_inp_format_idc indicates how to convert the sample values of the input picture into input values for NNPF. If nnpfc_inp_format_idc is 0, the input values of NNPF are real numbers, and the InpY( ) and InpC( ) functions are expressed as follows.
[0364] InpY(x) = x / ( ( 1 << BitDepthY ) - 1 )
[0365] InpC( x )= x / ( ( 1 << BitDepthC ) - 1 )
[0366] When nnpfc_inp_format_idc is 1, the input value of NNPF is an unsigned integer, and the InpY( ) and InpC( ) functions are expressed as follows.
[0367] shiftY = BitDepthY - inpTensorBitDepthY
[0368] if( inpTensorBitDepthY >= BitDepthY)
[0369] InpY(x) = x << ( inpTensorBitDepthY - BitDepthY )
[0370] otherwise
[0371] InpY(x) = Clip3(0, (1 << inpTensorBitDepthY ) - 1, (x + (1 << (shiftY - 1 ) ) ) >> shiftY )
[0372] shiftC = BitDepthC - inpTensorBitDepthC
[0373] If inpTensorBitDepthC >= BitDepthC
[0374] InpC(x) = x << ( inpTensorBitDepthC - BitDepthC )
[0375] otherwise
[0376] InpC(x) = Clip3(0, (1 << inpTensorBitDepthC ) - 1, (x + (1 << (shiftC - ) ) ) >> shiftC )
[0377] The variable inpTensorBitDepthY is derived from the syntax element nnpfc_inp_tensor_luma_bitdepth_minus8 specified below. The variable inpTensorBitDepthC is derived from the syntax element nnpfc_inp_tensor_chroma_bitdepth_minus8 specified below.
[0378] If the nnpfc_inp_format_idc value is greater than 1, it is reserved for future specifications in ITU-T | ISO / IEC and is not present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document will ignore NNPFC SEI messages containing reserved values for nnpfc_inp_format_idc.
[0379] A value of nnpfc_auxiliary_inp_idc greater than 0 indicates that the input tensor of NNPF contains auxiliary input data. A value of nnpfc_auxiliary_inp_idc of 0 indicates that the input tensor contains no auxiliary input data. A value of nnpfc_auxiliary_inp_idc of 1 indicates that the auxiliary input data is derived as specified in the formula (inpTensorBitDepthY = nnpfc_inp_tensor_luma_bitdepth_minus8 + 8).
[0380] The value of nnpfc_auxiliary_inp_idc is in the range 0 to 1 inclusive in bitstreams conforming to this version of this document. Values of nnpfc_auxiliary_inp_idc from 2 to 255 (inclusive) are reserved for future use in ITU-T | ISO / IEC and are not present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages for which nnpfc_auxiliary_inp_idc is in the range 2 to 255, inclusive. Values of nnpfc_auxiliary_inp_idc greater than 255 are not present in bitstreams conforming to this version of this document and are not reserved for future use.
[0381] nnpfc_inp_order_idc indicates how to form an input tensor for NNPF by sorting the sample array of input pictures.
[0382] The value of nnpfc_inp_order_idc is in the range 0 to 3 inclusive in bitstreams conforming to this version of this document. The value of nnpfc_inp_order_idc in the range 4 to 255 is reserved for future use in ITU-T | ISO / IEC and is not present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document shall ignore NNPFC SEI messages with nnpfc_inp_order_idc values in the range 4 to 255 inclusive. nnpfc_inp_order_idc values greater than 255 are not present in bitstreams conforming to this version of this document and are not reserved for future use.
[0383] If ChromaFormatIdc is not 1, nnpfc_inp_order_idc is not 3.
[0384] If ChromaFormatIdc is 0, nnpfc_inp_order_idc is not 0.
[0385] If chromaUpsamplingFlag is 1, nnpfc_inp_order_idc is not 0.
[0386] Table 7 (Description of nnpfc_inp_order_idc values) provides a description of the nnpfc_inp_order_idc values.
[0387] [Table 7]
[0388]
[0389] Figure 25 illustrates a process of deriving a luma channel from a luma component according to embodiments.
[0390] Figure 25 is an example of deriving four luma channels (right) from a luma component when nnpfc_inp_order_idc is 3.
[0391] nnpfc_inp_tensor_luma_bitdepth_minus8 + 8 represents the bit depth of the luma sample values in the input integer tensor. The value of inpTensorBitDepthY is derived as follows.
[0392] inpTensorBitDepthY = nnpfc_inp_tensor_luma_bitdepth_minus8 + 8 (85)
[0393] For bitstream conformance, the value of nnpfc_inp_tensor_luma_bitdepth_minus8 is in the range 0 to 24 (inclusive).
[0394] nnpfc_inp_tensor_chroma_bitdepth_minus8 + 8 represents the bit depth of the chroma sample values in the input integer tensor. The value of inpTensorBitDepthC is derived as follows.
[0395] inpTensorBitDepthC = nnpfc_inp_tensor_chroma_bitdepth_minus8 + 8
[0396] For bitstream conformance, the value of nnpfc_inp_tensor_chroma_bitdepth_minus8 is in the range of 0 to 24.
[0397] When nnpfc_auxiliary_inp_idc is 1, the variable strengthControlScaledVal is derived as follows.
[0398] for( i = 0; i < numInputPics; i++ )
[0399] if(nnpfc_inp_format_idc = = 1)
[0400] if( nnpfc_inp_order_idc = = 0 | | nnpfc_inp_order_idc = = 2 | |
[0401] nnpfc_inp_order_idc = = 3 )
[0402] strengthControlScaledVal[ i ] =
[0403] Floor ( StrengthControlVal[ i ] * ( ( 1 << inpTensorBitDepthY ) - 1 ) )
[0404] else if(nnpfc_inp_order_idc = = 1)
[0405] strengthControlScaledVal[ i ] =
[0406] Floor ( StrengthControlVal[ i ] * ( ( 1 << inpTensorBitDepthC ) - 1 ) )
[0407] otherwise
[0408] strengthControlScaledVal[i] = StrengthControlVal[i]
[0409] A patch is a rectangular array of samples extracted from a component of a picture (e.g., a luma or chroma component).
[0410] The DeriveInputTensors() process derives an input tensor inputTensor for the given vertical sample coordinates cTop and horizontal sample coordinates cLeft, which represents the top-left sample location of the sample patch contained in the input tensor, and is defined as follows:
[0411] for( i = 0; i < numInputPics; i++ ) {
[0412] if(nnpfc_inp_order_idc = = 0)
[0413] for( yP = -nnpfc_overlap; yP < inpPatchHeight + nnpfc_overlap; yP++)
[0414] for( xP = -nnpfc_overlap; xP < inpPatchWidth + nnpfc_overlap; xP++ ) {
[0415] inpVal = InpY( InpSampleVal( cTop + yP, cLeft + xP, CroppedHeight,
[0416] CroppedWidth, CroppedYPic[ i ], 0 ) )
[0417] yPovlp = yP + nnpfc_overlap
[0418] xPovlp = xP + nnpfc_overlap
[0419] if( !nnpfc_component_last_flag )
[0420] inputTensor
[0000] [ i ]
[0000] [ yPovlp ][ xPovlp ] = inpVal
[0421] else
[0422] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0000] = inpVal
[0423] if( nnpfc_auxiliary_inp_idc = = 1 )
[0424] if( !nnpfc_component_last_flag )
[0425] inputTensor
[0000] [ i ]
[0001] [ yPovlp ][ xPovlp ] = strengthControlScaledVal[ i ]
[0426] else
[0427] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0001] = strengthControlScaledVal[ i ]
[0428] }
[0429] else if( nnpfc_inp_order_idc = = 1 )
[0430] for( yP = -nnpfc_overlap; yP < inpPatchHeight + nnpfc_overlap; yP++)
[0431] for( xP = -nnpfc_overlap; xP < inpPatchWidth + nnpfc_overlap; xP++ ) {
[0432] inpCbVal = InpC( InpSampleVal( cTop + yP, cLeft + xP, CroppedHeight / SubHeightC,
[0433] CroppedWidth / SubWidthC, CroppedCbPic[ i ], 1 ) )
[0434] inpCrVal = InpC( InpSampleVal( cTop + yP, cLeft + xP, CroppedHeight / SubHeightC,
[0435] CroppedWidth / SubWidthC, CroppedCrPic[ i ], 2 ) )
[0436] yPovlp = yP + nnpfc_overlap
[0437] xPovlp = xP + nnpfc_overlap
[0438] if( !nnpfc_component_last_flag ) {
[0439] inputTensor
[0000] [ i ]
[0000] [ yPovlp ][ xPovlp ] = inpCbVal
[0440] inputTensor
[0000] [ i ]
[0001] [ yPovlp ][ xPovlp ] = inpCrVal
[0441] } else {
[0442] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0000] = inpCbVal
[0443] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0001] = inpCrVal
[0444] }
[0445] if( nnpfc_auxiliary_inp_idc = = 1 )
[0446] if( !nnpfc_component_last_flag )
[0447] inputTensor
[0000] [ i ]
[0002] [ yPovlp ][ xPovlp ] = strengthControlScaledVal[ i ]
[0448] else
[0449] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0002] = strengthControlScaledVal[ i ]
[0450] }
[0451] else if( nnpfc_inp_order_idc = = 2 )
[0452] for( yP = -nnpfc_overlap; yP < inpPatchHeight + nnpfc_overlap; yP++)
[0453] for( xP = -nnpfc_overlap; xP < inpPatchWidth + nnpfc_overlap; xP++ ) {
[0454] yY = cTop + yP
[0455] xY = cLeft + xP
[0456] yC = yY / SubHeightC
[0457] xC = xY / SubWidthC
[0458] inpYVal = InpY( InpSampleVal( yY, xY, CroppedHeight,
[0459] CroppedWidth, CroppedYPic[ i ], 0 ) )
[0460] inpCbVal = InpC( InpSampleVal( yC, xC, CroppedHeight / SubHeightC,
[0461] CroppedWidth / SubWidthC, CroppedCbPic[ i ], 1 ) )
[0462] inpCrVal = InpC( InpSampleVal( yC, xC, CroppedHeight / SubHeightC,
[0463] CroppedWidth / SubWidthC, CroppedCrPic[ i ], 2 ) )
[0464] yPovlp = yP + nnpfc_overlap
[0465] xPovlp = xP + nnpfc_overlap
[0466] if( !nnpfc_component_last_flag ) {
[0467] inputTensor
[0000] [ i ]
[0000] [ yPovlp ][ xPovlp ] = inpYVal
[0468]
[0469]
[0470] inputTensor
[0000] [ i ]
[0001] [ yPovlp ][ xPovlp ] = inpCbVal
[0471] inputTensor
[0000] [ i ]
[0002] [ yPovlp ][ xPovlp ] = inpCrVal
[0472] } else {
[0473] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0000] = inpYVal
[0474]
[0475]
[0476] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0001] = inpCbVal
[0477] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0002] = inpCrVal
[0478] }
[0479] if( nnpfc_auxiliary_inp_idc = = 1 )
[0480] if( !nnpfc_component_last_flag )
[0481] inputTensor
[0000] [ i ]
[0003] [ yPovlp ][ xPovlp ] = strengthControlScaledVal[ i ]
[0482] else
[0483] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0003] = strengthControlScaledVal[ i ]
[0484] }
[0485] else if( nnpfc_inp_order_idc = = 3 )
[0486] for( yP = -nnpfc_overlap; yP < inpPatchHeight + nnpfc_overlap; yP++)
[0487] for( xP = -nnpfc_overlap; xP < inpPatchWidth + nnpfc_overlap; xP++ ) {
[0488] yTL = cTop + yP * 2
[0489] xTL = cLeft + xP * 2
[0490] yBR = yTL + 1
[0491] xBR = xTL + 1
[0492] yC = cTop / 2 + yP
[0493] xC = cLeft / 2 + xP
[0494] inpTLVal = InpY( InpSampleVal( yTL, xTL, CroppedHeight,
[0495] CroppedWidth, CroppedYPic[ i ], 0 ) )
[0496] inpTRVal = InpY( InpSampleVal( yTL, xBR, CroppedHeight,
[0497] CroppedWidth, CroppedYPic[ i ], 0 ) )
[0498] inpBLVal = InpY( InpSampleVal( yBR, xTL, CroppedHeight,
[0499] CroppedWidth, CroppedYPic[ i ], 0 ) )
[0500] inpBRVal = InpY( InpSampleVal( yBR, xBR, CroppedHeight,
[0501] CroppedWidth, CroppedYPic[ i ], 0 ) )
[0502] inpCbVal = InpC( InpSampleVal( yC, xC, CroppedHeight / 2,
[0503] CroppedWidth / 2, CroppedCbPic[ i ], 1 ) )
[0504] inpCrVal = InpC( InpSampleVal( yC, xC, CroppedHeight / 2,
[0505] CroppedWidth / 2, CroppedCrPic[ i ], 2 ) )
[0506] yPovlp = yP + nnpfc_overlap
[0507] xPovlp = xP + nnpfc_overlap
[0508] if( !nnpfc_component_last_flag ) {
[0509] inputTensor
[0000] [ i ]
[0000] [ yPovlp ][ xPovlp ] = inpTLVal
[0510] inputTensor
[0000] [ i ]
[0001] [ yPovlp ][ xPovlp ] = inpTRVal
[0511] inputTensor
[0000] [ i ]
[0002] [ yPovlp ][ xPovlp ] = inpBLVal
[0512] inputTensor
[0000] [ i ]
[0003] [ yPovlp ][ xPovlp ] = inpBRVal
[0513] inputTensor
[0000] [ i ]
[0004] [ yPovlp ][ xPovlp ] = inpCbVal
[0514] inputTensor
[0000] [ i ]
[0005] [ yPovlp ][ xPovlp ] = inpCrVal
[0515] } else {
[0516] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0000] = inpTLVal
[0517] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0001] = inpTRVal
[0518] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0002] = inpBLVal
[0519] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0003] = inpBRVal
[0520] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0004] = inpCbVal
[0521] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0005] = inpCrVal
[0522] }
[0523] if( nnpfc_auxiliary_inp_idc = = 1 )
[0524] if( !nnpfc_component_last_flag )
[0525] inputTensor
[0000] [ i ]
[0006] [ yPovlp ][ xPovlp ] = strengthControlScaledVal[ i ]
[0526] else
[0527] inputTensor
[0000] [ i ][ yPovlp ][ xPovlp ]
[0006] = strengthControlScaledVal[ i ]
[0528] }
[0529] }
[0530] If nnpfc_out_format_idc is 0, it indicates that the sample values output by NNPF are real numbers, and that the range of values from 0 to 1 is linearly mapped to the range of unsigned integer values from 0 to (1 << bitDepth) - 1 for the desired bit depth bitDepth for subsequent postprocessing or display.
[0531] If nnpfc_out_format_idc is 1, it indicates that the luma sample values output from NNPF are unsigned integers from 0 to (1 << outTensorBitDepthY) - 1, and the chroma sample values output from NNPF are unsigned integers from 0 to (1 << outTensorBitDepthC) - 1.
[0532] nnpfc_out_format_idc values greater than 1 are reserved for future specifications in ITU-T | ISO / IEC and must not be included in bitstreams conforming to this version of this document. Decoders conforming to this version of this document ignore NNPFC SEI messages containing reserved nnpfc_out_format_idc values.
[0533] nnpfc_out_order_idc indicates the output order of samples generated by NNPF.
[0534] The value of nnpfc_out_order_idc is in the range 0 to 3 inclusive in bitstreams conforming to this version of this document. nnpfc_out_order_idc values from 4 to 255 (inclusive) are reserved for future use in ITU-T | ISO / IEC and are not present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document will ignore NNPFC SEI messages with nnpfc_out_order_idc in the range 4 to 255, inclusive. nnpfc_out_order_idc values greater than 255 are not present in bitstreams conforming to this version of this document and are not reserved for future use.
[0535] If chromaUpsamplingFlag is 1, nnpfc_out_order_idc cannot be 0 or 3.
[0536] If colorizationFlag is 1, nnpfc_out_order_idc cannot be 0.
[0537] Table 8 (Description of nnpfc_out_order_idc values) provides a description of the nnpfc_out_order_idc values.
[0538] [Table 8]
[0539]
[0540] nnpfc_out_tensor_luma_bitdepth_minus8 + 8 represents the bit depth of the luma sample values in the output integer tensor. The value of nnpfc_out_tensor_luma_bitdepth_minus8 ranges from 0 to 24. The value of outTensorBitDepthY is derived as follows.
[0541] outTensorBitDepthY = nnpfc_out_tensor_luma_bitdepth_minus8 + 8
[0542] nnpfc_out_tensor_chroma_bitdepth_minus8 + 8 represents the bit depth of the chroma sample values in the output integer tensor. The value of nnpfc_out_tensor_chroma_bitdepth_minus8 ranges from 0 to 24. The value of outTensorBitDepthC is derived as follows.
[0543] outTensorBitDepthC = nnpfc_out_tensor_chroma_bitdepth_minus8 + 8
[0544] If bitDepthUpsamplingFlag is 1, the value of nnpfc_out_format_idc must be 1 and meet one or more of the following conditions:
[0545] nnpfc_out_tensor_luma_bitdepth_minus8 exists and outTensorBitDepthY is greater than BitDepthY.
[0546] nnpfc_out_tensor_chroma_bitdepth_minus8 exists and outTensorBitDepthC is greater than BitDepthC.
[0547] If nnpfc_inp_tensor_luma_bitdepth_minus8, nnpfc_inp_tensor_chroma_bitdepth_minus8, nnpfc_out_tensor_luma_bitdepth_minus8, nnpfc_out_tensor_chroma_bitdepth_minus8 exist and outTensorBitDepthY is greater than inpTensorBitDepthY, then outTensorBitDepthC cannot be less than inpTensorBitDepthC. If nnpfc_inp_tensor_luma_bitdepth_minus8, nnpfc_inp_tensor_chroma_bitdepth_minus8, nnpfc_out_tensor_luma_bitdepth_minus8, nnpfc_out_tensor_chroma_bitdepth_minus8 exist and outTensorBitDepthC is greater than inpTensorBitDepthC, outTensorBitDepthY cannot be less than inpTensorBitDepthY.
[0548] The StoreOutputTensors() process, which derives sample values of filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor outputTensor for the given vertical sample coordinates cTop and the horizontal sample coordinates cLeft, which specify the upper-left sample location of the sample patch contained in the input tensor, is expressed as follows.
[0549] for( i = 0; i < numOutputPics; i++ ) {
[0550] if(nnpfc_out_order_idc = = 0)
[0551] for(yP = 0; yP < outPatchHeight; yP++)
[0552] for( xP = 0; xP < outPatchWidth; xP++ ) {
[0553] yY = cTop * outPatchHeight / inpPatchHeight + yP
[0554] xY = cLeft * outPatchWidth / inpPatchWidth + xP
[0555] if ( yY < nnpfcOutputPicHeight && xY < nnpfcOutputPicWidth )
[0556] if( !nnpfc_component_last_flag )
[0557] FilteredYPic[ i ][ xY ][yY ] = outputTensor
[0000] [ i ]
[0000] [ yP ][ xP ]
[0558] else
[0559] FilteredYPic[ i ][ xY ][ yY ] = outputTensor
[0000] [ i ][ yP ][ xP ]
[0000] }
[0560] else if( nnpfc_out_order_idc = = 1 ) (91)
[0561] for( yP = 0; yP < outPatchCHeight; yP++)
[0562] for( xP = 0; xP < outPatchCWidth; xP++ ) {
[0563] xSrc = cLeft * horCScaling + xP
[0564] ySrc = cTop * verCScaling + yP
[0565] if ( ySrc < nnpfcOutputPicHeight / outSubHeightC &&
[0566] xSrc < nnpfcOutputPicWidth / outSubWidthC )
[0567] if( !nnpfc_component_last_flag ) {
[0568] FilteredCbPic[ i ][ xSrc ][ ySrc ] = outputTensor
[0000] [ i ]
[0000] [ yP ][ xP ]
[0569] FilteredCrPic[ i ][ xSrc ][ ySrc ] = outputTensor
[0000] [ i ]
[0001] [ yP ][ xP ]
[0570] } else {
[0571] FilteredCbPic[ i ][ xSrc ][ ySrc ] = outputTensor
[0000] [ i ][ yP ][ xP ]
[0000]
[0572] FilteredCrPic[ i ][ xSrc ][ ySrc ] = outputTensor
[0000] [ i ][ yP ][ xP ]
[0001]
[0573] }
[0574] }
[0575] else if( nnpfc_out_order_idc = = 2 )
[0576] for( yP = 0; yP < outPatchHeight; yP++)
[0577] for( xP = 0; xP < outPatchWidth; xP++ ) {
[0578] yY = cTop * outPatchHeight / inpPatchHeight + yP
[0579] xY = cLeft * outPatchWidth / inpPatchWidth + xP
[0580] yC = yY / outSubHeightC
[0581] xC = xY / outSubWidthC
[0582] yPc = ( yP / outSubHeightC ) * outSubHeightC
[0583] xPc = ( xP / outSubWidthC ) * outSubWidthC
[0584] if ( yY < nnpfcOutputPicHeight && xY < nnpfcOutputPicWidth )
[0585] if( !nnpfc_component_last_flag ) {
[0586] FilteredYPic[ i ][ xY ][ yY ] = outputTensor
[0000] [ i ]
[0000] [ yP ][ xP ]
[0587] FilteredCbPic[ i ][ xC ][ yC ] = outputTensor
[0000] [ i ]
[0001] [ yPc ][ xPc ]
[0588] FilteredCrPic[ i ][ xC ][ yC ] = outputTensor
[0000] [ i ]
[0002] [ yPc ][ xPc ]
[0589] } else {
[0590] FilteredYPic[ i ][ xY ][ yY ] = outputTensor
[0000] [ i ][ yP ][ xP ]
[0000]
[0591] FilteredCbPic[ i ][ xC ][ yC ] = outputTensor
[0000] [ i ][ yPc ][ xPc ]
[0001]
[0592] FilteredCrPic[ i ][ xC ][ yC ] = outputTensor
[0000] [ i ][ yPc ][ xPc ]
[0002]
[0593] }
[0594] }
[0595] else if( nnpfc_out_order_idc = = 3 )
[0596] for( yP = 0; yP < outPatchHeight; yP++ )
[0597] for( xP = 0; xP < outPatchWidth; xP++ ) {
[0598] ySrc = cTop / 2 * outPatchHeight / inpPatchHeight + yP
[0599] xSrc = cLeft / 2 * outPatchWidth / inpPatchWidth + xP
[0600] if ( ySrc < nnpfcOutputPicHeight / 2 &&
[0601] xSrc < nnpfcOutputPicWidth / 2 )
[0602] if( !nnpfc_component_last_flag ) {
[0603] FilteredYPic[ i ][ xSrc * 2 ][ ySrc * 2 ] = outputTensor
[0000] [ i ]
[0000] [ yP ][ xP ]
[0604] FilteredYPic[ i ][ xSrc * 2 + 1 ][ ySrc * 2 ] = outputTensor
[0000] [ i ]
[0001] [ yP ][ xP ]
[0605] FilteredYPic[ i ][ xSrc * 2 ][ ySrc * 2 + 1 ] = outputTensor
[0000] [ i ]
[0002] [ yP ][ xP ]
[0606] FilteredYPic[ i ][ xSrc * 2 + 1][ ySrc * 2 + 1 ] = outputTensor
[0000] [ i ]
[0003] [ yP ][ xP ]
[0607] FilteredCbPic[ i ][ xSrc ][ ySrc ] = outputTensor
[0000] [ i ]
[0004] [ yP ][ xP ]
[0608] FilteredCrPic[ i ][ xSrc ][ ySrc ] = outputTensor
[0000] [ i ]
[0005] [ yP ][ xP ]
[0609] } else {
[0610] FilteredYPic[ i ][ xSrc * 2 ][ ySrc * 2 ] = outputTensor
[0000] [ i ][ yP ][ xP ]
[0000]
[0611] FilteredYPic[ i ][ xSrc * 2 + 1 ][ ySrc * 2 ] = outputTensor
[0000] [ i ][ yP ][ xP ]
[0001]
[0612] FilteredYPic[ i ][ xSrc * 2 ][ ySrc * 2 + 1 ] = outputTensor
[0000] [ i ][ yP ][ xP ]
[0002]
[0613] FilteredYPic[ i ][ xSrc * 2 + 1][ ySrc * 2 + 1 ] = outputTensor
[0000] [ i ][ yP ][ xP ]
[0003]
[0614] FilteredCbPic[ i ][ xSrc ][ ySrc ] = outputTensor
[0000] [ i ][ yP ][ xP ]
[0004]
[0615] FilteredCrPic[ i ][ xSrc ][ ySrc ] = outputTensor
[0000] [ i ][ yP ][ xP ]
[0005]
[0616] }
[0617] }
[0618] }
[0619] If nnpfc_separate_colour_description_present_flag is 1, it indicates that the unique combination of color primaries, transfer characteristics, matrix coefficients, scaling, and offset values applied with respect to the matrix coefficients for the pictures generated by NNPF is specified in the SEI message syntax structure. If nnpfc_separate_colour_description_present_flag is 0, it indicates that the combination of color primaries, transfer characteristics, matrix coefficients, scaling, and offset values applied with respect to the matrix coefficients for the pictures generated by NNPF is the same as that specified in the VUI parameter of CLVS.
[0620] nnpfc_colour_primaries has the same meaning as the vui_colour_primaries syntax element, but with the following differences:
[0621] nnpfc_colour_primaries represents the color primaries of the picture generated by applying the NNPF specified in the SEI message, rather than the color primaries used in CLVS.
[0622] If nnpfc_colour_primaries is not present in the NNPFC SEI message, the value of nnpfc_colour_primaries is inferred to be equal to vui_colour_primaries.
[0623] nnpfc_transfer_characteristics has the same meaning as specified for the vui_transfer_characteristics syntax element, except that:
[0624] nnpfc_transfer_characteristics represents the transfer characteristics of the picture generated by applying the NNPF specified in the SEI message, not the transfer characteristics used in CLVS.
[0625] If nnpfc_transfer_characteristics is not present in the NNPFC SEI message, the value of nnpfc_transfer_characteristics is inferred to be equal to vui_transfer_characteristics.
[0626] nnpfc_matrix_coeffs describes the equations used to derive the luma and chroma signals from the green, blue, red or Y, Z, X primaries. The semantics of this function are applied to the picture generated by applying the NNPF specified in this SEI message, as specified in MatrixCoefficients in Rec. ITU-T H.273 | ISO / IEC 23091-2, where BitDepthY and BitDepthC are equal to outTensorBitDepthY and outTensorBitDepthC, respectively.
[0627] If nnpfc_matrix_coeffs is not present in the NNPFC SEI message, the value of nnpfc_matrix_coeffs is inferred to be equal to vui_matrix_coeffs.
[0628] nnpfc_matrix_coeffs cannot be 0 unless both of the following conditions are true:
[0629] nnpfc_out_tensor_chroma_bitdepth_minus8 is the same as nnpfc_out_tensor_luma_bitdepth_minus8.
[0630] nnpfc_out_order_idc is 2, outSubHeightC is 1, and outSubWidthC is 1.
[0631] nnpfc_matrix_coeffs cannot be 8 unless one of the following conditions is true:
[0632] nnpfc_out_tensor_chroma_bitdepth_minus8 is the same as nnpfc_out_tensor_luma_bitdepth_minus8.
[0633] nnpfc_out_tensor_chroma_bitdepth_minus8 is equal to nnpfc_out_tensor_luma_bitdepth_minus8 + 1, nnpfc_out_order_idc is equal to 2, outSubHeightC is equal to 1, and outSubWidthC is equal to 1.
[0634] nnpfc_full_range_flag indicates the scaling and offset values applied with respect to the matrix coefficients specified in nnpfc_matrix_coeffs. The meaning of this value is the same as that specified in the VideoFullRangeFlag parameter of Rec. ITU-T H.273 | ISO / IEC 23091-2. If the nnpfc_full_range_flag value is absent, it is inferred to be 0.
[0635] If the value of nnpfc_chroma_loc_info_present_flag is 1, it indicates that the NNPFC SEI message contains the nnpfc_chroma_sample_loc_type_frame syntax element. If the value of nnpfc_chroma_loc_info_present_flag is 0, it indicates that the NNPFC SEI message does not contain the nnpfc_chroma_sample_loc_type_frame syntax element. If colorizationFlag is 0 or nnpfc_out_colour_format_idc is not 1, the value of nnpfc_chroma_loc_info_present_flag is equal to 0.
[0636] If nnpfc_chroma_sample_loc_type_frame is not 6 and nnpfc_out_colour_format_idc is 1, it indicates the chroma sample location of the output picture. If nnpfc_chroma_sample_loc_type_frame is 6 and nnpfc_out_colour_format_idc is 1, it indicates that the chroma sample location is unknown, unspecified, or specified in some other way not specified in this document. The value of nnpfc_chroma_sample_loc_type_frame is in the range 0 to 6, inclusive.
[0637] nnpfc_overlap indicates the number of horizontal and vertical samples that overlap between adjacent input tensors of NNPF. The nnpfc_overlap value ranges from 0 to 16,383.
[0638] If nnpfc_constant_patch_size_flag is 1, it indicates that NNPF accepts as input exactly the patch sizes specified by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. If nnpfc_constant_patch_size_flag is 0, NNPF accepts as input any patch size with width inpPatchWidth and height inpPatchHeight, where the width of the extended patch (i.e., the area overlapping the patch) is equal to inpPatchWidth + 2 * nnpfc_overlap, which is a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 * nnpfc_overlap, and the height of the extended patch is equal to inpPatchHeight + 2 * nnpfc_overlap, which is a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 * nnpfc_overlap.
[0639] The value of nnpfc_patch_width_minus1 plus 1 represents the horizontal sample of the patch size required for NNPF input when nnpfc_constant_patch_size_flag is 1. The value of nnpfc_patch_width_minus1 ranges from 0 to Min(32,766, CroppedWidth - 1).
[0640] The value of nnpfc_patch_height_minus1 plus 1 represents the number of vertical samples of the patch size required for NNPF input when nnpfc_constant_patch_size_flag is 1. The value of nnpfc_patch_height_minus1 ranges from 0 to Min(32,766, CroppedHeight - 1).
[0641] nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 * nnpfc_overlap represents the common divisor of all allowed values of extended patch widths required for NNPF input when nnpfc_constant_patch_size_flag is 0. The value of nnpfc_extended_patch_width_cd_delta_minus1 ranges from 0 to Min(32,766, CroppedWidth - 1), inclusive.
[0642] nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 * nnpfc_overlap represents the common divisor of all allowed values of extended patch heights required for NNPF input when nnpfc_constant_patch_size_flag is 0. The value of nnpfc_extended_patch_height_cd_delta_minus1 ranges from 0 to Min(32,766, CroppedHeight - 1).
[0643] Set the inpPatchWidth and inpPatchHeight variables to the patch size width and patch size height, respectively.
[0644] If nnpfc_constant_patch_size_flag is 0, the following applies:
[0645] The inpPatchWidth and inpPatchHeight values are provided through external means not specified in this document or are set by the postprocessor itself.
[0646] The value of inpPatchWidth + 2 * nnpfc_overlap must be a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 * nnpfc_overlap, where inpPatchWidth is less than or equal to CroppedWidth. The value of inpPatchHeight + 2 * nnpfc_overlap must be a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 * nnpfc_overlap, where inpPatchHeight is less than or equal to CroppedHeight.
[0647] Otherwise (nnpfc_constant_patch_size_flag is 1), the inpPatchWidth value is set to nnpfc_patch_width_minus1 + 1, and the inpPatchHeight value is set to nnpfc_patch_height_minus1 + 1.
[0648] The outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, and outPatchCHeight variables are derived as follows.
[0649] outPatchWidth = (nnpfcOutputPicWidth * inpPatchWidth) / CroppedWidth
[0650] outPatchHeight = (nnpfcOutputPicHeight * inpPatchHeight) / CroppedHeight
[0651] horCScaling = SubWidthC / outSubWidthC
[0652] verCScaling = SubHeightC / outSubHeightC
[0653] outPatchCWidth = outPatchWidth * horCScaling
[0654] outPatchCHeight = outPatchHeight * verCScaling
[0655] For bitstream conformance, outPatchWidth * CroppedWidth is equal to nnpfcOutputPicWidth * inpPatchWidth, and outPatchHeight * CroppedHeight is equal to nnpfcOutputPicHeight * inpPatchHeight.
[0656] nnpfc_padding_type indicates the padding process when referencing sample positions outside the input picture boundaries, as described in Table 9 (Informative description of nnpfc_padding_type values). The values of nnpfc_padding_type are in the range 0 to 4 inclusive in bitstreams conforming to this revision of this document. The values of nnpfc_padding_type in the range 5 to 15 are reserved for future use in ITU-T | ISO / IEC and are not present in bitstreams conforming to this revision of this document. Decoders conforming to this revision of this document will ignore NNPFC SEI messages with nnpfc_padding_type values between 5 and 15, inclusive. nnpfc_padding_type values greater than 15 are not present in bitstreams conforming to this revision of this document and are not reserved for future use.
[0657] [Table 9]
[0658]
[0659] nnpfc_luma_padding_val represents the luma value to be used for padding when nnpfc_padding_type is 4. The value of nnpfc_luma_padding_val ranges from 0 to (1 << BitDepthY) - 1.
[0660] nnpfc_cb_padding_val represents the Cb value to be used for padding when nnpfc_padding_type is 4. The value of nnpfc_cb_padding_val ranges from 0 to (1 << BitDepthC) - 1.
[0661] nnpfc_cr_padding_val represents the Cr value to be used for padding when nnpfc_padding_type is 4. The nnpfc_cr_padding_val value is in the range of 0 to (1 << BitDepthC) - 1.
[0662] The InpSampleVal(y, x, picHeight, picWidth, croppedPic, cIdx) function takes as input the vertical sample position y, the horizontal sample position x, the picture height picHeight, the picture width picWidth, the sample array croppedPic, and the component index cIdx (0 for luma, 1 for Cb, 2 for Cr) and returns the derived sampleVal value as follows:
[0663] For the input to the InpSampleVal() function, the vertical position is listed before the horizontal position for compatibility with the input tensor rules of some inference engines.
[0664] if(nnpfc_padding_type = = 0)
[0665] if( y < 0 | | x < 0 | | y >= picHeight | | x >= picWidth )
[0666] sampleVal = 0
[0667] else
[0668] sampleVal = croppedPic[x][y] (98)
[0669] else if(nnpfc_padding_type = = 1)
[0670] sampleVal = croppedPic[ Clip3( 0, picWidth - 1, x ) ][ Clip3( 0, picHeight - 1, y ) ]
[0671] else if(nnpfc_padding_type = = 2)
[0672] sampleVal = croppedPic[ Reflect( picWidth - 1, x ) ][ Reflect( picHeight - 1, y ) ]
[0673] else if(nnpfc_padding_type = = 3)
[0674] if( y >= 0 && y < picHeight )
[0675] sampleVal = croppedPic[ Wrap( picWidth - 1, x ) ][ y ]
[0676] else if(nnpfc_padding_type = = 4)
[0677] if( y < 0 | | x < 0 | | y >= picHeight | | x >= picWidth )
[0678] sampleVal = ( cIdx = = 0 ? nnpfc_luma_padding_val :
[0679] ( cIdx = = 1 ? nnpfc_cb_padding_val : nnpfc_cr_padding_val ) )
[0680] else
[0681] sampleVal = croppedPic[x][y]
[0682] NNPF PostProcessingFilter() is a target NNPF derived from the semantics of the NNPFA SEI message. The following example process can be used with NNPF PostProcessingFilter() to generate a filtered and / or interpolated image in a patch-wise manner. This image contains arrays of Y, Cb, and Cr samples named FilteredYPic, FilteredCbPic, and FilteredCrPic, respectively, indicated in nnpfc_out_order_idc.
[0683] if( nnpfc_inp_order_idc = = 0 | | nnpfc_inp_order_idc = = 2 )
[0684] for( cTop = 0; cTop < CroppedHeight; cTop += inpPatchHeight )
[0685] for( cLeft = 0; cLeft < CroppedWidth; cLeft += inpPatchWidth ) {
[0686] DeriveInputTensors( )
[0687] outputTensor = PostProcessingFilter( inputTensor )
[0688] StoreOutputTensors( )
[0689] }
[0690] else if( nnpfc_inp_order_idc = = 1 )
[0691] for( cTop = 0; cTop < CroppedHeight / SubHeightC; cTop += inpPatchHeight )
[0692] for( cLeft = 0; cLeft < CroppedWidth / SubWidthC; cLeft += inpPatchWidth ) {
[0693] DeriveInputTensors( )
[0694] outputTensor = PostProcessingFilter( inputTensor )
[0695] StoreOutputTensors( )
[0696] }
[0697] else if( nnpfc_inp_order_idc = = 3 )
[0698] for( cTop = 0; cTop < CroppedHeight; cTop += inpPatchHeight * 2 )
[0699] for( cLeft = 0; cLeft < CroppedWidth; cLeft += inpPatchWidth * 2 ) {
[0700] DeriveInputTensors( )
[0701] outputTensor = PostProcessingFilter( inputTensor )
[0702] StoreOutputTensors( )
[0703] }
[0704] The NNPF-generated image with index i, if any, contains sample arrays FilteredYPic[ i ], FilteredCbPic[ i ], and FilteredCrPic[ i ] derived by the above equation (cLeft = 0; cLeft < CroppedWidth / SubWidthC; cLeft += inpPatchWidth). The NNPF-generated image does not contain overlapping regions.
[0705] The NNPF process consists of the process defined in the above formula (cLeft = 0; cLeft < CroppedWidth / SubWidthC; cLeft += inpPatchWidth) and then the process that outputs the NNPF-generated images in index order. Here, all NNPF-generated images interpolated by NNPF are output, and the NNPF-generated images corresponding to the input images for NNPF are output as specified in the semantics of the NNPFA SEI message.
[0706] If nnpfc_complexity_info_present_flag is 1, it indicates the presence of one or more syntax elements representing the complexity of the NNPF associated with nnpfc_id. If nnpfc_complexity_info_present_flag is 0, it specifies that no syntax elements representing the complexity of the NNPF associated with nnpfc_id exist.
[0707] If nnpfc_parameter_type_idc is 0, it indicates that the network uses only integer parameters. If nnpfc_parameter_type_flag is 1, it indicates that the network can use either floating-point or integer parameters. If nnpfc_parameter_type_idc is 2, it indicates that the network uses only binary parameters. If nnpfc_parameter_type_idc is 3, it is reserved for future use in ITU-T | ISO / IEC and shall not be included in bitstreams conforming to this version of this document. Decoders conforming to this version of this document ignore NNPFC SEI messages with nnpfc_parameter_type_idc equal to 3.
[0708] If nnpfc_log2_parameter_bit_length_minus3 is 0, 1, 2, or 3, it indicates that the network does not use parameters with bit lengths greater than 8, 16, 32, or 64, respectively. If nnpfc_parameter_type_idc is present and nnpfc_log2_parameter_bit_length_minus3 is absent, the network does not use parameters with bit lengths greater than 1.
[0709] nnpfc_num_parameters_idc represents the maximum number of neural network parameters for NNPF, in powers of 2,048. A value of nnpfc_num_parameters_idc of 0 indicates that the maximum number of neural network parameters is unknown. The value of nnpfc_num_parameters_idc is in the range 0 to 52, inclusive. nnpfc_num_parameters_idc values greater than 52 are reserved for future use in ITU-T | ISO / IEC and should not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document ignore NNPFC SEI messages with nnpfc_num_parameters_idc greater than 52.
[0710] If the value of nnpfc_num_parameters_idc is greater than 0, the variable maxNumParameters is derived as follows.
[0711] maxNumParameters = ( 2 048 << nnpfc_num_parameters_idc ) - 1
[0712] It is a requirement of bitstream conformance that the number of neural network parameters in NNPF must be less than or equal to maxNumParameters.
[0713] If nnpfc_num_kmac_operations_idc is greater than 0, it indicates that the maximum number of multiply-accumulate operations per NNPF sample is less than or equal to nnpfc_num_kmac_operations_idc * 1 000. If nnpfc_num_kmac_operations_idc is 0, it indicates that the maximum number of multiply-accumulate operations in the network is unknown. The value of nnpfc_num_kmac_operations_idc ranges from 0 to 232-2(2).
[0714] If nnpfc_total_kilobyte_size is greater than 0, it indicates the total size (in kilobytes) required to store the uncompressed parameters of the neural network. The total size (in bits) is a number greater than or equal to the sum of the bits used to store each parameter. nnpfc_total_kilobyte_size is the total size (in bits) divided by 8,000 and rounded up. If nnpfc_total_kilobyte_size is 0, it indicates that the total size required to store the parameters of the neural network is unknown. The value of nnpfc_total_kilobyte_size ranges from 0 to 232-2(2).
[0715] If nnpfc_metadata_extension_num_bits is 0, it indicates the absence of nnpfc_reserved_metadata_extension. If nnpfc_metadata_extension_num_bits is greater than 0, it indicates the length (in bits) of nnpfc_reserved_metadata_extension. In this version of this document, nnpfc_metadata_extension_num_bits is 0. Values in the range 1 to 2,048 (inclusive) for nnpfc_metadata_extension_num_bits are reserved for future use in ITU-T | ISO / IEC and may not be present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document accept all nnpfc_metadata_extension_num_bits values in the range 0 to 2,048 (inclusive). If the nnpfc_metadata_extension_num_bits value is greater than 2,048, it is not present in bitstreams that follow this version of this document and is not reserved for future use.
[0716] The nnpfc_reserved_metadata_extension value is not present in bitstreams that conform to this version of this document. However, decoders that conform to this version of this document ignore the presence and value of nnpfc_reserved_metadata_extension. If nnpfc_reserved_metadata_extension is present, its length (in bits) is equal to nnpfc_metadata_extension_num_bits.
[0717] nnpfc_reserved_zero_bit_b is equal to 0 in bitstreams that follow this version of this document. The decoder ignores NNPFC SEI messages with nnpfc_reserved_zero_bit_b not equal to 0.
[0718] nnpfc_payload_byte[ i ] contains the ith byte of a bitstream conforming to ISO / IEC 15938-17. For all current values of i, the byte sequence nnpfc_payload_byte[ i ] is a complete bitstream conforming to ISO / IEC 15938-17.
[0719] Figure 26 illustrates the syntax of a neural network post-filter activation SEI (Supplemental enhancement information) message according to embodiments.
[0720] Referring to Figure 26, the neural network post filter activation (NNPFA) SEI message semantics is described.
[0721] The NNPFA SEI message enables or disables post-processing filtering for a set of pictures using a target neural network post-processing filter (NNPF), identified by nnpfa_target_id and nnpfa_target_base_flag. For a given picture with an NNPF enabled, the target NNPF is derived as follows:
[0722] If nnpfa_target_base_flag is 1, the target NNPF is the base NNPF whose nnpfc_id and nnpfa_target_id are the same.
[0723] Otherwise (nnpfa_target_base_flag is 0), the target NNPF is the NNPF specified in the last NNPFC SEI message with nnpfc_id equal to nnpfa_target_id that precedes the first VCL NAL unit of the current picture in decoding order, and is not a repeat of the NNPFC SEI message containing the base NNPF.
[0724] There may be multiple NNPFA SEI messages for the same picture, for example, if NNPF is used for different purposes or for filtering different color components.
[0725] nnpfa_target_id indicates the target NNPF associated with the current picture and specified by one or more NNPFC SEI messages whose nnpfc_id is equal to nnpfa_target_id. The value of nnpfa_target_id is in the range 0 to 232-2.
[0726] An NNPFA SEI message with a specific value of nnpfa_target_id does not exist in the current PU unless one or both of the following conditions are met:
[0727] Currently, there is an NNPFC SEI message with nnpfc_id equal to a specific value of nnpfa_target_id in a PU preceding the current PU in decoding order within the CLVS.
[0728] Currently, the PU has an NNPFC SEI message with nnpfc_id equal to a specific value of nnpfa_target_id.
[0729] If a PU contains both an NNPFC SEI message with a specific value of nnpfc_id and an NNPFA SEI message with an nnpfa_target_id equal to a specific value of nnpfc_id, the NNPFC SEI message precedes the NNPFA SEI message in decoding order.
[0730] If nnpfa_cancel_flag is 1, it indicates that the persistence of the target NNPF established by a previous NNPFA SEI message with the same nnpfa_target_id as the current SEI message is canceled. That is, the target NNPF is no longer used unless it is activated by another NNPFA SEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag is 0. If nnpfa_cancel_flag is 0, it indicates that nnpfa_target_base_flag, nnpfa_persistence_flag, and nnpfa_num_output_entries follow.
[0731] If nnpfa_target_base_flag is 1, it indicates that the target NNPF is a base NNPF whose nnpfc_id is equal to nnpfa_target_id. If nnpfa_target_base_flag is 0, it indicates that the target NNPF is an NNPF that is not a repeat of an NNPFC SEI message containing the base NNPF and whose nnpfc_id is equal to nnpfa_target_id in the last NNPFC SEI message preceding the first VCL NAL unit of the current picture in decoding order.
[0732] nnpfa_persistence_flag indicates the persistence of the target NNPF for the current layer.
[0733] If nnpfa_persistence_flag is 0, it indicates that the target NNPF can only be used for post-processing filtering on the current picture.
[0734] If nnpfa_persistence_flag is 1, it indicates that the target NNPF can be used for post-processing filtering for the current picture and all subsequent pictures in the current layer until one or more of the following conditions are met:
[0735] A new CLVS for the current layer begins. The bitstream ends. The picture of the current layer associated with the NNPFA SEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag set to 1 is output after the current picture in output order.
[0736] The target NNPF does not apply to subsequent pictures of the current layer associated with an NNPFA SEI message that has the same nnpfa_target_id as the current SEI message and has nnpfa_cancel_flag set to 1.
[0737] Let nnpfcTargetPictures be the set of pictures associated with the last NNPFC SEI message with nnpfc_id equal to nnpfa_target_id that precedes the current NNPFA SEI message in decoding order. Let nnpfaTargetPictures be the set of pictures whose target NNPFs are activated by the current NNPFA SEI message. For bitstream conformance, all pictures included in nnpfaTargetPictures must also be included in nnpfcTargetPictures.
[0738] nnpfa_num_output_entries represents the number of nnpfa_output_flag[ i ] syntax elements in the NNPFA SEI message. The value of nnpfa_num_output_entries ranges from 0 to NumInpPicsInOutputTensor, inclusive.
[0739] If nnpfa_output_flag[ i ] is 1, it indicates that the NNPF-generated picture corresponding to the input picture with index InpIdx[ i ] is output by the NNPF process activated by this NNPFA SEI message, where the NNPF process is specified in the semantics of the NNPFC SEI message. If nnpfa_output_flag[ i ] is 0, it indicates that the NNPF-generated picture corresponding to the input picture with index InpIdx[ i ] is not output by the NNPF process activated by this NNPFA SEI message. If nnpfa_num_output_entries is less than NumInpPicsInOutputTensor, nnpfa_output_flag[ i ] is inferred to be equal to 1 for each value of i in the range from nnpfa_num_output_entries to NumInpPicsInOutputTensor - 1, inclusive.
[0740] Figure 27 illustrates the syntax of a neural network post-filler group characteristic SEI message according to embodiments.
[0741] Referring to Figure 27, the semantics of the Neural-network post-filter group characteristics (NNPFGC) SEI message are explained.
[0742] The NNPFGC SEI message represents a neural network post-filter (NNPF) group. This SEI message indicates whether an NNPF group defines an NNPF cascade or whether it defines an NNPF group or an NNPF cascade that replaces each other. The use of an NNPF group in an NNPF cascade for a particular picture is indicated by the Neural Network Post-filter Group Activation (NNPFGA) SEI message.
[0743] nnpfgc_id contains an identification number that can be used to identify an NNPF group. nnpfgc_id values are in the range 0 to 2-2, inclusive. nnpfgc_id values 256 to 511, inclusive, and 231 to 2-2, inclusive, are reserved for future use by ITU-T | ISO / IEC. A decoder conforming to this version of this document shall ignore an NNPFGC SEI message with an nnpfgc_id in the range 256 to 511, inclusive, or 231 to 2-2, inclusive. An nnpfgc_id value must not be identical to the nnpfc_id value of an NNPFC SEI message in the same CLVS. If the nnpfgc_id value of NNPFGC SEI message nnpfgcSeiA is the same as the nnpfgc_id value of another NNPFGC SEI message nnpfgcSeiB in the same CLVS, nnpfgcSeiA and nnpfgcSeiB are equal.
[0744] If nnpfgc_grouping_type is 0, this SEI message specifies a group of cascaded neural network postprocessing filters.
[0745] If nnpfgc_grouping_type is 1, it indicates that the NNPF or NNPF group identified by nnpfgc_member_id[ i ] replaces each other, and the postprocessor should select only one of them to apply.
[0746] If nnpfgc_grouping_type is 2, this SEI message specifies a group of NNPFs that are intended to be used jointly, and are activated in rotation, indicating that at most one of these NNPFs is active for any picture.
[0747] If nnpfgc_grouping_type is 3, it indicates that the NNPF or NNPF group identified by nnpfgc_member_id[ i ] is intended to be used in parallel.
[0748] When nnpfgc_grouping_type is 4, it indicates that the NNPF or NNPF group identified by nnpfgc_member_id[ i ] is optional, i.e., it may or may not be applied by the postprocessor.
[0749] The nnpfgc_grouping_type value is in the range 0 to 255. nnpfgc_grouping_type values in the range 5 to 255 are reserved for future specification in ITU-T | ISO / IEC and are not present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document ignore NNPFGC SEI messages with nnpfgc_grouping_type in the range 5 to 255.
[0750] nnpfgc_purpose has the same meaning as nnpfc_purpose, except that it specifies the meaning for the NNPF group defined in this SEI message, not the NNPF defined in the NNPFC SEI message.
[0751] nnpfgc_num_members_minus2 + 2 represents the number of NNPFs or NNPF groups within the NNPF group defined by this SEI message.
[0752] nnpfgc_member_id[ i ] represents the ith member of the NNPF group defined by this SEI message, as follows:
[0753] If there is an NNPF with nnpfc_id equal to nnpfgc_member_id[ i ] defined in CLVS, then the ith member of the NNPF group defined by this SEI message is the NNPF with nnpfc_id equal to nnpfgc_member_id[ i ].
[0754] - Otherwise (if there is no NNPF with nnpfc_id equal to nnpfgc_member_id[ i ] defined in CLVS), the ith member of the NNPF group defined by this SEI message is the NNPF group with nnpfgc_id equal to nnpfgc_member_id[ i ].
[0755] For bitstream conformance, if the nnpfgc_member_id[ i ] value refers to the nnpfgc_id value of an NNPFGC SEI message nnpfgcSei, then the nnpfgc_grouping_type of the NNPFGC SEI message nnpfgcSei is 0. For bitstream conformance, if nnpfgc_grouping_type is 0 or 2, then there is an NNPF with an nnpfc_id value equal to nnpfgc_member_id[ i ] defined in CLVS. For bitstream conformance, if nnpfgc_grouping_type is 1, 3, or 4, then there is an NNPF with an nnpfc_id value equal to nnpfgc_member_id[ i ], or an NNPF group with an nnpfgc_id value equal to nnpfgc_member_id[ i ] defined in CLVS.
[0756] When nnpfgc_grouping_type is 0, NNPFs with nnpfc_id value equal to nnpfgc_member_id[ i ] are activated by NNPFGA SEI messages with nnpfga_target_id equal to nnpfgc_id, and are thus cascaded in increasing order of i.
[0757] nnpfgc_complexity_info_present_flag, nnpfgc_parameter_type_idc, nnpfgc_log2_parameter_bit_length_minus3, nnpfgc_num_parameters_idc, nnpfgc_num_kmac_operations_idc, nnpfgc_total_kilobyte_size have the semantics of nnpfc_complexity_info_present_flag, nnpfc_parameter_type_idc, nnpfc_log2_parameter_bit_length_minus3, nnpfc_num_parameters_idc, nnpfc_num_kmac_operations_idc, nnpfc_total_kilobyte_size, respectively, but differ in that these semantics are specified for the NNPF defined in this SEI message, not for the NNPF defined in the NNPFC SEI message. If nnpfgc_grouping_type is 1, nnpfgc_complexity_info_present_flag is equal to 0.
[0758] Figure 28 illustrates the syntax of a neural network post-filter group activation SEI message according to embodiments.
[0759] Referring to Figure 28, the semantics of the Neural-network post-filter group activation (NNPFGA) SEI message are described.
[0760] The NNPFGA SEI message enables or disables post-processing filtering of a set of pictures using a target neural network post-processing filter group (NNPFG) among the NNPF groups identified by nnpfga_target_id. The nnpfgc_grouping_type of the identified NNPF group is equal to 0 (cascade) or 1 (alternative). When nnpfgc_grouping_type is 1, each member of the group has the same number of input pictures and NNPF output pictures. For a given picture with NNPFG enabled, the target NNPFG precedes the first VCL NAL unit of the current picture in decoding order, and an NNPF in the target NNPFG is defined by an NNPFC SEI message with an nnpfc_id equal to the nnpfgc_member_id[ i ] of the target NNPFG and that exists in the current picture unit or precedes the current picture in decoding order.
[0761] Using this SEI message requires defining the following variables:
[0762] - Input picture width and height in luma sample units. Here, they are represented as InitCroppedWidth[idx] and InitCroppedHeight[idx], respectively, and are the width and height of candidate input pictures with indices idx from 0 to numCandInputPics - 1 (inclusive) that can be used as input to NNPFG.
[0763] - The luma sample array InitCroppedYPic[idx] and the chroma sample arrays InitCroppedCbPic[idx] and InitCroppedCrPic[idx] (if any), and the width and height of the candidate input pictures whose indices idx are from 0 to numCandInputPics - 1 (inclusive), which can be used as input to NNPFG.
[0764] - BitDepthY, the bit depth of the luma sample array of the candidate input picture.
[0765] - BitDepthC, the bit depth of the chroma sample array (if any) of the candidate input picture.
[0766] - Chroma format indicator represented by ChromaFormatIdc
[0767] - When nnpfc_auxiliary_inp_idc is 1, the filtering strength control value array StrengthControlVal[idx] contains real numbers in the range 0 to 1 (inclusive) for candidate input pictures whose index idx is in the range 0 to numCandInputPics - 1.
[0768] The candidate input picture with index 0 corresponds to the picture for which NNPFG is activated by this NNPFGA SEI message. Candidate input pictures with index i in the range from 1 to numCandInputPics-1 (inclusive) are output in order before the candidate input picture with index i-1. Assume candInputPicList[0] is the list of candidate input pictures output in reverse order.
[0769] nnpfga_target_id is specified by the NNPFGC SEI message and is associated with the current picture, and represents the target NNPFG whose nnpfgc_id is equal to nnpfga_target_id.
[0770] The value of nnpfga_target_id is in the range 0 to 232-2 (inclusive).
[0771] An NNPFGA SEI message with a particular nnpfga_target_id value is not present in the current PU unless an NNPFGC SEI message with an nnpfgc_id equal to a particular value of nnpfgc_target_id and an nnpfgc_grouping_type of 0 exists in the current PU or in a PU preceding the current PU in decoding order within the current CLVS.
[0772] If a PU contains both an NNPFGC SEI message with a specific nnpfgc_id value and an NNPFGA SEI message with an nnpfga_target_id equal to a specific value of nnpfgc_id, the NNPFGC SEI message precedes the NNPFGA SEI message in decoding order.
[0773] If nnpfga_cancel_flag is 1, it indicates that the persistence of the target NNPFG established by a previous NNPFGA SEI message with the same nnpfga_target_id as the current SEI message is canceled. That is, the target NNPFG is no longer used unless it is activated by another NNPFGA SEI message with the same nnpfga_target_id as the current SEI message and nnpfga_cancel_flag is 0. If nnpfga_cancel_flag is 0, it indicates that the target NNPFG is activated and available for use.
[0774] nnpfga_persistence_flag indicates the persistence of the target NNPFG for the current layer.
[0775] If nnpfga_persistence_flag is 0, it indicates that the target NNPFG can only be used for post-processing filtering on the current picture.
[0776] If nnpfga_persistence_flag is 1, it indicates that the target NNPFG can be used for post-processing filtering for the current picture and all subsequent pictures of the current layer in output order until one or more of the following conditions are true:
[0777] - A new CLVS for the current tier begins.
[0778] - The bitstream ends.
[0779] - The picture of the current layer associated with the NNPFGA SEI message that follows the current picture in output order and has the same nnpfga_target_id as the current SEI message.
[0780] NOTE - The target NNPFG does not apply to subsequent pictures of the current layer that are associated with an NNPFGA SEI message with the same nnpfga_target_id as the current SEI message.
[0781] Let nnpfgcTargetPictures be the set of pictures to which the last NNPFGC SEI message with nnpfgc_id equal to nnpfga_target_id, which comes before the current NNPFGA SEI message in decoding order, belongs. Let nnpfgaTargetPictures be the set of pictures whose target NNPFG is activated by the current NNPFGA SEI message. It is a requirement of bitstream conformance that all pictures included in nnpfgaTargetPictures must also be included in nnpfgcTargetPictures.
[0782] nnpfga_num_filters_minus2 + 2 represents the number of NNPFs in the NNPFG that this SEI message activates. The value of nnpfga_num_filters_minus2 is equal to the value of nnpfgc_num_members_minus2 in the NNPFGC SEI message where nnpfgc_id is equal to nnpfga_target_id.
[0783] If nnpfga_target_base_flag[ i ] is 1, it indicates that the ith NNPF in the target NNPFG is the base NNPF whose nnpfc_id is equal to nnpfgc_member_id[ i ] in an NNPFGC SEI message whose nnpfgc_id is equal to nnpfga_target_id. If nnpfga_target_base_flag[ i ] is 0, it indicates that the ith NNPF in the target NNPFG is the NNPF specified in the last NNPFC SEI message whose nnpfc_id is equal to nnpfgc_member_id[ i ] in an NNPFGC SEI message whose nnpfgc_id is equal to nnpfga_target_id, comes before the first VCL NAL unit of the current picture in decoding order, and is not a repeat of the NNPFC SEI message containing the base NNPF.
[0784] If nnpfga_input_all_pics_flag[ i ] is 1, it indicates that the input picture for the i-th NNPF is selected from the candidate input picture list candInputPicList[ i ] without skipping it. If nnpfga_input_all_pics_flag[ i ] is 0, the input picture for the i-th NNPF is selected from the candidate input picture list (candInputPicList[ i ]), and some candidate input pictures are skipped.
[0785] nnpfga_num_input_pics_minus1[ i ] represents the number of input pictures for the ith NNPF in the target NNPFG. If nnpfga_num_input_pics_minus1[ i ] is present, it is equal to nnpfc_num_input_pics_minus1 for the NNPF whose nnpfgc_id is equal to nnpfgc_member_id[ i ] of the NNPFGC SEI message equal to nnpfga_target_id. If not present, nnpfga_num_input_pics_minus1[ i ] is inferred to be equal to nnpfc_num_input_pics_minus1 for NNPFs whose nnpfc_id is equal to nnpfgc_member_id[ i ] in NNPFGC SEI messages where nnpfgc_id is equal to nnpfga_target_id.
[0786] nnpfga_input_pic_skip_count[ i ][ j ] represents the j-th picture count to skip in the candidate input picture list candInputPicList[ i ] when selecting an input picture for the NNPF activated by the ith loop item. If nnpfga_input_pic_skip_count[ i ][ j ] is absent, it is assumed to be 0 for all values of j in the range from 0 to nnpfga_num_input_pics_minus1[ i ]. The variable numCandInputPics, which represents the number of candidate input pictures for the NNPFG, is derived as follows.
[0787] numCandInputPics = 0
[0788] for( j = 0; j <= nnpfga_num_input_pics_minus1
[0000] ; j++ )
[0789] numCandInputPics += 1 + nnpfga_input_pic_skip_count
[0000] [ j ]
[0790] If candInputPicList[ m ] ranges from 1 to nnpfga_num_filters_minus2 + 1 (inclusive), it is initially empty and becomes a list of pictures in reverse output order formed in decreasing order of n in the range 0 to m-1 (inclusive). It contains each picture output from the NNPF process of the n-th loop item that does not already have a corresponding picture in candInputPicList[ m ], and finally each picture in candInputPicList
[0000] that does not already have a corresponding picture in candInputPicList[ m ]. If the candidate input picture candInputPicList[ m ][ idx ] is the NNPF output picture of the nth NNPF process for the value of m in the range from 1 to nnpfga_num_filters_minus2 + 1, and the value of n is less than the value of m, the width and height of the candidate input picture are equal to nnpfcOutputPicWidth and nnpfcOutputPicHeight of the NNPF output picture, respectively.
[0791] The list of input pictures inputPicList[ m ] for NNPF in the mth loop entry is derived as follows.
[0792] for( k = 0, candIdx = 0; k <= nnpfga_num_input_pics_minus1[ m ]; k++, candIdx++ ) {
[0793] candIdx += nnpfga_input_pic_skip_count[ m ][ k ]
[0794] inputPicList[ m ][ k ] = candInputPicList[ m ][ candIdx ]
[0795] }
[0796] For bitstream conformance, candIdx must not exceed the number of pictures in candInputPicList[ m ].
[0797] For bitstream conformance, the pictures in inputPicList[ m ] have the same width, height, bit depth, and chroma format for values of m in the range 1 to nnpfga_num_filters_minus2 + 1.
[0798] To interpret NNPFC SEI messages where nnpfc_id equals nnpfgc_member_id[ i ] in NNPFGC SEI messages where nnpfgc_id equals nnpfga_target_id, the following variables are specified in the ith loop entry:
[0799] - The BitDepthY, BitDepthC, and ChromaFormatIdc variables are used as provided for interpreting this SEI message.
[0800] - CroppedWidth and CroppedHeight are set to the same width and height, in luma samples, as the pictures in inputPicList[ i ].
[0801] - For each input picture k in the range 0 to nnpfga_num_input_pics_minus1[ i ], the following applies:
[0802] - If CroppedYPic[ k ], CroppedCbPic[ k ], and CroppedCrPic[ k ] exist, they are set equal to the corresponding sample arrays in inputPicList[ i ][ k ].
[0803] - If nnpfc_auxiliary_inp_idc is equal to 1 for an NNPF where nnpfc_id is equal to nnpfgc_member_id[ i ] in an NNPFGC SEI message where nnpfgc_id is equal to nnpfga_target_id, then the following applies:
[0804] - It is a bitstream conformance requirement that inputPicList[ i ][ k ] must be equal to candInputPicList
[0000] [ idx ] for all idx values in the range 0 to nnpfga_num_input_pics_minus1[ i ], inclusive. numCandInputPics - 1 (inclusive).
[0805] - StrengthControlVal[ k ] is set equal to InitStrengthControlVal[ idx ].
[0806] nnpfga_num_output_entries[ i ] represents the number of nnpfga_output_flag[ i ][ j ] syntax elements present in the NNPFGA SEI message. The value of nnpfga_num_output_entries[ i ] is in the range of 0 to NumInpPicsInOutputTensor for NNPFGC SEI messages where nnpfc_id is equal to nnpfgc_member_id[ i ] and nnpfgc_id is equal to nnpfga_target_id.
[0807] If nnpfga_output_flag[ i ][ j ] is 1, it indicates that the NNPF-generated picture corresponding to the input picture having the derived index InpIdx[ j ] for the ith NNPF of the target NNPFG is output by the NNPF process activated by this loop entry, where the NNPF process is specified in the semantics of the NNPFC SEI message. If nnpfga_output_flag[ i ][ j ] is 0, it indicates that the NNPF-generated picture corresponding to the input picture having the derived index InpIdx[ j ] for the ith NNPFG is not output by the NNPF process activated by this loop entry. If nnpfga_num_output_entries[ i ] is less than the NumInpPicsInOutputTensor derived for the ith NNPF of the target NNPFG, nnpfga_output_flag[ i ][ j ] is inferred to be 1 for each value of i in the range NumInpPicsInOutputTensor - 1 (inclusive) from nnpfga_num_output_entries[ i ].
[0808] NnpfgaOutputPicList, a list of pictures output in output order by the NNPF process of NNPFG, is initially empty and is formed in decreasing order of n in the range from 0 to nnpfga_num_filters_minus2 + 1, and contains each picture output by the NNPF process of the nth loop item for which there is no corresponding picture already in NnpfgaOutputPicList.
[0809] Figure 29 shows source picture timing information according to embodiments.
[0810] The Source Picture Timing Information (SPTI) SEI message indicates the temporal distance between the source picture associated with the corresponding decoded output picture before encoding. For example, for camera-captured content, the temporal distance between source pictures is the difference between the time the image sensor was exposed to produce the source picture associated with the currently decoded picture and the time the image sensor was exposed to produce the source picture associated with the previously decoded picture in output order. The information provided in the SPTI SEI message applies only to pictures in the current layer of the access unit containing the SPTI SEI message, and all subsequent pictures in the current layer in output order.
[0811] If spti_cancel_flag is 1, it indicates that the SPTI SEI message cancels the persistence of the previous SPTI SEI message in the output order applied to the current layer. If spti_cancel_flag is 0, it indicates that source picture timing information follows.
[0812] spti_persistence_flag indicates the persistence of SPTI SEI messages for the current layer.
[0813] If spti_persistence_flag is 0, it indicates that the SPTI SEI message applies only to the currently decoded picture.
[0814] If spti_persistence_flag is 1, it indicates that the SPTI SEI message applies only to the currently decoded picture and persists for all subsequent pictures in the current layer in output order until one or more of the following conditions are met: If spti_persistence_flag is 1, it indicates that it applies to multiple sublayers.
[0815] - A new CLVS for the current layer begins.
[0816] - The bitstream ends.
[0817] - The picture in the current layer of the AU associated with the SPTI SEI message is output after the current picture in output order.
[0818] If spti_source_timing_equals_output_timing_flag is 1, it indicates that the timing of the source picture is the same as the timing of the corresponding decoded output picture. If spti_source_timing_equals_output_timing_flag is 0, it indicates that the timing of the source picture may be different from the timing of the corresponding decoded output picture.
[0819] If spti_source_timing_equals_output_timing_flag is 1 and there is a picture timing SEI message for the current picture, the source picture timing can be determined from the information passed in the picture timing SEI message.
[0820] If spti_source_type_present_flag is 1, it indicates that the syntax element spti_source_type is present in the SEI message. If spti_source_type_present_flag is 0, it indicates that the syntax element spti_source_type is not present in the SEI message.
[0821] spti_source_type represents the timing relationship between the source picture and the corresponding decoded output picture specified in Table 10 (Interpretation of spti_source_type). Here, if (spti_source_type & bitMask) is not 0, it indicates that the timing relationship has an interpretation related to the bitMask value in Table 10 (Interpretation of spti_source_type). If spti_source_type is greater than 0 and (spti_source_type & bitMask) is 0, the interpretation related to the bitMask value does not apply to the SPTI SEI message. If spti_source_type is 0, the timing relationship can be specified by the application.
[0822] The value of spti_source_type is in the range 0 to 127 in bitstreams conforming to this version of this document. The values 128 to 255 (inclusive) for spti_source_type are reserved for future use in ITU-T | ISO / IEC and are not present in bitstreams conforming to this version of this document. Decoders conforming to this version of this document will ignore SPTI SEI messages with spti_source_type in the range 128 to 255.
[0823] [Table 10]
[0824]
[0825] The values of (spti_source_type & 0x04) and (spti_source_type & 0x08) are 0 (e.g., spti_source_type should not indicate both high-speed imaging and time-lapse imaging).
[0826] spti_time_scale represents the number of time units that elapse in one second. The value of spti_time_scale cannot be 0. For example, the spti_time_scale of a time coordinate system that measures time using a 27 MHz clock is 27,000,000.
[0827] spti_num_units_in_elemental_interval represents the number of time units of a clock operating at a frequency of spti_time_scale Hz corresponding to the specified element source picture interval of consecutive pictures in output order in CLVS. The value of spti_num_units_in_elemental_interval is non-zero.
[0828] The elemental source picture interval in seconds, represented by the ElementalSourcePictureInterval variable, is equal to the quotient of spti_num_units_in_elemental_interval divided by spti_time_scale. For example, to represent an element source picture interval of 0.04 seconds, spti_time_scale could be 27,000,000 and spti_num_units_in_elemental_interval could be 1,080,000.
[0829] The way to represent the elemental source picture interval is that in Rec. ITU-T H.266 | ISO / IEC 23090-3, spti_time_scale is similar to time_scale in that syntax, and spti_num_units_in_elemental_interval is similar to num_units_in_tick in that syntax, so the ElementalSourcePictureInterval variable is similar to the ClockTick variable in Rec. ITU-T H.266 | ISO / IEC 23090-3.
[0830] spti_max_sublayers_minus_1 + 1 indicates the maximum number of temporal sublayers for which the picture interval scale factor (spti_sublayer_interval_scale_factor[ i ]) and synthesis flag (spti_sublayer_synthesized_picture_flag[ i ]) information is signaled. If spti_max_sublayers_minus_1 is absent, it is inferred to be the same as TemporalId.
[0831] When spti_sublayer_interval_scale_factor[ i ] is present, it represents a scale factor used to determine the source picture interval of the corresponding picture in the CLVS with TemporalId i, relative to the previous output picture with TemporalId less than or equal to i. A value of 0 can be used to indicate that the source picture corresponding to the currently decoded output picture is the same as the source picture corresponding to the previous decoded output picture with TemporalId less than or equal to i.
[0832] The indicated source picture interval associated with an output picture with TemporalId i, in seconds compared to the previous output picture with TemporalId less than or equal to i, is denoted by the variable SourcePictureInterval[ i ] and is derived as follows:
[0833] SourcePictureInterval[ i ] = ElementalSourcePictureInterval * spti_sublayer_interval_scale_factor[ i ] *
[0834] ( 1 - 2 * temporalReversalFlag )
[0835] If spti_source_type_present_flag is 1, then the variable temporalReversalFlag is equal to ( spti_source_type & 0x10 )? 1 : 0. Otherwise (i.e., if spti_source_type_present_flag is 0), the variable temporalReversalFlag is equal to 0.
[0836] When computing SourcePictureInterval[ i ], ElementalSourcePictureInterval is multiplied by spti_sublayer_interval_scale_factor[ i ], so the same value of SourcePictureInterval[ i ] can be expressed in multiple ways by applying a scale factor to the spti_time_scale value and applying the same scale factor to spti_num_units_in_elemental_interval or spti_sublayer_interval_scale_factor[ i ]. There is no assumption that common scale factors have been removed or that the value of spti_sublayer_interval_scale_factor[ i ] is equal to 1 for the highest value of i. Allowing multiple ways to express the same value is at least in part to allow spti_time_scale to be chosen to match other timing-related factors used in the system environment, such as the 27 MHz clock rate used in some multimedia communication systems.
[0837] If spti_sublayer_synthesized_picture_flag[ i ] is present and set to 1, it indicates that the decoded output picture belonging to the i-th temporal sublayer is synthesized and does not match the original unmodified source picture. If spti_sublayer_synthesized_picture_flag[ i ] is set to 0, this indication is not provided. If absent, the value of spti_sublayer_synthesized_picture_flag[ i ] is inferred to be 0.
[0838] If the TemporalId of the SPTI SEI message is greater than 0 and the SPTI SEI message persists for one or more pictures with lower TemporalIds, the encoder can include the information of the SPTI SEI message in one or more SPTI SEI messages with lower TemporalIds to prevent information loss when pictures in the temporal lower layer are lost or removed.
[0839] Figures 30a and 30b illustrate object mask information SEI messages according to embodiments.
[0840] The Object Mask Information (OMI) SEI message provides object mask information for an object mask picture in an auxiliary picture layer associated with the base picture layer (now called the base picture layer) in which the SEI message exists. If an OMI SEI message exists, it is present in the base picture layer. A base picture layer may be associated with one or more auxiliary picture layers. For each associated auxiliary picture layer that contains an object mask picture whose nuh_layer_id is equal to sdi_layer_id[ i ], the value of sdi_aux_id[ i ] is equal to AUX_OBJECT_MASK for all values of i in the range 0 to sdi_max_layers_minus1, inclusive.
[0841] Using this SEI message requires defining the following variables:
[0842] - The cropped picture width and picture height in luma sample units are indicated as CroppedWidth and CroppedHeight, respectively.
[0843] - Conformity crop window left offset, ConfWinLeftOffset
[0844] - Top offset of the conformity crop window, ConfWinTopOffset
[0845] - Chroma format indicator, indicated by ChromaFormatIdc
[0846] The SubWidthC and SubHeightC variables are derived from ChromaFormatIdc.
[0847] If omi_cancel_flag is 1, it indicates that the SEI message cancels the persistence of the previous object mask information SEI message in the same layer in output order. If omi_cancel_flag is 0, it indicates that object mask information follows.
[0848] omi_persistence_flag indicates the persistence of the object mask information provided in this SEI message. If omi_persistence_flag is 0, the object mask information applies only to the current image. If omi_persistence_flag is 1, the object mask information applies to the current image and all subsequent images in the same layer in output order, until one or more of the following conditions are true:
[0849] - A new CLVS for the current layer begins.
[0850] - The bitstream ends.
[0851] - The picture in the current layer of the PU that includes the object mask information SEI message is output after the current picture in the output order.
[0852] If CVS does not contain an SDI SEI message with sdi_aux_id[ i ] equal to AUX_OBJECT_MASK for one or more of the i values, the OMI SEI message is ignored.
[0853] If an AU contains both an SDI SEI message and an OMI SEI message where sdi_aux_id[ i ] equals AUX_OBJECT_MASK for one or more of the i values, the SDI SEI message precedes the OMI SEI message in decoding order.
[0854] omi_num_aux_pic_layer represents the number of auxiliary picture layers associated with the current base picture layer. For bitstream conformance, the value of omi_num_aux_pic_layer must be equal to numAuxLayer, where the numAuxLayer variable is derived as follows.
[0855] omiPrimaryLayerId is represented as the nuh_layer_id value of the NAL unit containing the SEI message.
[0856] numAuxLayer = 0;
[0857] for( i = 0; i <= sdi_max_layers_minus1; i++ )
[0858] if( sdi_aux_id[ i ] = = AUX_OBJECT_MASK )
[0859] for( j = 0; j <= sdi_num_associated_primary_layers_minus1[ i ]; j++ )
[0860] if( sdi_layer_id[ sdi_associated_primary_layer_idx[ i ][ j ] ] = = omiPrimaryLayerId )
[0861] numAuxLayer++;
[0862] omi_mask_id_length_minus1 + 1 represents the length (in bits) of the omi_mask_id[ i ][ j ] syntax element.
[0863] omi_mask_sample_value_length_minus8 + 8 represents the length (in bits) of the omi_aux_sample_value[ i ][ j ] syntax element. The value of omi_mask_sample_value_length_minus8 is in the range 0 to 8.
[0864] If omi_mask_confidence_info_present_flag is 1, it indicates that the omi_mask_confidence[ i ][ j ] syntax element is present. If omi_mask_confidence_info_present_flag is 0, it indicates that the omi_mask_confidence[ i ][ j ] syntax element is not present.
[0865] omi_mask_confidence_length_minus1 + 1 represents the length (in bits) of the omi_mask_confidence[ i ][ j ] syntax element.
[0866] If omi_mask_depth_info_present_flag is 1, it indicates that the omi_mask_depth[ i ][ j ] syntax element exists. If omi_mask_depth_info_present_flag is 0, it indicates that the omi_mask_depth[ i ][ j ] syntax element does not exist.
[0867] omi_mask_depth_length_minus1 + 1 represents the length (in bits) of the omi_mask_depth[ i ][ j ] syntax element.
[0868] For bitstream conformance, the values of omi_num_aux_pic_layer, omi_mask_id_length_minus1, omi_mask_sample_value_length_minus8, omi_mask_confidence_info_present_flag, omi_mask_confidence_length_minus1 (if present), omi_mask_depth_info_present_flag, and omi_mask_depth_length_minus1 (if present) are the same across all object_mask_info( ) syntax constructs within a CLVS.
[0869] If omi_mask_label_info_present_flag is 1, it indicates that omi_mask_label_language_present_flag and omi_mask_label[ i ][ j ] syntax elements are present. If omi_mask_label_info_present_flag is 0, it indicates that omi_mask_label_language_present_flag and omi_mask_label[ i ][ j ] syntax elements are absent.
[0870] If omi_mask_label_language_present_flag is 1, it indicates that the omi_mask_label_language syntax element is present, and if omi_mask_label_language_present_flag is 0, it indicates that the omi_mask_label_language syntax element is not present.
[0871] omi_bit_equal_to_zero is equal to 0.
[0872] omi_mask_label_language contains the language tag specified in IETF RFC 5646 and a null-terminating byte such as 0x00. The length of the omi_mask_label_language syntax element is no more than 255 bytes, excluding the null-terminating byte. If this element is absent, the label's language is not specified.
[0873] If omi_mask_pic_update_flag[ i ] is 1, it indicates that the object mask information of the object mask picture in the ith auxiliary picture layer associated with the current base picture layer can be updated. If omi_mask_pic_update_flag[ i ] is 0, it indicates that the mask information of the object mask picture in the ith auxiliary picture layer associated with the current base picture layer is not changed. If omi_mask_pic_update_flag[ i ] is 0, the persistence mechanism is used. That is, the object mask information is inherited from the last OMI SEI message in the same layer in decoding order, which signals the mask information for the object mask picture in the ith auxiliary picture layer associated with the current base picture layer.
[0874] omi_num_mask_in_pic_update[ i ] represents the number of object masks in the object mask picture of the i-th auxiliary picture layer associated with the current base picture layer. omi_num_mask_in_pic_update[ i ] ranges from 0 to (1<<(omi_mask_id_length_minus1 + 1)) - 1.
[0875] omi_mask_id[ i ][ j ] represents the identifier of the jth object mask in the object mask picture of the ith auxiliary picture layer associated with the current base picture layer. The length of the omi_mask_id[ i ][ j ] syntax element is omi_mask_id_length_minus1 + 1 bits.
[0876] The variable maskId[ i ][ j ], which specifies the jth object mask identifier of the object mask picture of the ith auxiliary picture layer associated with the current base picture layer, is derived as follows.
[0877] for(i = 0; i < omi_num_aux_pic_layer; i++)
[0878] for(j = 0; j < omi_num_mask_in_pic_update[ i ]; j++)
[0879] maskId[ i ][ j ] = omi_mask_id[ i ][ j ] + (1<<(omi_mask_id_length_minus1 + 1))*i
[0880] omi_aux_sample_value[ i ][ j ] represents the sample value within the j-th object mask area of the object mask picture in the i-th auxiliary picture layer associated with the current base picture layer.
[0881] If omi_mask_cancel[ i ][ j ] is 1, it cancels the persistence range of the jth object mask of the object mask picture in the ith auxiliary picture layer associated with the current base picture layer. If omi_mask_cancel[ i ][ j ] is 0, it indicates that the jth object mask information of the object mask picture in the ith auxiliary picture layer associated with the current base picture layer is signaled.
[0882] As a requirement of bitstream conformance, when an omi_mask_id[ i ][ j ] with a particular value in the current CLVS is parsed for the first time, the value of the corresponding omi_mask_cancel[ i ][ j ] is equal to 0.
[0883] If omi_mask_bounding_box_present_flag[ i ][ j ] is equal to 1, it indicates that the syntax elements omi_mask_top[ i ][ j ], omi_mask_left[ i ][ j ], omi_mask_width[ i ][ j ], and omi_mask_height[ i ][ j ] exist. If omi_mask_bounding_box_present_flag[ i ][ j ] is equal to 0, it indicates that the syntax elements omi_mask_top[ i ][ j ], omi_mask_left[ i ][ j ], omi_mask_width[ i ][ j ], and omi_mask_height[ i ][ j ] do not exist.
[0884] omi_mask_top[ i ][ j ], omi_mask_left[ i ][ j ], omi_mask_width[ i ][ j ], omi_mask_height[ i ][ j ] represent the coordinates of the upper left corner and the width and height, respectively, of the bounding box of the j-th object mask in the cropped decoded object mask picture of the i-th auxiliary picture layer associated with the current base picture layer based on the suitability cropping window specified in the active SPS.
[0885] The values of omi_mask_left[ i ][ j ] must be in the range 0 to ( CroppedWidth / SubWidthC - 1 ), where CroppedWidth and SubWidthC are associated with the object mask photo of the ith auxiliary photo layer associated with the current main photo layer. If the value is missing, the value of omi_mask_left[ i ][ j ] is inferred to be 0.
[0886] The values of omi_mask_top[ i ][ j ] must be in the range 0 to ( CroppedHeight / SubHeightC - 1 ), where CroppedHeight and SubHeightC are associated with the object mask photo of the ith auxiliary photo layer associated with the current main photo layer. If a value is missing, the value of omi_mask_top[ i ][ j ] is inferred to be 0.
[0887] The value of omi_mask_width[ i ][ j ] ranges from 0 to (CroppedWidth / SubWidthC - omi_mask_left[ i ][ j ] ). If the value is missing, the value of omi_mask_width [ i ][ j ] is inferred to be (CroppedWidth / SubWidthC - omi_mask_left[ i ][ j ] ).
[0888] The value of omi_mask_height[ i ][ j ] ranges from 0 to (CroppedHeight / SubHeightC - omi_mask_top[ i ][ j ] ). If the value is missing, the value of omi_mask_height [ i ][ j ] is inferred to be (CroppedHeight / SubWidthC - omi_mask_top[ i ][ j ] ).
[0889] The identified object mask is within a bounding box that contains luma samples whose horizontal coordinates are from SubWidthC * ( ConfWinLeftOffset + omi_mask_left[ i ][ j ] ) to SubWidthC * ( ConfWinLeftOffset + omi_mask_left[ i ][ j ] + omi_mask_width[ i ][ j ] ) - 1 (inclusive), and whose vertical coordinates are from SubHeightC * ( ConfWinTopOffset + omi_mask_top[ i ][ j ] ) to SubHeightC * ( ConfWinTopOffset + omi_mask_top[ i ][ j ] + omi_mask_height[ i ][ j ] ) - 1 (inclusive).
[0890] The variable pI[ i ] [ x ][ y ] is the decoded value of the sample at the relative sample position (x, y) in the cropped object mask picture of the i-th auxiliary picture layer associated with the current base picture layer. The next process is to determine the mask area in the auxiliary picture.
[0891] for( i = 0; i < omi_num_aux_pic_layer; i++ )
[0892] for( j = 0; j < omi_num_mask_in_pic_update[ i ]; j++ )
[0893] if( pI[ i ][ x ][ y ] = = omi_aux_sample_value [ i ][ j ]
[0894] && x >= omi_mask_left[ i ][ j ]
[0895] && x < omi_mask_left[ i ][ j ] + omi_mask_width[ i ][ j ]
[0896] && y >= omi_mask_top[ i ][ j ]
[0897] && y < omi_mask_top[ i ][ j ] + omi_mask_height[ i ][ j ] )
[0898] the sample at location (x, y) is associated with the object mask with the identifier of maskId[ i ][ j ]
[0899] omi_mask_confidence[ i ][ j ] represents the confidence associated with the jth object mask of the object mask picture of the ith auxiliary picture layer associated with the current base picture layer, in units of 2 - ( omi_mask_confidence_length_minus1 + 1 ). Therefore, a higher value of omi_mask_confidence[ i ][ j ] indicates a higher confidence. The length of the omi_mask_confidence[ i ][ j ] syntax element is omi_mask_confidence_length_minus1 + 1 bits.
[0900] omi_mask_depth[ i ][ j ] represents the object depth associated with the jth object mask of the object mask picture of the ith auxiliary picture layer associated with the current base picture layer. A smaller omi_mask_depth value indicates a shorter distance to the object. The length of the omi_mask_depth[ i ][ j ] syntax element is omi_mask_depth_length_minus1 + 1 bit.
[0901] omi_mask_label[ i ][ j ] represents the contents of the label associated with the jth object mask of the object mask picture of the ith auxiliary picture layer associated with the current base picture layer. The length of the omi_mask_label[ i ][ j ] syntax element is 255 bytes or less, excluding the terminating null byte.
[0902] Figure 31 shows an SEI processing order SEI message according to embodiments.
[0903] SEI Processing Order (SPO) The SEI message carries information indicating the preferred processing order determined by the encoder (i.e., content producer) for the different types of SEI messages that may exist in the CVS.
[0904] The semantics of SPO SEI messages use the concept of SEI message types. SEI messages with different payloadType values are considered to be different types of SEI messages. Additionally, different SEI messages with the same payloadType value but distinguished by the values of the syntax elements in their SEI payload are considered to be different types of SEI messages. This distinction based on the values of the syntax elements in their SEI payload is made by comparing the values transmitted using the po_sei_prefix_data_bit[ i ][ j ] syntax elements (if present) or by comparing the values transmitted within the SEI message within a processing-order nested SEI message (if present). For example, neural network post-filter feature (NNPFC) SEI messages can be distinguished by having different nnpfc_id values.
[0905] If both po_sei_wrapping_flag[ i ] and po_sei_prefix_flag[ i ] of the ith SEI message seiA in all SPO SEI messages are 0, then no other SEI message seiB is included in the same SPO SEI message or in another SPO SEI message in the current CVS for which all of the following conditions are true:
[0906] - The po_sei_payload_type[ i ] value of seiB is the same as the value of seiA.
[0907] - The value of po_sei_wrapping_flag[ i ] of seiB is 0.
[0908] - The value of po_sei_prefix_flag[ i ] of seiB is 1.
[0909] If an SPO SEI message with a specific po_id value exists in an access unit of CVS, the SPO SEI message with that po_id value exists in the first access unit of CVS in decoding order. The number of SEI messages and the payloadType code indicated in each SPO SEI message with the same po_id value are maintained in decoding order from the current access unit to the end of CVS (in output order).
[0910] An SPO SEI message may contain one or more SEI prefix indications for a particular payloadType. Each SEI prefix indication is a bit string that follows the SEI payload syntax for the corresponding payloadType value, starting with the first syntax element of the SEI payload and containing multiple complete syntax elements. These SEI prefix indications provide sufficient information to determine the specific processing order for SEI message types with the same payloadType value but different preferred processing orders.
[0911] po_id contains an identification number that identifies the SPO SEI message.
[0912] A processing chain consists of a list of SEI message types identified by an SPO SEI message, in the preferred processing order specified in the SPO SEI message.
[0913] Each SEI message type in the processing chain represented by a SPO SEI message is identified by the syntax elements po_sei_payload_type[ i ], po_sei_wrapping_flag[ i ], po_sei_processing_order[ i ], and, if present, po_num_bits_in_prefix_indication_minus1[ i ] and po_prefix_data_bit[ i ][ j ].
[0914] An SEI message type does not need to belong to any processing chain, and may belong to multiple processing chains identified by SPO SEI messages with different po_id values.
[0915] Each SEI message of an SEI message type identified within a SPO SEI message has the same persistence scope as if that SEI message were transmitted outside of the SPO SEI message and not identified within the SPO SEI message.
[0916] Processing chains can be alternatives, meaning that at most one processing chain can be selected for application. Alternatively, they can be complementary, meaning that two or more processing chains can be selected and applied independently, each producing a single output.
[0917] po_num_sei_messages_minus2 + 2 represents the number of SEI message types for which preferred processing order is specified in SPO SEI messages.
[0918] If po_sei_wrapping_flag[ i ] is 1, it indicates that there must be at least one processing-order nested SEI message that satisfies both of the following constraints:
[0919] - pon_target_po_id[ j ] with all values of j is equal to po_id.
[0920] - There is a kth loop item in the processing order nested SEI message, so the payloadType of the kth nested SEI message is equal to po_sei_payload_type[ i ] and pon_processing_order[ k ] is equal to po_sei_processing_order[ i ].
[0921] If po_sei_wrapping_flag[ i ] is 0, then the SEI message whose payloadType is po_sei_payload_type[ i ] (and if po_sei_prefix_flag[ i ] is 1, then the prefix data matching the value of po_sei_prefix_data_bit[ i ][ j ]) is outside the processing order nested SEI message.
[0922] When po_sei_wrapping_flag[ i ] is 1, it allows SEI messages to be passed within a processing-order nested SEI message, preventing decoders that do not process SPO SEI messages from misinterpreting the SEI message. Therefore, when po_sei_wrapping_flag[ i ] is 0, unintended results may be generated in the decoder, so po_sei_wrapping_flag[ i ] is 1.
[0923] po_sei_importance_flag[ i ] indicates the importance determined by the encoder for the SEI message type with index i.
[0924] If the decoding system cannot interpret or does not support the functionality indicated in an SEI message with po_sei_importance_flag[ i ] set to 1, it ignores the entire SPO SEI message.
[0925] po_sei_payload_type[ i ] represents the payloadType value of the ith type of the SEI message.
[0926] If po_sei_prefix_flag[ i ] is 1, it indicates the presence of po_num_bits_in_prefix_indication_minus1[ i ] and some po_sei_prefix_data_bit[ i ][ j ] syntax elements. If po_sei_prefix_flag[ i ] is 0, it indicates the absence of these syntax elements.
[0927] SeiProcessingOrderSeiList is set to consist of payloadType values 3, 4, 5, 19, 137, 142, 144, 147, 148, 149, 165, 177, 210, 211. The value of po_sei_payload_type[ i ] for each i in the range 0 to po_num_sei_messages_minus2 + 1 is equal to the value in SeiProcessingOrderSeiList.
[0928] po_sei_processing_order[ i ] represents the preferred processing order for the i-th type of SEI message for which preferred processing order information is provided in the SPO SEI message. For two different integer values of m and n, a po_sei_processing_order[ m ] less than po_sei_processing_order[ n ] indicates that the SEI message type associated with index m should be processed before the SEI message type associated with index n, and a po_sei_processing_order[ m ] equal to po_sei_processing_order[ n ] indicates that there is no preferred processing order between the SEI message types associated with indices m and n (e.g., they may both represent different attributes applicable at that stage, or alternative processes applicable, or one may represent an attribute and the other a process).
[0929] If i is greater than 0, po_sei_processing_order[ i ] is greater than or equal to po_sei_processing_order[ i - 1 ].
[0930] po_num_bits_in_prefix_indication_minus1[ i ] and po_sei_prefix_data_bit[ i ][ j ], if present, have the same semantics as the num_bits_in_prefix_indication_minus1[ i ] and sei_prefix_data_bit[ i ][ j ] syntax elements of an SEI prefix indication SEI message, and prefix_sei_payload_type is replaced by po_sei_payload_type[ i ].
[0931] If there is more than one SPO SEI message with a particular po_id value in CVS, the value of po_num_sei_messages_minus2 and the values of po_sei_wrapping_flag[ i ], po_sei_prefix_flag[ i ], po_sei_importance_flag[ i ], po_sei_payload_type[ i ], po_sei_processing_order[ i ] for each value of i are the same as for other SPO SEI messages with the same po_id value in CVS.
[0932] po_byte_alignment_bit_equal_to_one is equal to 1.
[0933] Figure 32 illustrates a processing order nesting SEI message according to embodiments.
[0934] A Processing Order Overlap (PON) SEI message contains one or more SEI messages that must be applied only as part of the processing chain identified by the associated SEI Processing Order SEI message, and must not be applied in a manner that contradicts the processing chain identified by the associated SEI Processing Order SEI message.
[0935] An SEI message included in a PON SEI message is called a PON nested SEI message.
[0936] An encoder can include multiple PON SEI messages in the same access unit. For example, the first PON SEI message in an access unit can include a PON nested SEI message that applies to multiple processing chains, and one or more other PON SEI messages in the same access unit that apply only to a single processing chain.
[0937] As a requirement of bitstream conformance, the meaning and effect of non-PON nested SEI messages must not depend on the meaning or effect of PON nested SEI messages. This restriction has the following specific consequences, where associated SEI messages are considered SEI messages that affect the meaning or effect of a specific SEI message:
[0938] If there is a neural network post-filter feature SEI message with a particular value of nnpfc_id that is a PON nested SEI message, then the associated neural network post-filter activation SEI message with nnpfa_target_id equal to that nnpfc_id value is also a PON nested SEI message.
[0939] - If nnpfa_persistence_flag is 1 and there is a neural network post-filter activation (NNPFA) SEI message with a particular value of nnpfa_target_id that is not a PON overlap SEI message, the next picture in the output order with an NNPFA SEI message with the same nnpfa_target_id value (if any) in the same CLVS does not have an associated NNPFA SEI message that is a PON overlap SEI message.
[0940] - If fg_characteristics_persistence_flag is 1 and a film grain characteristics SEI message other than a PON overlap SEI message exists, there is no associated film grain characteristics SEI message in the same CLVS that is a PON overlap SEI message.
[0941] - If there is a frame packing array SEI message with fp_arrangement_persistence_flag equal to 1 and is not a PON nested SEI message, and there is no associated frame packing array SEI message with fp_arrangement_cancel_flag equal to 1 or with the same fp_arrangement_id value in the same CLVS, then this message is a PON nested SEI message.
[0942] - If a content color volume SEI message with ccv_persistence_flag set to 1 is not a PON nested SEI message, there is no associated frame packing array SEI message in the same CLVS that is a PON nested SEI message.
[0943]
[0944] If erp_persistence_flag is 1 and there is an isometric projection SEI message that is not a PON overlapping SEI message, then there is no associated isometric projection SEI message that is a PON overlapping SEI message in the same CLVS.
[0945] - If gcmp_persistence_flag is 1 and a generalized cubemap project SEI message exists that is not a PON nested SEI message, then there is no associated generalized cubemap project SEI message in the same CLVS that is a PON nested SEI message.
[0946] - If sphere_rotation_persistence_flag is 1 and there is a spherical rotation SEI message that is not a PON overlap SEI message, there is no associated spherical rotation SEI message that is a PON overlap SEI message in the same CLVS.
[0947] - If rwp_persistence_flag is 1 and there is a regional packing SEI message that is not a PON nested SEI message, there is no associated regional packing SEI message that is a PON nested SEI message in the same CLVS.
[0948] - If omni_viewport_persistence_flag is 1 and an omni-viewport SEI message other than a PON nested SEI message exists, there is no associated omni-viewport SEI message in the same CLVS that is a PON nested SEI message.
[0949] - If sari_persistence_flag is 1 and there is a sample aspect ratio SEI message that is not a PON overlap SEI message, there is no associated sample aspect ratio SEI message that is a PON overlap SEI message in the same CLVS.
[0950] - If there is an annotated local SEI message that is not a PON nested SEI message, there is no associated annotated local SEI message in the same CLVS that is a PON nested SEI message.
[0951] If an alpha channel information SEI message other than a PON overlap SEI message exists, then there is no associated alpha channel information SEI message in the same CLVS that is a PON overlap SEI message.
[0952] - If a display direction SEI message other than a PON overlap SEI message exists, there is no associated display direction SEI message that is a PON overlap SEI message in the same CLVS.
[0953] - If there is a color transformation indication SEI message with color_transform_persistence_flag equal to 1 that is not a PON nested SEI message, there is no associated color transformation indication SEI message with color_transform_cancel_flag equal to 1 or with the same color_transform_id value that is a PON nested SEI message in the same CLVS.
[0954] pon_num_po_ids_minus1 + 1 represents the number of SEI processing order SEI messages associated with this PON SEI message.
[0955] pon_target_po_id[ i ] represents the po_id of the SEI message of the i-th associated SEI processing order.
[0956] pon_num_seis_minus1 + 1 represents the number of PON nested SEI messages contained in this PON SEI message.
[0957] pon_processing_order[ i ] represents the position of the ith processing order nested SEI message within the processing order defined in the associated SEI processing order SEI message. If i is greater than 0, pon_processing_order[ i ] is greater than or equal to pon_processing_order[ i - 1 ].
[0958] Each associated SEI processing order SEI message must have at least one value i in the range 0 through pon_num_seis_minus1, inclusive, and there is an item k in the associated SEI processing order SEI message for which all of the following are true:
[0959] - po_sei_processing_order[ k ] is equal to pon_processing_order[ i ].
[0960] - po_sei_payload_type[ k ] is equal to the payloadType value of the ith PON nested SEI message.
[0961] - When po_sei_prefix_flag[ k ] is 1, po_sei_prefix_data_bit[ k ][ j ] for j in the range from 0 to po_num_bits_in_prefix_indication_minus1[ k ] contains the same content as po_num_bits_in_prefix_indication_minus1[ k ] of the SEI message payload of the ith PON nested SEI message plus one initial bit.
[0962] The ith PON nested SEI message is applied as the kth loop item of the associated SEI processing sequence SEI message.
[0963] Processing of the processing chain
[0964] Processing chains are interchangeable, meaning that the decoding system can select and apply at most one processing chain at a time.
[0965] The decoding system can select and apply the processing chain as follows:
[0966] - First, decode the bitstream, set the PoPicList list to a list of decoded images cut in output order resulting from the bitstream decoding, and select the processing chain.
[0967] - (Option 1: Picture-by-picture, zig-zag, breadth-first for one filter) For each SEI message type in the selected processing chain, the following are applied in non-decreasing order of their po_sei_processing_order[ i ] values:
[0968] - If the SEI message associated with the i-th SEI message type persists for a picture picA or for a picture whose NNPF generated picA was activated by a previous process in the processing chain, then the following applies to each picture picA in PoPicList in output order:
[0969] - If picA is not a truncated decoded picture, the following exception applies to the interpretation of SEI messages:
[0970] - Interface variables for interpreting SEI messages are derived from picA instead of syntax elements representing properties of the corresponding truncated decoded picture.
[0971] The meaning of an SEI message, or the meaning of an SEI message, and if the SEI message is an NNPFA SEI message, the meaning of the associated NNPFC SEI message, applies to the pictures in PoPicList instead of the truncated decoded pictures.
[0972] - If the i-th SEI message type is present in SpoProcessingList, the process included in the SEI message is performed, the corresponding picture is replaced with the corresponding processed picture (if any) generated as a result of the process, and another picture (if any) generated as a result of the process is inserted into PoPicList, thereby updating PoPicList by observing the output order.
[0973] (Option 2: Filter-by-filter, jag compression, depth-first for a single picture) The following is applied recursively for each picture picA in PoPicList in output order. If the set of SEI messages associated with the SEI message types in the SpoProcessingList of the selected processing chain persists for picA, the following is applied:
[0974] - The following are applied to each SEI message set in non-decreasing order of the corresponding po_sei_processing_order[ i ] values.
[0975] - If the current SEI message is not the first message in the SEI message set, the following exceptions apply to the interpretation of SEI messages:
[0976] - Interface variables for interpreting SEI messages are derived from the pictures in the updated PoPicList instead of the syntax elements representing the properties of the corresponding truncated decoded picture.
[0977] - The meaning of the SEI message, or the meaning of the SEI message, and if the SEI message is an NNPFA SEI message, the meaning of the associated NNPFC SEI message, applies to the pictures in the PoPicList instead of the decoded cropped pictures.
[0978] The process included in the SEI message is repeatedly called for each picture in picA and PoPicList in output order. These pictures are either interpolated or extrapolated pictures generated by applying the process included in the previous SEI message to picA or correspond to them. Each time the process is called, PoPicList is updated to respect the output order by replacing the picture with the corresponding processed picture (if any) generated as a result of that process and inserting another picture (if any) into PoPicList.
[0979] Figure 33 shows the syntax of an encoder optimization information SEI message according to embodiments.
[0980] The Encoder Optimization Information SEI message is used to indicate whether the video is optimized for human viewing or machine analysis, and what type of optimization was applied during preprocessing or encoding.
[0981] If eoi_cancel_flag is 1, it indicates that the persistence of the encoder optimization information SEI message included in the previous PU in the output order is canceled. If eoi_cancel_flag is 0, it indicates that the optimization information applied during preprocessing or encoding is applied next.
[0982] eoi_persistence_flag indicates the persistence of the optimization information provided in this SEI message. If eoi_persistence_flag is 0, the optimization information applies only to the current picture. If eoi_persistence_flag is 1, the optimization information applies to the current picture and all subsequent pictures in the current layer in output order until one or more of the following conditions are met:
[0983] - A new CLVS for the current tier begins.
[0984] - The bitstream ends.
[0985] - The picture of the current layer associated with the encoder optimization information SEI message is output after the current picture in the output order.
[0986] A value of eoi_for_human_viewing_idc of 3 indicates that the applied optimizations include human viewing. A value of eoi_for_human_viewing_idc of 2 indicates that the video is suitable but is not specifically optimized for human viewing. A value of eoi_for_huma_viewing_idc of 1 indicates that the video is not suitable for human viewing. A value of eoi_for_human_viewing_idc of 0 indicates that it is not known whether the video is suitable for human viewing.
[0987] If eoi_for_machine_analysis_idc is 3, it indicates that machine analysis is included in the purpose of the applied optimization. If eoi_for_machine_analysis_idc is 2, it indicates that the video is suitable but is not specifically optimized for machine analysis. If eoi_for_machine_analysis_idc is 1, it indicates that the video is not suitable for machine analysis. If eoi_for_machine_analysis_idc is 0, it is not known whether the video is suitable for machine analysis.
[0988] As a requirement of bitstream conformance, the values of eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc are not both 1. eoi_type represents the type of optimization method specified in Table x1. Here, if (eoi_type & bitMask) is not 0, it indicates that the optimization type using the bitMask value in Table 11 (Definition of eoi_type) is applied. If eoi_type is greater than 0 and (eoi_type & bitMask) is 0, the optimization type using the bitMask value is not applied. If eoi_type is 0, the optimization determined by the application is used.
[0989] [Table 11]
[0990]
[0991] EoiTemporalQualityFlag, EoiSpatialQualityFlag, and EoiPrivacyProtectionFlag specify whether eoi_type represents an optimization type, including object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, and privacy-preserving optimization, and are derived as follows:
[0992]
[0993] EoiObjectBasedFlag = ( ( eoi_type & 0x01 ) > 0 ) ? 1:0
[0994] EoiTemporalResamplingFlag = ( ( eoi_type & 0x02 ) > 0 ) ? 1:0
[0995] EoiSpatialResamplingFlag = ( ( eoi_type & 0x04 ) > 0 ) ? 1 : 0 (xx)
[0996] EoiTemporalQualityFlag = ( ( eoi_type & 0x08 ) > 0 ) ? 1:0
[0997] EoiSpatialQualityFlag = ( (eoi_type & 0x10 ) > 0 ) ? 1:0
[0998] EoiPrivacyProtectionFlag = ( (eoi_type & 0x20 ) > 0 ) ? 1:0
[0999] For example, if a particular top temporal sublayer is encoded with coarse quantization that makes quality fluctuations unpleasant for human viewers, but does not affect machine task performance, eoi_for_human_viewing_flag and eoi_for_machine_analaysis_flag can be set to 0 and 1, respectively, and eoi_type can be set to a value that sets EoiTemporalQualityFlag to 1.
[1000] When eoi_persistence_flag is 0, EoiTemporalResamplingFlag is 0 and EoiTemporalQualityFlag is 0 according to bitstream conformance requirements.
[1001] eoi_object_based_idc, if present, indicates the object-based optimization type specified in Table 12 (Definition of eoi_object_based_idc). Where (eoi_object_based_idc & bitMask) is non-zero, it indicates that the object-based optimization type associated with the bitMask value in Table 12 is applied. If eoi_object_based_idc is greater than 0 and (eoi_object_based_idc & bitMask) is 0, the object-based optimization type associated with the bitMask value is not applied. If eoi_object_based_idc is 0, the application-defined object-based optimization type is applied. The value of eoi_object_based_idc is in the range 0 to 7 in bitstreams conforming to this version of this specification. The values 8 to 65,535 for eoi_object_based_idc are reserved for future use by ITU-T. Defined by ISO / IEC and is not present in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification ignore eoi_object_based_idc if the value of eoi_object_based_idc is between 8 and 65,535 (inclusive).
[1002] [Table 12]
[1003]
[1004] If eoi_temporal_resampling_type_flag is 0, it indicates that the temporal resampling optimization is a subsampling operation. If eoi_temporal_resampling_type_flag is 1, it indicates that the temporal resampling optimization is an upsampling operation.
[1005] If eoi_num_int_pics is greater than 0, it indicates that the encoding system has a constant number of pictures to exclude between each pair of coded pictures in output order (if eoi_temporal_resampling_type_flag is 0) or to add for encoding between each pair of source pictures within the duration of this SEI message (if eoi_temporal_resampling_type_flag is 1). If eoi_temporal_resampling_type_flag is 0 and eoi_num_int_pics is greater than 0, eoi_num_int_pics indicates the number of pictures that the encoding system has excluded between each pair of coded pictures in output order. If eoi_temporal_resampling_type_flag is 1 and eoi_num_int_pics is greater than 0, eoi_num_int_pics indicates the number of pictures that the encoding system has added between each pair of source pictures for encoding.
[1006] If eoi_num_int_pics is 0, it indicates that the encoding system, within the persistence of this SEI message, does not know or is variable the number of pictures it has excluded between each pair of coded pictures in output order (if eoi_temporal_resampling_type_flag is 0) or added between each pair of source pictures for encoding (if eoi_temporal_resampling_type_flag is 1).
[1007] The eoi_num_int_pics value ranges from 0 to 63.
[1008] If eoi_spatial_resampling_type_flag is 0, it indicates that the spatial resampling optimization is a subsampling operation. If eoi_spatial_resampling_type_flag is 1, it indicates that the spatial resampling optimization is an upsampling operation.
[1009] If eoi_privacy_protection_type_idc exists, it indicates the privacy optimization type specified in Table 13 (Definition of eoi_privacy_protection_type_idc).
[1010] [Table 13]
[1011]
[1012] eoi_privacy_protected_info_type, if present, indicates the type of information being protected, as specified in Table 14(). Where eoi_privacy_protected_info_type is greater than 0 and (eoi_privacy_protected_info_type & bitMask) is non-zero, this indicates that an information type with a bitMask value from Table 14 (Definition of eoi_privacy_protection_info_type) is protected. If eoi_privacy_protected_info_type is 0, information of an application-defined type is protected. The values of eoi_privacy_protection_info_type are in the range 0 to 7 inclusive in a bitstream conforming to this version of this specification. The values 8 to 255 (inclusive) for eoi_privacy_protected_info_type are reserved for future use in ITU-T | ISO / IEC and are not present in a bitstream conforming to this version of this specification. Decoders conforming to this version of this specification ignore eoi_privacy_protected_info_type if the value of eoi_privacy_protected_info_type is in the range 8 to 255.
[1013] [Table 14]
[1014]
[1015] Figure 34 shows the syntax of a text description information SEI message according to embodiments.
[1016] The Text Description Information SEI message provides a text description of one or more pictures.
[1017] txt_descr_id represents the identifier value of this text description information SEI message. The txt_descr_id value must be in the range 1 to 16383. The value 0 is reserved.
[1018] If txt_cancel_flag is 1, it indicates that the text description information SEI message cancels the persistence of the previous text description information SEI message with the same txt_descr_id in the output order applied to the current layer. If txt_cancel_flag is 0, it indicates that the text description information follows.
[1019] txt_persistence_flag indicates the persistence of the text information description SEI message for the current layer.
[1020] If txt_persistence_flag is 0, it indicates that the text description information applies only to the currently decoded picture.
[1021] If txt_persistence_flag is 1, it indicates that the text description information SEI message is applied to the currently decoded picture and persists in output order for all subsequent pictures in the current layer until one or more of the following conditions are met:
[1022] - A new CLVS for the current tier begins.
[1023] - The bitstream ends.
[1024] - The picture in the current layer of the AU associated with the text description information SEI message with the same txt_descr_id is output after the current picture in output order.
[1025] txt_descr_purpose indicates the purpose of the text description SEI, as specified in Table 15 (Definition of txt_descr_purpose). Values for text_descr_purpose range from 0 to 5, inclusive. Values in the range 6 to 255 for text_descr_purpose are reserved for future use in ITU-T | ISO / IEC and are not present in bitstreams conforming to this version of this specification. Decoders conforming to this version of this specification accept any text_descr_purpose value in the range 0 to 255, inclusive.
[1026] [Table 15]
[1027]
[1028] txt_num_strings_minus1 + 1 represents the number of items in txt_descr_string_lang[ i ] and the txt_descr_string[ i ] that follows it.
[1029] txt_descr_string_lang[ i ] represents the language of txt_descr_string[ i ]. The language of txt_descr_string[ i ] is specified by a language tag defined in IETF RFC 5646. The length of txt_descr_string_lang[ i ] is from 0 to 49 (inclusive).
[1030] txt_descr_string[ i ] represents the ith text description information string interpreted as the value specified in txt_descr_purpose.
[1031] If txt_descr_purpose is 0, the interpretation of the information contained in txt_descr_string is defined by the application.
[1032] If txt_descr_purpose is 1, txt_descr_string[ i ] represents copyright information related to photos within the persistence range defined by txt_cancel_flag and txt_persistence_flag.
[1033] If txt_descr_purpose is 2, txt_descr_string[ i ], if not a null string, represents AI display information related to photos within the persistence scope of this SEI message.
[1034] Note: If txt_descr_purpose is 2, the string may contain information about machine learning-based processing, the intended use of the decoded photo, or other aspects related to the photo.
[1035] If txt_descr_purpose is 3, txt_descr_string[ i ] represents a plain text label description associated with the photo within the persistence range defined by txt_cancel_flag and txt_persistence_flag.
[1036] If txt_descr_purpose is 4, txt_descr_string[ i ] represents content recommendation rating information compliant with the US and Canadian Rating Territory Tables (RRT) for photos within the persistence range defined by txt_cancel_flag and txt_persistence_flag.
[1037] If txt_descr_purpose is 5, txt_descr_string[ i ] contains a tag URI with the syntax and semantics specified in IETF RFC 4151, which identifies the CLVS.
[1038] Figure 35 shows the syntax of photosensitive content information according to embodiments.
[1039] Regarding the technical problem addressed in the examples, paper JVET-AI0060, submitted to the 35th JVET conference, proposes signaling photosensitive content information, such as strobe lights / flashing lights, in images and image sequences. The proposer proposes two approaches for signaling this information.
[1040] Solution 1: Signaling a new type of text description in the SEI message.
[1041] Option 2: Define a new SEI message as follows.
[1042] FIG. 35 illustrates the photosensitive content information SEI message syntax in relation to the photosensitive content information SEI message. The semantics of the photosensitive content information SEI message are described with reference to FIG. 35.
[1043] If pci_cancel_flag is 1, it indicates that the SEI message cancels the persistence of the previous photosensitive content information SEI message in the output order. If pci_cancel_flag is 0, it indicates that the photosensitive content information follows.
[1044] pci_persistence_flag indicates the persistence of the photosensitive content information SEI message for the current layer.
[1045] If pci_persistence_flag is 0, it indicates that the photosensitive content information SEI message applies only to the currently decoded picture.
[1046] If pci_persistence_flag is 1, it indicates that the photosensitive content information SEI message applies to the currently decoded picture and persists for all subsequent pictures in the current layer in output order until one or more of the following conditions are met:
[1047] - A new CLVS for the current tier begins.
[1048] - The bitstream ends.
[1049] - The picture in the current layer of the AU associated with the SEI message for photosensitive content information is output after the current picture in the output order.
[1050] pci_severity_idc indicates the severity level of photosensitive content in the associated decoded picture. The defined values are as follows: 01 indicates low severity, 10 indicates medium severity, and 11 indicates high severity. 00 is reserved.
[1051] pci_light_types indicates the light types specified in Table 16 (Definition of pci_light_types) that are present in the associated decoded picture and cause photosensitivity. With respect to Table 16, if (pci_light_types & bitmask) is non-zero, it indicates that the specified light type is present in the associated decoded picture, and if (pci_light_types & bitmask) is 0, it indicates that the specified light type is not present in the associated decoded picture. If pci_light_types is 0, no information about the light type is specified.
[1052] [Table 16]
[1053]
[1054] pci_num_descr_strings indicates the number of entries for the following pci_descr_string_lang[ i ] and pci_descr_string[ i ].
[1055] pci_descr_string_lang[ i ] represents the language of pci_descr_string[ i ]. The language of pci_descr_string[ i ] is specified by a language tag defined in IETF RFC 5646.
[1056] pci_descr_string[ i ] represents textual description information about the photosensitive content of the associated decoded picture.
[1057] If pci_rectangle_info_present_flag is 0, it indicates that the photosensitive content information provided in the SEI message applies to the entire associated decoded picture. If pci_rectangle_info_present_flag is 1, it indicates that the rectangle syntax elements (pci_rect_top_left_x, pci_rect_top_left_y, pci_rect_width, pci_rect_height) follow. If pci_rectangle_info_present_flag is 1, the photosensitive content information is included only in the rectangle specified by (pci_rect_top_left_x, pci_rect_top_left_y, pci_rect_width, pci_rect_height) in the associated decoded picture.
[1058] pci_rect_top_left_x and pci_rect_top_left_y specify the luma sample position in horizontal and vertical pixel coordinates, respectively, representing the upper-left corner of the rectangle containing the photosensitive content in the associated decoded picture. pci_rect_top_left_x must be less than CroppedWidth, and pci_rect_top_left_y must be less than CroppedHeight.
[1059] pci_rect_width represents the width of the rectangle containing the photosensitive content in the associated decoded image. pci_rect_width is less than or equal to (CroppedWidth- pci_rect_top_left_x).
[1060] pci_rect_height represents the height of the rectangle containing the photosensitive content in the associated decoded image. pci_rect_height is less than or equal to (CroppedHeight- pci_rect_top_left_y).
[1061] Regarding the proposed new SEI (i.e., Option 2), it could be argued that the scope is too narrow, as it could include not only photosensitive content in the video but also other special content / events that the content sender wishes to notify the decoder / receiver of. For example, the types of content / events that the receiver should be notified of could include violent content or scenes of sexual or nudity.
[1062] Figure 36 shows a content event information SEI according to embodiments.
[1063] The methods and devices according to the embodiments can provide solutions to the technical problems described above as follows. Each embodiment can be applied individually or in combination.
[1064] 1. The method and device according to the embodiments can generate, encode, and decode a new SEI message called Content Event Information (CEI). The CEI according to the embodiments includes information about an event that may appear in a video associated with the SEI message.
[1065] 2. The CEI SEI message may contain the following information: - persistence information; - the number of events described in the SEI; - a list of event types. For example, the event types may be based on existing classifications, such as those defined by the Camera Ratings and Administration (CARA) or the Motion Picture Association of America; - optionally, the frame / picture location where the event occurs, etc.
[1066] 3. Additional signaling for other types of events may be added in the future.
[1067] 4. Event types may include: - sudden lighting changes; - sexual content; - violent content; - indications of illegal drugs; - and others.
[1068] The Content Event Information SEI message provides information about events that may exist in the picture associated with the SEI message. This information can be used to warn about content types that may raise concerns when displayed to certain groups / types of viewers.
[1069] If cei_cancel_flag is 1, it indicates that the SEI message cancels the persistence of the previous content event information SEI message in the output order. If cei_cancel_flag is 0, it indicates that the event information syntax element follows.
[1070] cei_persistence_flag indicates the persistence of content event information SEI messages for the current layer.
[1071] If cei_persistence_flag is 0, it indicates that the content event information SEI message applies only to the currently decoded picture.
[1072] If cei_persistence_flag is 1, it indicates that the content event information SEI message applies to the currently decoded picture and persists in output order for all subsequent pictures in the current layer until one or more of the following conditions are met:
[1073] - A new CLVS for the current tier begins.
[1074] - The bitstream ends.
[1075] - The picture in the current layer of the AU associated with the content event information SEI message is output after the current picture in the output order.
[1076] cei_num_events_minus1 + 1 represents the number of content events currently present. The value of cei_num_events_minus1 ranges from 0 to 255.
[1077] cei_event_type[ i ] represents the type of the ith event. The value of cei_event_type[ i ] is one of the values specified in Table 17 (List of event typesValue) below.
[1078] [Table 17]
[1079]
[1080] cei_severity_idc[ i ] represents the severity level of the ith event in the associated decoded picture. The defined values are as follows: 01 represents low severity, 10 represents medium severity, and 11 represents high severity. 00 is reserved.
[1081] If cei_location_info_present_flag[ i ] is 0, it indicates that the location of the ith event in the associated decoded picture is unknown. If cei_location_info_present_flag is 1, it indicates that the ith event exists within the region specified by the syntax elements cei_top_left_x[ i ], cei_top_left_y[ i ], cei_width[ i ], and cei_height[ i ].
[1082] cei_top_left_x[ i ] and cei_top_left_y[ i ] represent the luma sample positions in horizontal and vertical pixel coordinates, respectively, of the upper-left corner of the region containing the i-th event. cei_top_left_x[ i ] must be less than CroppedWidth, and cei_top_left_y[ i ] must be less than CroppedHeight.
[1083] cei_width[ i ] represents the width of the area containing the ith event. cei_width[ i ] is less than or equal to (CroppedWidth - cei_top_left_x[ i ]).
[1084] cei_height[ i ] represents the height of the area containing the ith event. cei_height[ i ] is less than or equal to (CroppedHeight - cei_top_left_y[ i ]).
[1085] Figure 37 shows an encoding method according to embodiments.
[1086] The encoding method according to the embodiments may include a step of deriving a SEI (Supplemental Enhancement Information) message for pictures (S3800); and / or a step of encoding pictures (S3810);
[1087] The step (S3800) of deriving an SEI (Supplemental Enhancement Information) message may be performed in the encoding process described above with reference to FIGS. 1 to 34. For example, the SEI message according to the embodiments may include HLS information, as illustrated in FIGS. 23 to 36.
[1088] The step of encoding pictures (S3810) may refer to the encoding process described above in FIGS. 1 to 34.
[1089] SEI messages may contain content event information about events that exist within pictures.
[1090] Content event information includes: information indicating the type of event, and the type of event may include information indicating a sudden light change within the pictures.
[1091] The type may further include information indicating other special contents or events including at least one of mature themes, language, or violence.
[1092] Content event information may further include information indicating whether location information of the event exists within the associated decoded picture.
[1093] The method of FIG. 37 may be performed by an encoding device. The encoding device according to embodiments includes a memory; and at least one processor connected to the memory; wherein the at least one processor may be configured to: derive a Supplemental Enhancement Information (SEI) message for pictures; and encode the pictures.
[1094] The embodiments further include a computer-readable storage medium storing a bitstream generated by the method according to FIG. 37.
[1095] Embodiments may include a step of obtaining a bitstream for image information, the bitstream being generated based on a step of deriving a Supplemental Enhancement Information (SEI) message for pictures; and a step of encoding the pictures; and a step of transmitting data including the bitstream.
[1096] Figure 38 shows a decryption method according to embodiments.
[1097] A decoding method according to embodiments may include a step of obtaining a SEI (Supplemental Enhancement Information) message for pictures (S3900); and / or a step of decoding pictures (S3910);
[1098] The decoding method of Fig. 38 and the encoding method of Fig. 37 are complementary to each other and can have an inverse relationship with each other.
[1099] The step (S3900) of obtaining an SEI (Supplemental Enhancement Information) message refers to the description of obtaining HLS information from a bitstream including encoded image information, as described in FIGS. 1 to 34, etc., as shown in FIGS. 23 to 36.
[1100] The step of decoding pictures (S3910) follows the decoding operation described in FIGS. 1 to 34.
[1101] With respect to the syntax element Content_event_information, an SEI message may contain Content Event Information about events present within pictures.
[1102] With respect to the syntax elements cei_event_type and / or pci_light_types, etc., the content event information: includes information indicating the type of the event, and the type of the event may include information indicating a sudden light change within the pictures.
[1103] Compared to syntax elements such as 'cei_event_type' and 'pci_light_types', the event type may further include information indicating that the event contains additional special contents or at least one of mature themes, language, or violence.
[1104] With respect to the syntax element cei_location_info_present_flag, the content event information may further include information indicating whether location information of the event within the associated decoded picture is present.
[1105] With respect to syntax elements cei_top_left_x[ i ], cei_top_left_y[ i ], cei_width[ i ], cei_height[ i ], etc., content event information may further include: a horizontal position of a luma sample of an area containing an event (e.g., an i-th event), a vertical position of a luma sample of an area containing an event, a width of an area containing an event, and a height of an area containing an event, based on information indicating whether position information of an event (e.g., an i-th event) exists.
[1106] The method of FIG. 38 may be performed by a decoder device. The decoder device according to embodiments includes a memory; and at least one processor connected to the memory; wherein the at least one processor may be configured to: obtain a Supplemental Enhancement Information (SEI) message for pictures; and decode the pictures.
[1107] The method and device according to the embodiments provide the following technical effects.
[1108] The method and device according to the embodiments has the effect of addressing the limitation that the range of a video image is too narrow to include not only photosensitive content but also other special content / events that the content sender wants to notify the decoder / receiver of, for example, the type of content / event that the receiver needs to be notified of can provide additional signaling regarding violent content or sexual or nudity scenes.
[1109] The embodiments have been described in terms of methods and / or devices, and the descriptions of methods and devices may be applied complementarily.
[1110] For the convenience of explanation, each drawing has been described separately, but it is also possible to design a new embodiment by combining the embodiments described in each drawing. In addition, designing a computer-readable recording medium having a program recorded thereon for executing the previously described embodiments, as needed by a person skilled in the art, also falls within the scope of the embodiments. The devices and methods according to the embodiments are not limited to the configurations and methods of the embodiments described above, but the embodiments may be configured by selectively combining all or part of the embodiments so that various modifications can be made. Although preferred embodiments of the embodiments have been illustrated and described, the embodiments are not limited to the specific embodiments described above, and various modifications can be made by a person skilled in the art to which the present invention pertains without departing from the gist of the embodiments claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the embodiments.
[1111] The various components of the devices of the embodiments may be implemented by hardware, software, firmware, or a combination thereof. The various components of the embodiments may be implemented by a single chip, for example, a single hardware circuit. According to embodiments, the components according to the embodiments may be implemented by separate chips. According to embodiments, at least one of the components of the devices of the embodiments may be configured with one or more processors capable of executing one or more programs, and the one or more programs may perform, or include instructions for performing, one or more of the operations / methods according to the embodiments. The executable instructions for performing the methods / operations of the devices of the embodiments may be stored in non-transitory CRMs or other computer program products configured to be executed by one or more processors, or may be stored in temporary CRMs or other computer program products configured to be executed by one or more processors. In addition, the memory according to the embodiments may be used as a concept including not only volatile memory (e.g., RAM, etc.), but also non-volatile memory, flash memory, PROM, etc. Additionally, it may be implemented in the form of a carrier wave, such as transmission via the Internet. Furthermore, the processor-readable recording medium may be distributed across network-connected computer systems, allowing the processor-readable code to be stored and executed in a distributed manner.
[1112] In this document, “ / ” and “,” are interpreted as “and / or”. For example, “A / B” is interpreted as “A and / or B”, and “A, B” is interpreted as “A and / or B”. Additionally, “A / B / C” means “at least one of A, B, and / or C”. Also, “A, B, C” means “at least one of A, B, and / or C”. Additionally, “or” in this document is interpreted as “and / or”. For example, “A or B” can mean 1) “A” only, 2) “B” only, or 3) “A and B”. In other words, “or” in this document can mean “additionally or alternatively”.
[1113] Terms such as "first," "second," etc. may be used to describe various components of the embodiments. However, the various components according to the embodiments should not be interpreted as limited by these terms. These terms are merely used to distinguish one component from another. For example, a first user input signal may be referred to as a "second user input signal." Similarly, a second user input signal may be referred to as a "first user input signal." The use of these terms should be interpreted as not departing from the scope of the various embodiments. Although "first user input signal" and "second user input signal" are both user input signals, they do not mean the same user input signals unless the context clearly indicates otherwise.
[1114] The terminology used to describe the embodiments is for the purpose of describing particular embodiments and is not intended to be limiting of the embodiments. As used in the description of the embodiments and in the claims, the singular is intended to include the plural unless the context clearly dictates otherwise. The expressions “and / or” are used to mean all possible combinations of terms. The expression “includes” describes the presence of features, numbers, steps, elements, and / or components, but does not mean that additional features, numbers, steps, elements, and / or components are not included. Conditional expressions such as “if” or “when” used to describe the embodiments are not intended to be limited to only optional cases. When a specific condition is satisfied, a related action is performed in response to a specific condition, or a related definition is intended to be interpreted.
[1115] Additionally, the operations according to the embodiments described in this document may be performed by a transceiver device including a memory and / or a processor according to the embodiments. The memory may store programs for processing / controlling the operations according to the embodiments, and the processor may control various operations described in this document. The processor may be referred to as a controller, etc. The operations according to the embodiments may be performed by firmware, software, and / or a combination thereof, and the firmware, software, and / or a combination thereof may be stored in the processor or in the memory.
[1116] Meanwhile, the operations according to the embodiments described above may be performed by a transmitting device and / or a receiving device according to the embodiments. The transmitting / receiving device may include a transmitting / receiving unit for transmitting and receiving media data, a memory for storing instructions (program code, algorithm, flowchart, and / or data) for a process according to the embodiments, and a processor for controlling the operations of the transmitting / receiving device.
[1117] The processor may be referred to as a controller or the like, and may correspond to, for example, hardware, software, and / or a combination thereof. The operations according to the above-described embodiments may be performed by the processor. Furthermore, the processor may be implemented as an encoder / decoder or the like for the operations of the above-described embodiments.
[1118] As described above, the relevant contents have been described in the best form for carrying out the embodiments.
[1119] As described above, the embodiments may be applied in whole or in part to an image encoding method, an image encoding device, an image decoding method, an image decoding device, a method for transmitting a bitstream, and a recording medium storing a bitstream.
[1120] Those skilled in the art may make various changes or modifications to the embodiments within the scope of the embodiments.
[1121] Embodiments may include modifications / changes, which do not depart from the scope of the claims and their equivalents.
Claims
A step of obtaining SEI (Supplemental Enhancement Information) messages for pictures; and A step of decoding the above pictures; comprising; method. In the first paragraph, The above SEI message includes content event information about events existing in the pictures. method. In the second paragraph, The above content event information is: Contains information indicating the type of the above event, The type of the above event includes information indicating a sudden light change within the above pictures. method. In the third paragraph, The type of the above event further includes information indicating that the event includes additional special content or at least one of mature theme, language, or violence. method. In the second paragraph, The above content event information is: Further including information indicating whether there is location information of an event within the associated decoded picture, method. In paragraph 5, The above content event information is: Based on information indicating whether location information of the above event exists, Further comprising a horizontal position of a luma sample of an area containing the event, a vertical position of a luma sample of an area containing the event, a width of an area containing the event, and a height of an area containing the event. method. memory; and At least one processor connected to the memory; wherein the at least one processor comprises: Obtaining SEI (Supplemental Enhancement Information) messages for pictures; and configured to decode the above pictures; device. A step of deriving SEI (Supplemental Enhancement Information) messages for pictures; and A step of encoding the above pictures; comprising: method. In paragraph 8, The above SEI message includes content event information about events existing in the pictures. method. In paragraph 9, The above content event information is: Contains information indicating the type of the above event, The type of the above event includes information indicating a sudden light change within the above pictures. method. In Article 10, The type of the above event further includes information indicating that the event includes additional special content or at least one of mature theme, language, or violence. method. In paragraph 9, The above content event information is: Further including information indicating whether there is location information of an event within the associated decoded picture, method. memory; and At least one processor connected to a memory; wherein the at least one processor comprises: Deriving SEI (Supplemental Enhancement Information) messages for pictures; and Encoding the above pictures; configured to do so; device. A computer-readable storage medium storing a bitstream generated by the method according to Article 8. A step of obtaining a bitstream for video information, The bitstream is generated based on the steps of: deriving SEI (Supplemental Enhancement Information) messages for pictures; and encoding the pictures; and A step of transmitting data including the bitstream; comprising: method.
Citation Information
Patent Citations
Image processing system and image processing method
JP2018186528A
Video information decoding method, video decoding method and apparatus using the same
KR1020180035760A
Insulin Syringe Replacement System Applicable to Insulin Pump
KR1020230165958A
Method, program, and apparatus for building model optimized for analysis of bio signal
KR102893316B1