Video encoding method, video encoding apparatus, video decoding method, video decoding apparatus, method for transmitting bitstream, and recording medium on which bitstream is stored

WO2026197755A1PCT designated stage Publication Date: 2026-09-24LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/004286
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-17
Filing Date
2026-03-17
Publication Date
2026-09-24

Smart Images

  • Figure KR2026004286_24092026_PF_FP_ABST
    Figure KR2026004286_24092026_PF_FP_ABST
Patent Text Reader

Abstract

A method according to embodiments may comprise the steps of: acquiring a bitstream; acquiring at least one SEI message from the bitstream; and decoding a picture in the bitstream. A method according to embodiments may comprise the steps of: generating at least one SEI message; encoding a picture; and generating a bitstream including the at least one SEI message and the picture.
Need to check novelty before this filing date? Find Prior Art

Description

Image encoding method, image encoding device, image decoding method, image decoding device, method for transmitting a bitstream and a recording medium storing a bitstream

[0001] The embodiments relate to an image encoding method, an image encoding device, an image decoding method, an image decoding device, a method for transmitting a bitstream, and a recording medium storing a bitstream.

[0002] Recently, the demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition), has been increasing across various fields. As video data becomes higher in resolution and quality, the relative amount of information or bits transmitted increases compared to conventional video data. This increase in the amount of transmitted information or bits leads to higher transmission and storage costs.

[0003] Accordingly, high-efficiency video compression technology is required to effectively transmit, store, and play back high-resolution, high-quality video information.

[0004] The embodiments provide an image encoding method, an image encoding device, an image decoding method, an image decoding device, a method for transmitting a bitstream, and a recording medium storing a bitstream.

[0005] The embodiments provide an image encoding method with improved encoding and decoding efficiency, an image encoding device, an image decoding method, an image decoding device, a method for transmitting a bitstream, and a recording medium storing a bitstream.

[0006] However, the scope of rights of the embodiments is not limited to the technical problems described above, and may be extended to other technical problems that a person skilled in the art can infer based on the entire content described.

[0007] A method according to the embodiments may include the steps of: acquiring a bitstream; acquiring at least one SEI message from the bitstream; and decoding a picture within the bitstream. A method according to the embodiments may include the steps of: generating at least one SEI message; encoding a picture; and generating a bitstream comprising at least one SEI message and a picture.

[0008] The embodiments provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0009] The embodiments provide a non-transient computer-readable recording medium that stores a bitstream generated by an image encoding method.

[0010] The embodiments provide a non-transient computer-readable recording medium that stores a bitstream received and decoded by an image decoding device and used for image restoration.

[0011] The embodiments provide a method for transmitting a bitstream generated by an image encoding method.

[0012] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs from the description below.

[0013] Drawings are included to further understand the embodiments, and the drawings illustrate the embodiments along with descriptions related to the embodiments. For a better understanding of the various embodiments described below, one must refer to the description of the embodiments below in relation to the following drawings, which include parts corresponding to similar reference numerals throughout the drawings.

[0014] FIG. 1 shows a video and / or image coding system according to embodiments.

[0015] FIG. 2 shows an encoding device according to embodiments.

[0016] FIG. 3 shows a decoding device according to embodiments.

[0017] Figure 4 shows the structure of a content streaming system according to embodiments.

[0018] FIG. 5 shows an example of a picture divided into Coding Tree Units (CTUs) according to embodiments.

[0019] FIG. 6 shows an example of a picture partitioned into tiles and raster-scan slices according to embodiments.

[0020] FIG. 7 shows an example of a picture partitioned into tiles and raster-scan slices according to embodiments.

[0021] FIG. 8 shows an example of a picture partitioned into tiles, bricks, and rectangular slices according to embodiments.

[0022] FIG. 9 shows an example of a picture including subpictures according to embodiments.

[0023] FIG. 10 shows an example of a picture including tiles and CTUs according to embodiments.

[0024] FIG. 11 shows a multi-type tree splitting mode according to embodiments.

[0025] FIG. 12 shows splitting flags within a quad tree of a multi-type tree coding structure according to embodiments.

[0026] FIG. 13 shows an example of a quad tree of a multi-type tree coding block structure according to embodiments.

[0027] FIG. 14 shows the prohibition of TT (Ternary Tree) division for coding blocks according to embodiments.

[0028] FIG. 15 shows the transform and inverse transform according to the embodiments.

[0029] FIG. 16 shows a Low-Frequency Non-Separable Transform (LFNST) according to embodiments.

[0030] FIG. 17 shows CABAC (Context Adaptive Binary Arithmetic Coding) encoding according to embodiments.

[0031] FIG. 18 illustrates an entropy encoding method according to embodiments.

[0032] FIG. 19 illustrates an entropy decoding method according to embodiments.

[0033] FIG. 20 illustrates a picture decoding method according to embodiments.

[0034] FIG. 21 illustrates a picture encoding method according to embodiments.

[0035] FIG. 22 shows a hierarchical structure for a coded image according to embodiments.

[0036] FIGS. 23a, FIGS. 23b, FIGS. 23c, FIGS. 23d, and FIGS. 23e show picture header structures according to embodiments.

[0037] FIGS. 24a, FIGS. 24b, and FIGS. 24c show neural-network post-filter characteristics SEI message syntax according to embodiments.

[0038] FIG. 25 illustrates the process of inducing a luma channel in a luma component according to the embodiments.

[0039] FIG. 26 shows the neural-network post-filter activation SEI message syntax according to the embodiments.

[0040] FIGS. 27a, FIGS. 27b, FIGS. 27c, and FIGS. 27d illustrate constituent rectangles SEI message syntax according to embodiments.

[0041] FIGS. 28a and FIGS. 28b show display rectangles SEI message syntax according to embodiments.

[0042] FIG. 29 shows the display rectangles SEI message syntax according to the embodiments.

[0043] FIGS. 30a and FIG. 30b show display rectangles SEI message syntax according to embodiments.

[0044] FIG. 31 shows a encoding method according to embodiments.

[0045] FIG. 32 illustrates a decoding method according to embodiments.

[0046] Preferred embodiments of the embodiments are described in detail, and examples thereof are shown in the accompanying drawings. The detailed description below, with reference to the accompanying drawings, is intended to describe preferred embodiments of the embodiments rather than merely embodiments that may be implemented according to the embodiments. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it is obvious to those skilled in the art that the embodiments can be practiced without these details.

[0047] Most terms used in the embodiments are selected from those commonly used in the field, but some terms are chosen at the applicant's discretion, and their meanings are described in detail in the following description as necessary. Accordingly, the embodiments should be understood based on the intended meaning of the terms, rather than their mere names or meanings.

[0048] Related technical fields: Versatile Video Coding (VVC), Versatile supplemental enhancement information messages for coded video bitstreams (VSEI), Additional SEI messages for VSEI (Draft 3), SEI processing order and processing order nesting SEI messages in VVC (draft 7), Technologies under consideration for future extensions of VSEI (draft 4), SEI messages for VSEI version 4 (Draft 2).

[0049] FIG. 1 shows a video and / or image coding system according to embodiments.

[0050] As shown in FIG. 1, a video / image coding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image information or data to the receiving device in the form of a file or streaming via a digital storage medium or a network.

[0051] The source device may include a video source, an encoding device, and a transmission unit. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.

[0052] A video source may acquire video / images through processes such as video / image capture, synthesis, or generation. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include, for example, a computer, a tablet, and a smartphone, etc., and may generate video / images (electronically). For example, virtual video / images may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process in which related data is generated.

[0053] The encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0054] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device in the form of a file or streaming via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmission unit may include elements for creating a media file through a predetermined file format and elements for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.

[0055] The decoding device can decode video / images by performing a series of procedures such as inverse quantization, inverse transform, and prediction corresponding to the operation of the encoding device.

[0056] The renderer can render the decoded video / image. The rendered video / image can be displayed through the display unit.

[0057] This document relates to video / video coding. For example, the methods / executions disclosed in this document may be applied to methods disclosed in the VVC (versatile video coding) standard, EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard) or next-generation video / video coding standards (e.g., H.267 or H.268).

[0058] This document presents various embodiments regarding video / image coding, and unless otherwise noted, the embodiments may be performed in combination with one another.

[0059] In this document, "video" may refer to a set of images over time. "Picture" generally refers to a unit representing a single image at a specific time, and "slice" or "tile" are units that constitute a part of a picture in coding. A slice or tile may contain one or more CTUs (coding tree units). A single picture may consist of one or more slices or tiles. A single picture may consist of one or more tile groups. A tile group may contain one or more tiles. A "brick" may represent a rectangular area of ​​rows of CTUs within a tile in a picture.

[0060] A brick can represent a rectangular area of ​​a row of CTUs within a tile in a picture. A tile can be divided into multiple bricks, and each brick consists of one or more rows of CTUs within the tile. A tile that is not divided into multiple bricks is also referred to as a brick. A brick scan is a specific sequential order of CTUs that divides a picture. In a brick, CTUs are arranged sequentially as CTU raster scans; in a brick within a tile, bricks are arranged sequentially as brick raster scans of the tile; and in a tile within a picture, tiles are arranged sequentially as tile raster scans of the picture. A tile is a rectangular area of ​​CTUs within a specific tile column and a specific tile row in a picture. A tile column is a rectangular area of ​​CTUs that is equal to the height of the picture and has a width specified by the syntax element of the picture parameter set. A tile row is a rectangular area of ​​CTUs that has a height specified by the syntax element of the picture parameter set and a width equal to the width of the picture. A tile scan is a specific sequential order of CTUs that divides a picture. In tiles, CTUs are continuously aligned by CTU raster scans, and in pictures, tiles are continuously aligned by tile raster scans. A slice contains an integer number of bricks of a picture, which can be contained exclusively in a single NAL unit. A slice can consist of multiple complete tiles or a sequence in which the complete bricks of a single tile are arranged continuously.

[0061] In this document, tile group and slice may be used interchangeably. For example, in this document, tile group / tile group header may be referred to as slice / slice header.

[0062] A pixel or pel can refer to the smallest unit that constitutes a picture (or image). Additionally, the term 'sample' may be used as a counterpart to pixel. Generally, a sample can represent a pixel or its value, and it may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chroma component.

[0063] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of ​​a picture and information related to that area. A unit may include one luminance block and two chroma (e.g., cb, cr) blocks. Depending on the case, the term unit may be used interchangeably with terms such as block or area. In general, an MxN block may include samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows.

[0064] FIG. 2 shows an encoding device according to embodiments.

[0065] FIG. 2 shows a schematic block diagram of an encoding device to which the embodiment(s) of the present document can be applied and to which video / image signal encoding is performed.

[0066] As shown in FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a predictor (220), a residual processor (230), an entropy encoder (240), an adder (250), a filter (260), and a memory (270). The predictor (220) may include an inter-predictor (221) and an intra-predictor (222). The residual processor (230) may include a transformer (232), a quantizer (233), a dequantizer (234), and an inverse transformer (235). The residual processor (230) may further include a subtractor (231). The addition unit (250) may be referred to as a reconstructor or a reconstructed block generator. The above-described image segmentation unit (210), prediction unit (220), residual processing unit (230), entropy encoding unit (240), addition unit (250), and filtering unit (260) may be configured by one or more hardware components (e.g., an encoder chipset or processor) according to the embodiment. Additionally, the memory (270) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.

[0067] The image segmentation unit (210) can divide an input image (or picture, frame) input to an encoding device (200) into one or more processing units. For example, a processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, a single coding unit may be divided into multiple coding units of a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first and the binary-tree structure and / or ternary structure may be applied later. Or, the binary-tree structure may be applied first. A coding procedure according to this document may be performed based on the final coding unit that is no longer divided. In this case, based on coding efficiency according to image characteristics, the maximum coding unit may be used directly as the final coding unit, or, if necessary, the coding unit may be recursively divided into lower-depth coding units so that a coding unit of the optimal size is used as the final coding unit. Here, the term "coding procedure" may include procedures such as prediction, transformation, and restoration described below. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may each be divided or partitioned from the aforementioned final coding unit. The prediction unit may be a unit for sample prediction, and the transformation unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from transformation coefficients.

[0068] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the context. In general, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chroma component. A sample may be used to refer to a single picture (or image) as a term corresponding to a pixel or pel.

[0069] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (predicted block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input image signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, as illustrated, the unit that subtracts the prediction signal (predicted block, prediction sample array) from the input image signal (original block, original sample array) within the encoder (200) may be called a subtraction unit (231). The prediction unit performs a prediction for a block to be processed (hereinafter referred to as the current block) and can generate a predicted block containing prediction samples for the current block. The prediction unit can determine whether intra prediction is applied or inter prediction is applied at the current block or CU level. The prediction unit can generate various information regarding prediction, such as prediction mode information, as described below in the description of each prediction mode, and transmit it to the entropy encoding unit (240). The information regarding prediction can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.

[0070] The intra prediction unit (222) can predict the current block by referencing samples within the current picture. The referenced samples may be located near the current block or away from it, depending on the prediction mode. In intra prediction, the prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a Planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is merely an example, and depending on the settings, more or fewer directional prediction modes may be used. The intra prediction unit (222) may also determine the prediction mode applied to the current block by using the prediction mode applied to the surrounding blocks.

[0071] The inter prediction unit (221) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference picture indices. Motion information may further include information on inter prediction directions (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture containing the reference blocks and the reference picture containing the temporal neighboring blocks may be the same or different. Temporal surrounding blocks may be referred to by names such as collocated reference block, collocated CU (colCU), etc., and a reference picture containing temporal surrounding blocks may be referred to as a collocated picture (colPic). For example, the inter prediction unit (221) may construct a list of motion information candidates based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of surrounding blocks as motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0072] The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit may apply intra prediction or inter prediction for a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called combined inter and intra prediction (CIIP). Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, for example, screen content coding (SCC). IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. IBC may utilize at least one of the inter prediction techniques described in this document. The palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, sample values ​​within the picture can be signaled based on information regarding the palette table and palette index.

[0073] The prediction signal generated through the prediction unit (including the inter prediction unit (221) and / or the intra prediction unit (222)) may be used to generate a restored signal or to generate a residual signal. The transformation unit (232) may generate transform coefficients by applying a transformation technique to the residual signal. For example, the transformation technique may include at least one of the Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loeve Transform (KLT), Graph-Based Transform (GBT), or Conditionally Non-linear Transform (CNT). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transformation obtained based on generating a prediction signal using all previously reconstructed pixels. In addition, the transformation process can be applied to square pixel blocks of the same size, or to non-square blocks of variable size.

[0074] The quantization unit (233) quantizes the transformation coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information regarding the quantized transformation coefficients) and output it as a bitstream. The information regarding the quantized transformation coefficients may be called residual information. The quantization unit (233) can rearrange the block-shaped quantized transformation coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information regarding the quantized transformation coefficients based on the one-dimensional vector-shaped quantized transformation coefficients. The entropy encoding unit (240) can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit (240) may encode information necessary for video / image restoration (e.g., values ​​of syntax elements) together or separately, in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information regarding various parameter sets, such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Additionally, the video / image information may further include general constraint information. In this document, information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded through the encoding procedure described above and included in the bitstream.The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) that transmits the signal output from the entropy encoding unit (240) and / or a storage unit (not shown) that stores it may be configured as internal / external elements of the encoding device (200), or the transmission unit may be included in the entropy encoding unit (240).

[0075] Quantized transformation coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, a residual signal (residual block or residual samples) can be restored by applying inverse quantization and inverse transformation to the quantized transformation coefficients through the inverse quantization unit (234) and the inverse transformation unit (235). An adder (155) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter-prediction unit (221) or the intra-prediction unit (222). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The adder (250) may be called a reconstruction unit or a reconstruction block generation unit. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, and can also be used for inter prediction of the next picture after filtering as described below.

[0076] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.

[0077] The filtering unit (260) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (260) can generate a modified restored picture by applying various filtering methods to the restored picture, and can store the modified restored picture in memory (270), specifically in the DPB of memory (270). Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (240), as described below in the description of each filtering method. The information regarding filtering can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.

[0078] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter-prediction unit (221). Through this, when inter-prediction is applied, the encoding device can avoid prediction mismatches between the encoding device (100) and the decoding device, and can also improve encoding efficiency.

[0079] The memory (270) DPB can store the modified restored picture to be used as a reference picture in the inter-prediction unit (221). The memory (270) can store motion information of blocks from which motion information is derived (or encoded) within the current picture and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (221) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (270) can store restoration samples of the blocks restored within the current picture and transmit them to the intra-prediction unit (222).

[0080] FIG. 3 shows a decoding device according to embodiments.

[0081] FIG. 3 shows a schematic block diagram of a decoding device to which the embodiment(s) of the present document can be applied and to which decoding of a video / image signal is performed.

[0082]

[0083] As shown in FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-predictor (331) and an intra-predictor (332). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321). The aforementioned entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) may be configured by a single hardware component (e.g., a decoder chipset or a processor) according to an embodiment. Additionally, the memory (360) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.

[0084] When a bitstream containing video / image information is input, the decoding device (300) can restore the image in correspondence with the process in which the video / image information is processed by the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied by the encoding device. Thus, the processing unit for decoding may be, for example, a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a binary tree structure. One or more conversion units may be derived from the coding unit. And, the restored image signal decoded and output through the decoding device (300) can be played back through a playback device.

[0085] The decoding device (300) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through the entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may further include general constraint information. The decoding device can decode the picture based further on information regarding the parameter sets and / or general constraint information. The signaling / receiving information and / or syntax elements described below in this document can be obtained from the bitstream by decoding through a decoding procedure. For example, the entropy decoding unit (310) can decode information within a bitstream based on coding methods such as exponential chord coding, CAVLC, or CABAC, and output values ​​of syntax elements required for image restoration and quantized values ​​of transformation coefficients regarding residuals. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, determine a context model using information of the syntax element to be decoded and decoding information of surrounding and decoding target blocks or information of a symbol / bin decoded in the previous step, predict the probability of occurrence of the bin according to the determined context model, and perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model.Information regarding prediction among the information decoded in the entropy decoding unit (310) is provided to the prediction unit (inter prediction unit (332) and intra prediction unit (331)), and residual values ​​for which entropy decoding has been performed in the entropy decoding unit (310), for example, quantized transformation coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive residual signals (residual blocks, residual samples, residual sample arrays). Additionally, information regarding filtering among the information decoded in the entropy decoding unit (310) can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (300), or the receiving unit may be a component of the entropy decoding unit (310). Meanwhile, the decoding device according to the present document may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoding unit (310), and the sample decoder may include at least one of an inverse quantization unit (321), an inverse transform unit (322), an adder (340), a filtering unit (350), a memory (360), an inter prediction unit (332), and an intra prediction unit (331).

[0086] In the inverse quantization unit (321), the quantized transformation coefficients can be inversely quantized to output transformation coefficients. The inverse quantization unit (321) can rearrange the quantized transformation coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain transformation coefficients.

[0087] In the inverse conversion unit (322), the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).

[0088] The prediction unit can perform a prediction for the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit (310), the prediction unit can determine whether an intra prediction or an inter prediction is applied to the current block and can determine a specific intra / inter prediction mode.

[0089] The prediction unit (320) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit may apply intra prediction or inter prediction for a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called combined inter and intra prediction (CIIP). Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, for example, screen content coding (SCC). IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. IBC may utilize at least one of the inter prediction techniques described in this document. The palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and palette index may be included in the video / image information and signaled.

[0090] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The referenced samples may be located near the current block or away from it, depending on the prediction mode. In intra prediction, the prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit (331) may determine the prediction mode applied to the current block by using the prediction mode applied to the surrounding blocks.

[0091] The inter prediction unit (332) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. Motion information may include a motion vector and a reference picture index. Motion information may further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) may construct a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction may be performed based on various prediction modes, and information regarding the prediction may include information indicating the mode of inter prediction for the current block.

[0092] The adder (340) can generate a restoration signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the prediction signal (predicted block, predicted sample array) output from the prediction unit (including the inter prediction unit (332) and / or the intra prediction unit (331)). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the restoration block.

[0093] The addition unit (340) may be called a restoration unit or a restoration block generation unit. The generated restoration signal may be used for intra-predicting the next block to be processed within the current picture, may be output after filtering as described below, or may be used for inter-predicting the next picture.

[0094] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.

[0095] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can generate a modified restored picture by applying various filtering methods to the restored picture, and can transmit the modified restored picture to memory (360), specifically to the DPB of memory (360). Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0096] The (modified) restored picture stored in the DPB of the memory (360) can be used as a reference picture in the inter-prediction unit (332). The memory (360) can store motion information of blocks from which motion information within the current picture has been derived (or decoded) and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (260) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (360) can store restoration samples of blocks restored within the current picture and transmit them to the intra-prediction unit (331).

[0097] In this specification, the embodiments described in the filtering unit (260), inter prediction unit (221), and intra prediction unit (222) of the encoding device (100) may be applied to the filtering unit (350), inter prediction unit (332), and intra prediction unit (331) of the decoding device (300) in the same or corresponding manner.

[0098] Implementation and Application Examples:

[0099] The embodiments described in this document may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each figure may be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for implementation may be stored on a digital storage medium.

[0100] In addition, the decoding device and encoding device to which the embodiment(s) of this document apply may be included in multimedia broadcasting transmission and reception devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video conversation devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, Video on Demand (VoD) service providers, Over-the-top video (OTT) devices, internet streaming service providers, 3D video devices, virtual reality (VR) devices, augmented reality (AR) devices, video phone video devices, transportation terminals (e.g., vehicle terminals (including autonomous vehicles), airplane terminals, ship terminals, etc.), and medical video devices, and may be used to process video signals or data signals. For example, Over-the-top video (OTT) devices may include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, Digital Video Recorders (DVRs), etc.

[0101] Additionally, the processing method to which the embodiment(s) of this document are applied may be produced in the form of a program that is executed by a computer and may be stored on a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of this document may also be stored on a computer-readable recording medium. A computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. A computer-readable recording medium may include, for example, a Blu-ray disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. Additionally, a computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission over the Internet). Additionally, a bitstream generated by an encoding method may be stored on a computer-readable recording medium or transmitted via a wired or wireless communication network.

[0102] Additionally, the embodiment(s) of this document may be implemented as a computer program product by program code, and the program code may be executed on a computer by the embodiment(s) of this document. The program code may be stored on a computer-readable carrier.

[0103] Figure 4 shows the structure of a content streaming system according to embodiments.

[0104] A content streaming system to which the embodiment(s) of this document apply may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0105] The encoding server compresses content input from multimedia input devices, such as smartphones, cameras, and camcorders, into digital data to generate a bitstream, and transmits it to the streaming server. As another example, if multimedia input devices, such as smartphones, cameras, and camcorders, generate the bitstream directly, the encoding server can be omitted.

[0106] A bitstream may be generated by an encoding method or a bitstream generation method to which the embodiment(s) of this document are applied, and a streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0107] The streaming server transmits multimedia data to the user's device based on user requests made through the web server, while the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server forwards the request to the streaming server, which then transmits the multimedia data to the user. In this process, the content streaming system may include a separate control server, which plays the role of managing commands and responses between devices within the content streaming system.

[0108] A streaming server can receive content from a media storage and / or an encoding server. For example, if content is received from an encoding server, it can be received in real time. In this case, to provide a seamless streaming service, the streaming server may store the bitstream for a certain period of time.

[0109] Examples of user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs (head-mounted displays)), digital TVs, desktop computers, digital signage, etc.

[0110] Each server within the content streaming system can be operated as a distributed server, and in this case, data received from each server can be processed in a distributed manner.

[0111] Partitioning structure:

[0112] The video / image coding method according to this document can be performed based on the following partitioning structure. Specifically, the procedures described below, such as prediction, residual processing ((inverse)transform, (inverse)quantization, etc.), syntax element coding, and filtering, can be performed based on CTU and CU (and / or TU, PU) derived based on the partitioning structure. The block partitioning procedure is performed in the image splitting unit (210) of the encoding device described above, and the partitioning-related information can be processed (encoded) in the entropy encoding unit (240) and transmitted to the decoding device in the form of a bitstream. The entropy decoding unit (310) of the decoding device can derive the block partitioning structure of the current picture based on the partitioning-related information obtained from the bitstream, and perform a series of procedures for image decoding (e.g., prediction, residual processing, block / picture restoration, in-loop filtering, etc.) based thereon. The CU size and the TU size may be the same, or multiple TUs may exist within the CU area. Meanwhile, the term CU size generally refers to the luminance component (sample) CB size. The term TU size generally refers to the luminance component (sample) TB size. The chroma component (sample) CB or TB size can be derived based on the luminance component (sample) CB or TB size according to the component ratio based on the color format (chroma format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.) of the picture / image. The TU size can be derived based on maxTbSize. For example, if the CU size is greater than maxTbSize, multiple TUs (TBs) of maxTbSize are derived from C, and conversion / inverse conversion can be performed in TU (TB) units. In addition, for example, when intra prediction is applied, the intra prediction mode / type is derived in units of CU (or CB), and the procedure for deriving surrounding reference samples and generating prediction samples can be performed in units of TU (or TB).In this case, one or more TUs (or TBs) may exist within a single CU (or CB) region, and in this case, the multiple TUs (or TBs) may share the same intra prediction mode / type.

[0113] Additionally, in the coding of video / images according to this document, image processing units may have a hierarchical structure. A picture may be divided into one or more tiles, bricks, slices, and / or tile groups. A slice may contain one or more bricks. A brick may contain one or more CTU rows within a tile. A slice may contain an integer number of bricks in a picture. A tile group may contain one or more tiles. A tile may contain one or more CTUs. A CTU may be divided into one or more CUs. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. A tile group may contain an integer number of tiles based on a tile raster scan within a picture. A slice header may carry information / parameters that can be applied to the corresponding slice (blocks within the slice). If the encoding / decoding device has a multi-core processor, the encoding / decoding procedures for tiles, slices, bricks, and / or tile groups may be processed in parallel. In this document, the terms slice and tile group may be used interchangeably. A tile group header may be referred to as a slice header. Here, a slice may have one of the slice types, including intra (I) slice, predictive (P) slice, and bi-predictive (B) slice. For blocks within an I slice, inter-prediction is not used for prediction, and only intra-prediction may be used. Of course, even in this case, the original sample value may be coded and signaled without prediction.For blocks within a P slice, intra prediction or inter prediction may be used, and if inter prediction is used, only uni prediction may be used. Meanwhile, for blocks within a B slice, intra prediction or inter prediction may be used, and if inter prediction is used, up to bi prediction may be used.

[0114] In the encoder, tile / tile group, brick, slice, and maximum and minimum coding unit sizes are determined based on video characteristics (e.g., resolution) or by considering coding efficiency or parallel processing, and information regarding this or information that can derive it may be included in the bitstream.

[0115] The decoder can obtain information indicating whether the tile / tile group, brick, slias, and CTU within the tile of the current picture have been divided into multiple coding units. Efficiency can be increased by obtaining (transmitting) this information only under specific conditions.

[0116] A slice header (slice header syntax) may include information / parameters that can be applied commonly to slices. An APS (APS syntax) or PPS (PPS syntax) may include information / parameters that can be applied commonly to one or more pictures. An SPS (SPS syntax) may include information / parameters that can be applied commonly to one or more sequences. A VPS (VPS syntax) may include information / parameters that can be applied commonly to multiple layers. A DPS (DPS syntax) may include information / parameters that can be applied commonly across the video. A DPS may include information / parameters related to the concatenation of a CVS (coded video sequence).

[0117] In this document, the term "higher-level syntax" may include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, and slice header syntax.

[0118] In addition, for example, information regarding the division and configuration of tiles / tile groups / bricks / slices can be configured at the encoding stage through high-level syntax and transmitted to a decoding device in the form of a bitstream.

[0119] FIG. 5 shows an example of a picture divided into Coding Tree Units (CTUs) according to embodiments.

[0120] Partitioning of picture into CTUs:

[0121] Pictures can be divided into a sequence of coding tree units (CTUs). A CTU may correspond to a coding tree block (CTB). Alternatively, a CTU may include a coding tree block of luminance samples and two coding tree blocks of corresponding chroma samples. In other words, for a picture containing three sample arrays, a CTU may include an NxN block of luminance samples and two corresponding blocks of chroma samples. FIG. 5 illustrates an example in which a picture is divided into CTUs.

[0122] The maximum allowable size of a CTU for coding and prediction, etc., may differ from the maximum allowable size of a CTU for transformation. For example, the maximum allowable size of a luminance block within a CTU may be 128x128 (even though the maximum size of luminance ring blocks is 64x64).

[0123] FIG. 6 shows an example of a picture partitioned into tiles and raster-scan slices according to embodiments.

[0124] Partitioning of pictures into subpictures, slices, and tiles:

[0125] A picture is divided into one or more tile rows and one or more tile columns. A tile is a sequence of CTUs that includes a rectangular area of ​​the picture. The CTUs within a tile are scanned in the raster scan order within that tile.

[0126] A slice consists of an integer number of complete tiles or an integer number of consecutive complete CTU rows within a picture tile.

[0127] Two modes are supported for slicing: raster scan slice mode and rectangular slice mode. In raster scan slice mode, a slice contains a complete sequence of tiles from a tile raster scan of the picture. In rectangular slice mode, a slice contains multiple complete tiles that make up a rectangular area of ​​the picture, or multiple consecutive rows of complete CTUs of a single tile that make up a rectangular area of ​​the picture. The tiles within a rectangular slice are scanned in the tile raster scan order within the rectangular area corresponding to that slice.

[0128] A sub-picture consists of one or more slices that cover the entire rectangular area of ​​the picture.

[0129] Figure 6 shows an example of splitting a picture into raster scan slices. Here, the picture is divided into 12 tiles and 3 raster scan slices.

[0130] FIG. 7 shows an example of a picture partitioned into tiles and raster-scan slices according to embodiments.

[0131] Figure 7 shows an example of dividing a picture into rectangular slices. Here, the picture is divided into 24 tiles (6 tile columns and 4 tile rows) and 9 rectangular slices.

[0132] FIG. 8 shows an example of a picture partitioned into tiles, bricks, and rectangular slices according to embodiments.

[0133] Figure 8 shows an example of a picture divided into tiles and rectangular slices. Here, the picture is divided into 4 tiles (2 tile columns and 2 tile rows) and 4 rectangular slices.

[0134] FIG. 9 shows an example of a picture including subpictures according to embodiments.

[0135] Fig. 9 shows an example of sub-picture division of a picture. Here, the picture is divided into 28 sub-pictures of various dimensions.

[0136] FIG. 10 shows an example of a picture including tiles and CTUs according to embodiments.

[0137] If the picture is coded using three separate color planes (where separate_colour_plane_flag is 1), the slice contains only one CTU of a color component identified by its color_plane_id value, and each array of color components in the picture consists of slices having the same color_plane_id value. Coded slices with different color_plane_id values ​​within the picture may be interleaved with each other under the constraint that for each color_plane_id value, the coded slice NAL unit having that color_plane_id value must be in ascending order of CTU addresses in the tile scan order for the first CTU of each coded slice NAL unit.

[0138] Note - If separate_colour_plane_flag is 0, each CTU of the picture is contained in exactly one slice. If separate_colour_plane_flag is 1, each CTU of the color component is contained in exactly one slice (information for each CTU of the picture exists in exactly three slices, and these three slices have different colour_plane_id values).

[0139] Tile changes the order of CTUs in a picture. If the picture is divided into two or more tiles, the order of CTUs is the raster scan order within each tile, as shown in FIG. 10. In FIG. 10, the picture is divided into two tiles, and each tile has eight CTUs. The order of CTUs within the tile is the raster scan order.

[0140] FIG. 11 shows a multi-type tree splitting mode according to embodiments.

[0141] Partitioning of the CTUs using a tree structure

[0142] A CTU can be partitioned into CUs based on a quad-tree (QT) structure. The quad-tree structure can be referred to as a quaternary tree structure. This is intended to reflect various local characteristics. Meanwhile, in this document, a CTU can be partitioned based on a multitype tree structure partitioning that includes not only quad-trees but also binary trees (BT) and ternary trees (TT). Hereinafter, the term QTBT structure may include quad-tree and binary tree-based partitioning structures, and QTBTTT may include quad-tree, binary tree, and ternary tree-based partitioning structures. Alternatively, the QTBT structure may include quad-tree, binary tree, and ternary tree-based partitioning structures. In a coding tree structure, CUs can have a square or rectangular shape. A CTU can first be partitioned into a quad-tree structure. Subsequently, the leaf nodes of the quad-tree structure can be further partitioned by a multitype tree structure. For example, as shown in FIG. 11, a multitype tree structure may include four partition types in a schematic manner.

[0143] The four splitting types may include vertical binary splitting (SPLIT_BT_VER), horizontal binary splitting (SPLIT_BT_HOR), vertical ternary splitting (SPLIT_TT_VER), and horizontal ternary splitting (SPLIT_TT_HOR). Leaf nodes of a multitype tree structure may be called CUs. These CUs can be used for prediction and transformation procedures. In this document, CUs, PUs, and TUs generally have the same block size. However, if the maximum supported transform length is smaller than the width or height of the color component of the CU, the CU and TU may have different block sizes.

[0144] FIG. 12 shows splitting flags within a quad tree of a multi-type tree coding structure according to embodiments.

[0145] FIG. 12 exemplarily illustrates the signaling mechanism of partition splitting information in a quadtree with nested multi-type tree structure.

[0146] Here, the CTU is treated as the root of the quadtree and is initially partitioned into a quadtree structure. Each quadtree leaf node can subsequently be further partitioned into a multitype tree structure. In the multitype tree structure, a first flag (e.g., mtt_split_cu_flag) is signaled to indicate whether the node is further partitioned. If the node is further partitioned, a second flag (e.g., mtt_split_cu_vertical_flag) may be signaled to indicate the splitting direction. Subsequently, a third flag (e.g., mtt_split_cu_binary_flag) may be signaled to indicate whether the splitting type is binary or binary. For example, based on mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the multi-type tree splitting mode (MttSplitMode) of CU can be derived as shown in Table 1 (MttSplitMode derviation based on multi-type tree syntax elements).

[0147] [Table 1]

[0148]

[0149] FIG. 13 shows an example of a quad tree of a multi-type tree coding block structure according to embodiments.

[0150] FIG. 13 exemplarily illustrates a CTU being divided into multiple CUs based on a quadtree and nested multi-type tree structure.

[0151] Here, bold block edges represent quadtree partitioning, and the remaining edges represent multitype tree partitioning. Quadtree partitioning involving a multitype tree can provide a content-adapted coding tree structure. A CU can correspond to a coding block (CB). Alternatively, a CU may include a coding block of luminance samples and two coding blocks of corresponding chroma samples. The size of a CU may be as large as a CTU, or it may be 4x4 in luminance sample units. For example, in the case of a 4:2:0 color format (or chroma format), the maximum chroma CB size may be 64x64 and the minimum chroma CB size may be 2x2.

[0152] For example, in this document, the maximum allowable luma TB size may be 64x64 and the maximum allowable chroma TB size may be 32x32. If the width or height of a CB partitioned according to the tree structure is greater than the maximum conversion width or height, the CB may be automatically (or implicitly) partitioned until the horizontal and vertical TB size limits are satisfied.

[0153] Meanwhile, for a quadtree coding tree scheme involving a multitype tree, the following parameters can be defined and identified as SPS syntax elements.

[0154] CTU size: Size of the root node of a 4th-order tree

[0155] MinQTSize: Minimum allowed 4th-order tree leaf node size

[0156] MaxBtSize: Maximum allowed binary tree root node size

[0157] MaxTtSize: Maximum allowed ternary tree root node size

[0158] MaxMttDepth: The maximum allowed hierarchy depth of a multi-type tree splitting at a 4th-order tree leaf.

[0159] MinBtSize: Minimum allowed binary tree leaf node size

[0160] MinTtSize: Minimum allowed tertiary tree leaf node size

[0161] As an example of a quadtree coding tree structure involving a multitype tree, the CTU size can be set to 64x64 blocks of 128x128 luminance samples and two corresponding chroma samples (in the 4:2:0 chroma format). In this case, MinOTSize can be set to 16x16, MaxBtSize to 128x128, MaxTtSize to 64x64, MinBtSize and MinTtSize (for both width and height) to 4x4, and MaxMttDepth to 4. Quadtree partitioning can be applied to the CTU to create quadtree leaf nodes. Quadtree leaf nodes can be called leaf QT nodes. Quadtree leaf nodes can have sizes ranging from 16x16 (i.e., the MinOTSize) to 128x128 (i.e., the CTU size). If the leaf QT node is 128x128, it may not be further split into a binary tree / binary tree. This is because even if it were split in this case, it would exceed MaxBtsize and MaxTtsize (i.e., 64x64). Otherwise, the leaf QT node may be further split into a multitype tree. Therefore, the leaf QT node is the root node of the multitype tree, and the leaf QT node can have a multitype tree depth (mttDepth) value of 0. If the multitype tree depth reaches MaxMttdepth (e.g., 4), further splitting may not be considered. If the width of the multitype tree node is equal to MinBtSize and is less than or equal to 2xMinTtSize, further horizontal splitting may not be considered. If the height of a multitype tree node is equal to MinBtSize and less than or equal to 2xMinTtSize, no further vertical splitting may be considered.

[0162] FIG. 14 shows the prohibition of TT (Ternary Tree) division for coding blocks according to embodiments.

[0163] In order to allow 64x64 luminance block and 32x32 chroma pipeline designs in a hardware decoder, TT splitting may be forbidden in certain cases. For example, if the width or height of the luminance coding block is greater than 64, TT splitting may be forbidden, as shown in FIG. 14. Also, for example, if the width or height of the chroma coding block is greater than 32, TT splitting may be forbidden.

[0164] In this document, the coding tree scheme may support Luma and Chroma (component) blocks having separate block tree structures. If Luma and Chroma blocks within a single CTU have the same block tree structure, it may be denoted as SINGLE_TREE. If Luma and Chroma blocks within a single CTU have separate block tree structures, it may be denoted as DUAL_TREE. In this case, the block tree type for the Luma component may be called DUAL_TREE_LUMA, and the block tree type for the Chroma component may be called DUAL_TREE_CHROMA. For P and B slice / tile groups, Luma and Chroma CTBs within a single CTU may be restricted to having the same coding tree structure. However, for I slice / tile groups, Luma and Chroma blocks may have separate block tree structures. If individual block tree mode is applied, the Luma CTB may be divided into CUs based on a specific coding tree structure, and the Chroma CTB may be divided into Chroma CUs based on a different coding tree structure. This may mean that CUs within an I slice / tile group may consist of coding blocks of the Luma component or coding blocks of two Chroma components, and CUs within a P or B slice / tile group may consist of blocks of three color components. In this document, a slice may be referred to as a tile / tile group, and a tile / tile group may be referred to as a slice.

[0165] In the aforementioned "Partitioning of the CTUs using a tree structure," a quadtree coding tree structure involving a multitype tree was described, but the structure in which the CU is partitioned is not limited to this. For example, the BT structure and the TT structure can be interpreted as concepts included in the Multiple Partitioning Tree (MPT) structure, and the CU can be interpreted as being partitioned through the QT structure and the MPT structure. In an example where the CU is partitioned through the QT structure and the MPT structure, the partitioning structure can be determined by signaling a syntax element (e.g., MPT_split_type) containing information regarding how many blocks the leaf node of the QT structure is partitioned into, and a syntax element (e.g., MPT_split_mode) containing information regarding whether the leaf node of the QT structure is partitioned vertically or horizontally.

[0166] In another example, the CU may be divided in a way different from the QT structure, BT structure, or TT structure. That is, unlike when a lower-depth CU is divided into 1 / 4 the size of an upper-depth CU according to the QT structure, or a lower-depth CU is divided into 1 / 2 the size of an upper-depth CU according to the BT structure, or a lower-depth CU is divided into 1 / 4 or 1 / 2 the size of an upper-depth CU according to the TT structure, the lower-depth CU may, in some cases, be divided into 1 / 5, 1 / 3, 3 / 8, 3 / 5, 2 / 3, or 5 / 8 the size of an upper-depth CU, and the method of dividing the CU is not limited thereto.

[0167] Transformation / Inverse Transformation:

[0168] As described above, the encoding device can derive residual blocks (residual samples) based on blocks (predicted samples) predicted through intra / inter / IBC prediction, etc., and can derive quantized transformation coefficients by applying transformation and quantization to the derived residual samples. Information regarding the quantized transformation coefficients (residual information) can be included in the residual coding syntax and output in the form of a bitstream after encoding. The decoding device can obtain information regarding the quantized transformation coefficients (residual information) from the bitstream and derive the quantized transformation coefficients by decoding. The decoding device can derive residual samples by undergoing inverse quantization / inverse transformation based on the quantized transformation coefficients. As described above, at least one of quantization / inverse quantization and / or transformation / inverse transformation may be omitted. When the transform / inverse transform is omitted, the transform coefficients may be called coefficients or residual coefficients, or they may still be called transform coefficients for the sake of consistency in representation. Whether the transform / inverse transform is omitted can be signaled based on transform_skip_flag.

[0169] Transformation / inverse transformation can be performed based on transformation kernel(s). For example, according to this document, a multiple transform selection (MTS) scheme may be applied. In this case, some of the sets of multiple transformation kernels may be selected and applied to the current block. Transformation kernels may be referred to by various terms, such as transformation matrix or transformation type. For example, a set of transformation kernels may represent a combination of vertical transformation kernels and horizontal transformation kernels.

[0170] For example, MTS index information (or tu_mts_idx syntax elements) may be generated / encoded in an encoding device and signaled to a decoding device to indicate one of the sets of transformation kernels. For example, the sets of transformation kernels based on the values ​​of the MTS index information may be derived as shown in Table 2 (Specification of trTypeHor and trTypeVer depending on tu_mts_idx[ x ][ y ]), Table 3 (Specification of trTypeHor and trTypeVer depending on cu_sbt_horizontal_flag and cu_sbt_pos_flag), and / or Table 4 (Specification of trTypeHor and trTypeVer depending on predModeIntra).

[0171] [Table 2]

[0172]

[0173] The set of transformation kernels may be determined, for example, based on cu_sbt_horizontal_flag and cu__sbt_pos_flag.

[0174] If cu_sbt_horizontal_flag is 1, it indicates that the current coding unit is divided horizontally into two transformation units. If cu_sbt_horizontal_flag[ x0 ][ y0 ] is 0, it indicates that the current coding unit is divided vertically into two transformation units. If cu_sbt_pos_flag is 1, it indicates that tu_cbf_luma, tu_cbf_cb, and tu_cbf_cr of the first transformation unit of the current coding unit are not in the bitstream. If cu_sbt_pos_flag is 0, it indicates that tu_cbf_luma, tu_cbf_cb, and tu_cbf_cr of the second transformation unit of the current coding unit are not in the bitstream.

[0175] [Table 3]

[0176]

[0177] The set of transformation kernels may be determined, for example, based on the intra prediction mode for the current block.

[0178] [Table 4]

[0179]

[0180] In the tables above, trTypeHor can represent a horizontal direction conversion kernel, and trTypeVer can represent a vertical direction conversion kernel. Here, a trTypeHor / trTypeVer value of 0 can represent DCT2, a trTypeHor / trTypeVer value of 1 can represent DST7, and a trTypeHor / trTypeVer value of 2 can represent DCT8. However, this is merely an example, and by convention, other values ​​may be mapped to different DCTs / DSTs.

[0181] The following Table 5 (Transform basis functions of DCT-II / VIII and DSTVII for N-point input) illustrates exemplary basis functions for the aforementioned DCT2, DCT8, and DST7.

[0182] [Table 5]

[0183]

[0184] FIG. 15 shows the transform and inverse transform according to the embodiments.

[0185] In this document, the MTS-based transformation is applied as a primary transform, and a secondary transform may be applied. The secondary transform may be applied only to the coefficients in the upper-left wxh region of the coefficient block to which the primary transform is applied, and may be called the Reduced Secondary Transform (RST). For example, w and / or h may be 4 or 8. In the transformation, the primary and secondary transforms may be applied sequentially to the secondary block, and in the inverse transformation, the inverse secondary transform and the inverse primary transform may be applied sequentially to the transform coefficients. The secondary transform (RST transform) may be called the low frequency coefficients transform (LFCT) or low frequency non-separable transform (LFNST). The inverse secondary transform may be called the inverse LFCT or inverse LFNST.

[0186] FIG. 16 shows a Low-Frequency Non-Separable Transform (LFNST) according to embodiments.

[0187] The LFNST (Low Frequency Inseparable Transform), also known as the Reduced Secondary Transform, is applied between the forward first-order transform and quantization (encoder side), and between the inverse quantization and the inverse first-order transform (decoder side), as shown in FIG. 16. In the LFNST, a 4x4 inseparable transform or an 8x8 inseparable transform is applied depending on the block size. For example, a 4x4 LFNST is applied to small blocks (i.e., minimum value (width, height) < 8), and an 8x8 LFNST is applied to large blocks (i.e., minimum value (width, height) > 4).

[0188] The application of the inseparable transformation used in LFNST is explained as follows, using the input as an example. To apply a 4x4 LFNST, the 4x4 input block X is represented as a vector as follows.

[0189]

[0190]

[0191] Inseparable transformations are as follows: It is calculated as. Here represents the transformation coefficient vector, and T is a 16x16 transformation matrix. 16x1 coefficients The vector is then reconstructed into a 4x4 block using the scan order (horizontal, vertical, or diagonal) of the corresponding block. Factors with smaller indices are placed at smaller scan indices in the 4x4 factor block.

[0192] Transform / inverse transformation can be performed in units of CU or TU. That is, transformation / inverse transformation can be applied to residual samples within a CU or residual samples within a TU. The CU size and the TU size may be the same, or multiple TUs may exist within the CU area. Meanwhile, the term CU size generally refers to the luminance component (sample) CB size. The term TU size generally refers to the luminance component (sample) TB size. The chroma component (sample) CB or TB size can be derived based on the luminance component (sample) CB or TB size according to the component ratio based on the color format (chroma format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.). The TU size can be derived based on maxTbSize. For example, if the CU size is greater than maxTbSize, multiple TU(TB) of maxTbSize are derived from the CU, and conversion / inverse conversion can be performed in TU(TB) units. maxTbSize can be considered for determining whether to apply various intra-prediction types, such as ISP. Information regarding maxTbSize may be determined in advance, or it may be generated and encoded by an encoding device and signaled to a decoding device.

[0193] Quantization / Dequantization:

[0194] As described above, the quantization unit of the encoding device can derive quantized conversion coefficients by applying quantization to conversion coefficients, and the inverse quantization unit of the encoding device or the inverse quantization unit of the decoding device can derive conversion coefficients by applying inverse quantization to quantized conversion coefficients.

[0195] In general, in video / image coding, the quantization rate can be varied, and compression can be adjusted using the varied quantization rate. From an implementation perspective, considering complexity, quantization parameters (QP) can be used instead of directly using the quantization rate. For example, quantization parameters can be integer values ​​from 0 to 63, and each quantization parameter value can correspond to an actual quantization rate. The quantization parameter (QPY) for the luminance component (luma sample) and the quantization parameter (QPC) for the chroma component (chroma sample) can be set differently.

[0196] The quantization process takes a transform coefficient (C) as input and divides it by a quantization rate (Qstep) to obtain a quantized transform coefficient (C'). In this case, considering computational complexity, the quantization rate can be multiplied by a scale to form an integer, and a shift operation can be performed by an amount corresponding to the scale value. A quantization scale can be derived based on the product of the quantization rate and the scale value. In other words, the quantization scale can be derived according to QP. Alternatively, the quantization scale can be applied to the transform coefficient (C) to derive the quantized transform coefficient (C').

[0197] The inverse quantization process is the reverse of the quantization process; by multiplying the quantized transformation coefficients (C') by the quantization rate (Qstep), the reconstructed transformation coefficients (C'') can be obtained based on this. In this case, a level scale can be derived depending on the quantization parameters, and the reconstructed transformation coefficients (C'') can be derived by applying this level scale to the quantized transformation coefficients (C''). The reconstructed transformation coefficients (C'') may differ slightly from the original transformation coefficients (C) due to losses during the transformation and / or quantization processes. Therefore, the encoding device performs inverse quantization in the same manner as the decoding device.

[0198] Meanwhile, adaptive frequency-weighted quantization technology, which adjusts the quantization intensity according to frequency, may be applied. Adaptive frequency-weighted quantization is a method of applying different quantization intensities for each frequency. Adaptive frequency-weighted quantization can apply different quantization intensities for each frequency by utilizing a predefined quantization scaling matrix. That is, the aforementioned quantization / de-quantization process can be performed based further on the quantization scaling matrix. For example, different quantization scaling matrices may be used depending on whether the prediction mode applied to the current block to generate the size of the current block and / or the residual signal of the current block is inter-prediction or intra-prediction. The quantization scaling matrix may be referred to as a quantization matrix or a scaling matrix. The quantization scaling matrix may be predefined. Additionally, for frequency-adaptive scaling, frequency-specific quantization scale information regarding the quantization scaling matrix may be configured / encoded in the encoding device and signaled to the decoding device. Frequency-specific quantization scale information can be referred to as quantization scaling information. Frequency-specific quantization scale information may include scaling list data (scaling_list_data). A (modified) quantization scaling matrix can be derived based on the scaling list data. Additionally, frequency-specific quantization scale information may include present flag information indicating the existence of scaling list data. Alternatively, it may further include information indicating whether scaling list data is modified at a lower level (e.g., PPS or tile group header, etc.) when scaling list data is signaled at a higher level (e.g., SPS).

[0199] Entropy Coding:

[0200] As described above in the description of FIG. 2, part or all of the video / image information may be entropied by the entropy encoding unit (240), and part or all of the video / image information described above in the description of FIG. 3 may be entropied by the entropy decoding unit (310). In this case, the video / image information may be encoded / decoded in units of syntax elements. In this document, the term "information is encoded / decoded" may include encoding / decoding by the method described in this paragraph.

[0201] FIG. 17 shows CABAC (Context Adaptive Binary Arithmetic Coding) encoding according to embodiments.

[0202] Figure 17 shows a block diagram of a CABAC for encoding a single syntax element. The encoding process of the CABAC first converts the input signal into a binary value through binarization if the input signal is a syntax element rather than a binary value. If the input signal is already a binary value, it is bypassed without undergoing binarization. Here, each binary digit 0 or 1 constituting the binary value is called a bin. For example, if the binary string after binarization (bin string) is 110, each of 1, 1, and 0 is called a bin. The bin(s) for a single syntax element can represent the value of the corresponding syntax element.

[0203] Binary bins are input into a regular coding engine or a bypass coding engine. The regular coding engine assigns a context model reflecting probability values ​​to the corresponding bin and encodes the bin based on the assigned context model. The regular coding engine can update the probability model for each bin after performing coding for it. Bins coded in this way are called context-coded bins. The bypass coding engine omits the procedure of estimating probabilities for input bins and the procedure of updating the probability model applied to the bin after coding. Instead of assigning context, it improves coding speed by coding input bins using a uniform probability distribution (e.g., 50:50). Bins coded in this way are called bypass bins. The context model can be assigned and updated per context-coded (regularly coded) bin, and the context model can be indicated based on ctxidx or ctxInc. ctxidx can be derived based on ctxInc. Specifically, for example, the context index (ctxidx) pointing to the context model for each normally coded bean can be derived as the sum of the context index increment (ctxInc) and the context index offset (ctxIdxOffset). Here, ctxInc can be derived differently for each bean. ctxIdxOffset can be represented as the lowest value of ctxIdx. The lowest value of ctxIdx can be called the initial value (initValue) of ctxIdx. ctxIdxOffset is a value generally used to distinguish context models for other syntax elements, and the context model for a single syntax element can be distinguished / derived based on ctxinc.

[0204] In the entropy encoding procedure, it is determined whether to perform encoding through a regular coding engine or a bypass coding engine, and the coding path can be switched. Entropy decoding performs the same process as entropy encoding in reverse order.

[0205] FIG. 18 illustrates an entropy encoding method according to embodiments.

[0206] The entropy coding described in Fig. 17 can be performed, for example, as shown in Fig. 18.

[0207] Referring to FIG. 18, an encoding device (entropy encoding unit) performs an entropy coding procedure regarding image / video information. The image / video information may include partitioning-related information, prediction-related information (e.g., inter / intra prediction distinction information, intra prediction mode information, inter prediction mode information, etc.), residual information, in-loop filtering-related information, etc., or may include various syntax elements related thereto. Entropy coding may be performed on a syntax element basis. S600 to S610 may be performed by the entropy encoding unit (240) of the encoding device of FIG. 2 described above.

[0208] The encoding device performs binarization on the target syntax element (S600). Here, the binarization may be based on various binarization methods, such as the Truncated Rice binarization process and the Fixed-length binarization process, and the binarization method for the target syntax element may be predefined. The binarization procedure may be performed by the binarization unit (242) within the entropy encoding unit (240).

[0209] The encoding device performs entropy encoding on the target syntax element (S610). The encoding device may encode the empty string of the target syntax element based on a regular coding-based (context-based) or bypass coding-based method, such as CABAC (context-adaptive arithmetic coding) or CAVLC (context-adaptive variable length coding), and the output may be included in a bitstream. The entropy encoding procedure may be performed by an entropy encoding processing unit (243) within the entropy encoding unit (240). As previously mentioned, the bitstream may be transmitted to a decoding device via a (digital) storage medium or a network.

[0210] FIG. 19 illustrates an entropy decoding method according to embodiments.

[0211] As shown in FIG. 19, a decoding device (entropy decoding unit) can decode encoded image / video information. The image / video information may include partitioning-related information, prediction-related information (e.g., inter / intra prediction distinction information, intra prediction mode information, inter prediction mode information, etc.), residual information, in-loop filtering-related information, etc., or may include various syntax elements related thereto. Entropy coding can be performed on a syntax element basis. S700 to S710 can be performed by the entropy decoding unit (310) of the decoding device of FIG. 3 described above.

[0212] The decoding device performs binarization on the target syntax element (S700). Here, the binarization may be based on various binarization methods, such as the Truncated Rice binarization process and the Fixed-length binarization process, and the binarization method for the target syntax element may be predefined. The decoding device may derive available empty strings (empty string candidates) for the available values ​​of the target syntax element through the binarization procedure. The binarization procedure may be performed by the binarization unit (312) within the entropy decoding unit (310).

[0213] The decoding device performs entropy decoding for the target syntax element (S710). The decoding device sequentially decodes and parses each bin for the target syntax element from the input bit(s) in the bitstream, and compares the derived bin string with the available bin strings for the corresponding syntax element. If the derived bin string is equal to one of the available bin strings, the value corresponding to the bin string is derived as the value of the corresponding syntax element. If not, the next bit in the bitstream is parsed further, and the procedure described above is performed again. Through this process, information (specific syntax element) can be signaled using variable-length bits without using start bits or end bits for specific information within the bitstream. Through this, relatively fewer bits can be allocated to low values, and overall coding efficiency can be increased.

[0214] The decoding device can decode each bin within a bin string from a bitstream in a context-based or bypass-based manner based on an entropy coding technique such as CABAC or CAVLC. The entropy decoding procedure can be performed by an entropy decoding processing unit (313) within the entropy decoding unit (310). As described above, the bitstream may contain various information for image / video decoding. As previously stated, the bitstream may be transmitted to the decoding device via a (digital) storage medium or a network.

[0215] In this document, a table containing syntax elements (syntax table) may be used to represent the signaling of information from an encoding device to a decoding device. The order of the syntax elements in the table containing syntax elements used in this document may represent the parsing order of the syntax elements from the bitstream. The encoding device may configure and encode the syntax table so that the syntax elements can be parsed by the decoding device in the parsing order, and the decoding device may obtain the values ​​of the syntax elements by parsing and decoding the syntax elements of the corresponding syntax table from the bitstream according to the parsing order.

[0216] FIG. 20 illustrates a picture decoding method according to embodiments.

[0217] General Video / Video Coding Procedures:

[0218] In video coding, the pictures constituting the video can be encoded / decoded according to a series of decoding orders. The picture order corresponding to the output order of the decoded pictures can be set differently from the decoding order, and based on this, not only forward prediction but also reverse prediction can be performed during inter-prediction.

[0219] FIG. 20 illustrates an example of a schematic picture decoding procedure to which the embodiment(s) of the present document are applicable. In FIG. 20, S900 may be performed in the entropy decoding unit (310) of the decoding device described in FIG. 3, S910 may be performed in the prediction unit (330), S920 may be performed in the residual processing unit (320), S930 may be performed in the addition unit (340), and S940 may be performed in the filtering unit (350). S900 may include the information decoding procedure described in the present document, S910 may include the inter / intra prediction procedure described in the present document, S920 may include the residual processing procedure described in the present document, S930 may include the block / picture restoration procedure described in the present document, and S940 may include the in-loop filtering procedure described in the present document.

[0220] As shown in FIG. 20, the picture decoding procedure may include, schematically as described in FIG. 3, a procedure for obtaining image / video information (through decoding) from a bitstream (S900), a picture restoration procedure (S910–S930), and an in-loop filtering procedure for the restored picture (S940). The picture restoration procedure may be performed based on prediction samples and residual samples obtained through the inter / intra prediction (S910) and residual processing (S920, inverse quantization and inverse transformation of quantized transformation coefficients) described in this document. A modified restored picture may be generated through an in-loop filtering procedure for the restored picture generated through the picture restoration procedure, and the modified restored picture may be output as a decoded picture and may also be stored in the decoded picture buffer or memory (360) of the decoding device and used as a reference picture in the inter prediction procedure when decoding the picture thereafter. In some cases, the in-loop filtering procedure may be omitted, in which case the restored picture may be output as a decoded picture and may also be stored in the decoded picture buffer or memory (360) of the decoding device and used as a reference picture in the inter-prediction procedure during subsequent decoding of the picture. The in-loop filtering procedure (S940) may include a deblocking filtering procedure, a sample adaptive offset (SAO) procedure, an adaptive loop filter (ALF) procedure, and / or a bilateral filter procedure, as described above, and some or all of these may be omitted. Additionally, one or some of the deblocking filtering procedure, the sample adaptive offset (SAO) procedure, the adaptive loop filter (ALF) procedure, and the bilateral filter procedure may be applied sequentially, or all of them may be applied sequentially. For example, the SAO procedure may be performed after the deblocking filtering procedure is applied to the restored picture.Alternatively, for example, the ALF procedure may be performed after a deblocking filtering procedure has been applied to the restored picture. This can be done in the same way on the encoding device.

[0221] FIG. 21 illustrates a picture encoding method according to embodiments.

[0222] FIG. 21 illustrates an example of a schematic picture encoding procedure to which the embodiment(s) of the present document are applicable. In FIG. 21, S800 may be performed in the prediction unit (220) of the encoding device described above in FIG. 2, S810 may be performed in the residual processing unit (230), and S820 may be performed in the entropy encoding unit (240). S800 may include the inter / intra prediction procedure described in the present document, S810 may include the residual processing procedure described in the present document, and S820 may include the information encoding procedure described in the present document.

[0223] As shown in FIG. 21, the picture encoding procedure may include not only a procedure for encoding information for picture restoration (e.g., prediction information, residual information, partitioning information, etc.) in a general manner as described in FIG. 02 and outputting it in the form of a bitstream, but also a procedure for generating a restored picture for the current picture and a procedure for applying in-loop filtering to the restored picture (optional). The encoding device may derive (modified) residual samples from quantized transform coefficients through the inverse quantization unit (234) and the inverse transform unit (235), and may generate a restored picture based on the prediction samples and (modified) residual samples which are the outputs of S800. The restored picture thus generated may be identical to the restored picture generated by the decoding device described above. A modified restored picture can be generated through an in-loop filtering procedure for the restored picture, which can be stored in a decoded picture buffer or memory (270), and, as in the case of a decoding device, can be used as a reference picture in an inter-prediction procedure during the encoding of the picture thereafter. As described above, in some cases, part or all of the in-loop filtering procedure may be omitted. When the in-loop filtering procedure is performed, filtering-related information (parameters) can be encoded in the entropy encoding unit (240) and output in the form of a bitstream, and the decoding device can perform the in-loop filtering procedure in the same way as the encoding device based on the filtering-related information.

[0224] Through this in-loop filtering procedure, noise generated during video coding, such as blocking and ringing artifacts, can be reduced, and subjective and objective visual quality can be enhanced. Furthermore, by performing the in-loop filtering procedure in both the encoding and decoding devices, they can derive identical prediction results, increase the reliability of picture coding, and reduce the amount of data that must be transmitted for picture coding.

[0225] As described above, the picture restoration procedure can be performed not only in the decoding device but also in the encoding device. Restoration blocks can be generated based on intra-prediction / inter-prediction on a block-by-block basis, and a restored picture containing the restoration blocks can be generated. If the current picture / slice / tile group is the I picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based solely on intra-prediction. Meanwhile, if the current picture / slice / tile group is the P or B picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based on intra-prediction or inter-prediction. In this case, inter-prediction may be applied to some blocks within the current picture / slice / tile group, and intra-prediction may be applied to some remaining blocks. The color components of the picture may include luminance components and chroma components, and unless explicitly limited in this document, the methods and embodiments proposed in this document may be applied to luminance components and chroma components.

[0226] Examples of coding hierarchy and structure:

[0227] The coded video / image according to this document can be processed according to, for example, the coding layers and structures described below.

[0228] FIG. 22 shows a hierarchical structure for a coded image according to embodiments.

[0229] FIG. 22 is a diagram illustrating the hierarchical structure of a coded image.

[0230] The coded video is divided into a video coding layer (VCL) that handles the decoding processing of the video and the video itself, a subsystem that transmits and stores the encoded information, and a network abstraction layer (NAL) that exists between the VCL and the subsystem and is responsible for network adaptation functions.

[0231] In VCL, VCL data containing compressed image data (slice data) can be generated, or parameter sets containing information such as Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS), or SEI (Supplemental Enhancement Information) messages that are additionally required in the decoding process of the image can be generated.

[0232] In NAL, a NAL unit can be created by adding header information (NAL unit header) to the Raw Byte Sequence Payload (RBSP) generated in VCL. In this case, the RBSP refers to slice data, parameter sets, SEI messages, etc. generated in VCL. The NAL unit header may include NAL unit type information specified according to the RBSP data included in the NAL unit.

[0233] NAL units can be classified into VCL NAL units and Non-VCL NAL units depending on the RBSP generated in VCL. A VCL NAL unit may refer to a NAL unit containing information about an image (slice data), and a Non-VCL NAL unit may refer to a NAL unit containing information necessary to decode an image (parameter set or SEI message).

[0234] The aforementioned VCL NAL unit and Non-VCL NAL unit can be transmitted over a network by attaching header information according to the data specifications of the underlying system. For example, the NAL unit can be transformed into a data format of a specified specification, such as H.266 / VVC file format, RTP (Real-time Transport Protocol), TS (Transport Stream), etc., and transmitted over various networks.

[0235] The NAL unit type can be determined according to the RBSP data structure included in the NAL unit, and information about this NAL unit type can be stored in the NAL unit header and signaled.

[0236] For example, NAL units can be broadly classified into VCL NAL unit types and Non-VCL NAL unit types depending on whether they contain information about the image (slice data). VCL NAL unit types can be classified according to the properties and types of the picture included in the VCL NAL unit, while Non-VCL NAL unit types can be classified according to the types of parameter sets.

[0237] The following is an example of a NAL unit type specified based on the type of parameter set included by the Non-VCL NAL unit type: APS (Adaptation Parameter Set) NAL unit: A type for a NAL unit containing APS. DPS (Decoding Parameter Set) NAL unit: A type for a NAL unit containing DPS. VPS (Video Parameter Set) NAL unit: A type for a NAL unit containing VPS. SPS (Sequence Parameter Set) NAL unit: A type for a NAL unit containing SPS. PPS (Picture Parameter Set) NAL unit: A type for a NAL unit containing PPS.

[0238] The above-described NAL unit types have syntax information for the NAL unit type, and the syntax information can be stored in the NAL unit header and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit types can be specified by the nal_unit_type value.

[0239] A slice header (slice header syntax) may include information / parameters that can be applied commonly to slices. An APS (APS syntax) or a PPS (PPS syntax) may include information / parameters that can be applied commonly to one or more slices or pictures. An SPS (SPS syntax) may include information / parameters that can be applied commonly to one or more sequences. A VPS (VPS syntax) may include information / parameters that can be applied commonly to multiple layers. A DPS (DPS syntax) may include information / parameters that can be applied commonly across the video. A DPS may include information / parameters related to the concatenation of a CVS (coded video sequence). In this document, High-level syntax (HLS) may include at least one of an APS syntax, a PPS syntax, an SPS syntax, a VPS syntax, a DPS syntax, and a slice header syntax.

[0240] In this document, the image / video information that is encoded from an encoding device to a decoding device and signaled in the form of a bitstream includes not only information related to partitioning within a picture, intra / inter prediction information, residual information, in-loop filtering information, etc., but may also include information included in a slice header, information included in an APS, information included in the PPS, information included in an SPS, and / or information included in a VPS.

[0241] Coding descriptors:

[0242] The following descriptors represent the parsing process for each syntax element: ae(v): Context-adaptive arithmetic entropy-coded syntax element. b(8): A byte containing a bit string of arbitrary patterns (8 bits). The parsing process for this descriptor is specified by the return value of the read_bits(8) function. f(n): A fixed-pattern bit string of n bits written from left to right with the left bit coming first. The parsing process for this descriptor is specified by the return value of the read_bits(n) function. i(n): A signed integer using n bits. If n is "v" in the syntax table, the number of bits depends on the values ​​of other syntax elements. The parsing process for this descriptor is specified by the return value of the read_bits(n) function and is interpreted as a two's complement integer representation with the most significant bit written first. se(v): A signed integer zero-ordered Exp-Golomb-coded syntax element, with the left bit coming first. The parsing process for this descriptor is specified by the order of k being zero. st(v): A null-terminated string encoded in Universal Coded Character Set (UCS) Transfer Format-8 (UTF-8) characters as specified in ISO / IEC 10646. The parsing process is as follows: st(v) moves the bitstream pointer (stringLength + 1) * 8 bit positions starting from the current position in the bitstream's byte alignment position to the next byte alignment byte, such as 0x00 (excluding that byte), where stringLength is equal to the number of bytes returned. The st(v) syntax descriptor is used in this specification only when the current position in the bitstream is the byte alignment position. tu(v): A truncated unary operator using up to maxVal bits. maxVal is defined in the semantics of the symtax ​​element. u(n): An unsigned integer using n bits.In the syntax table, if n is "v", the number of bits depends on the values ​​of other syntax elements. The parsing process of this descriptor is specified by the return value of the function read_bits(n) and is interpreted as a binary representation of an unsigned integer with the most significant bit written first. ue(v): An unsigned integer of a zero-order exponent Colomb-coded syntax element with the left bit written first. The parsing process of this descriptor is specified by setting the order of k to 0.

[0243] High-level syntax signaling and semantics are described below with reference to each figure.

[0244] FIGS. 23a, FIGS. 23b, FIGS. 23c, FIGS. 23d, and FIGS. 23e show picture header structures according to embodiments.

[0245] Picture header and slice header:

[0246] A coded picture may consist of one or more slices. Parameters describing the coded picture are passed within the picture header (PH), and parameters describing the slice are passed within the slice header. The PH is passed as its own NAL unit type. The SH is located at the beginning of the NAL unit containing the slice's payload (e.g., slice data). For details on the syntax and semantics of the PH and SH, refer to Section 7 of the VVC specification.

[0247] SEI Messages:

[0248] Neural-network post-filter SEI messages

[0249] General post-processing filtering processes using NNPFs

[0250]

[0251] The input to this process is a bitstream BitstreamToFilter. The output of this process is a list of NNPF output pictures, ListNnpfOutputPics.

[0252] First, BitstreamToFilter is decoded, and the CroppedDecodedPictures list is set as a list of decoded pictures cropped in the order of the BitstreamToFilter decoding output.

[0253] Second, a filtering process for one picture is in CroppedDecodedPictures and is repeatedly called in output order for each cropped decoded picture with one or more NNPFs enabled.

[0254] The order of the pictures in ListNnpfOutputPics is the output order.

[0255] There must be only one picture associated with a specific output time instance within ListNnpfOutputPics. If there are multiple NNPFs enabled for a specific picture in CroppedDecodedPictures and only one NNPF can be selected to apply (other NNPFs can also be selected), the above constraint applies regardless of which NNPF is applied to the specific picture.

[0256] Single Picture Filtering Process Using NNPF:

[0257] The filtering process is applied to each cropped decoded picture (referred to as the current picture) that belongs to CroppedDecodedPictures and has one or more NNPFs enabled.

[0258] When applying NNPF to the current picture, the filtered and / or interpolated picture is generated by NNPF by applying the NNPF process specified in the semantics of the NNPFFC SEI message to the current picture in a patch manner.

[0259] When applying NNPF to the current picture, the order of the picture generated by NNPF by applying the NNPF process is the same as the output order stored in the output tensor of NNPF.

[0260] If the applied NNPF is the last NNPF applied to the current picture, the picture generated by the NNPF and the picture output from the NNPF process are included in ListNnpfOutputPics, in the same order as when the picture is stored in the output tensor of the NNPF.

[0261] FIGS. 24a, FIGS. 24b, and FIGS. 24c show neural-network post-filter characteristics SEI message syntax according to embodiments.

[0262] The syntax of the NNPFC SEI message associated with the Neural-network post-filter characteristics SEI message (NNPFC) is as shown in FIGS. 24a, 24b, and 24c.

[0263] The NNPFC SEI message represents a neural network that can be used as a post-processing filter. The use of a specified neural network post-processing filter (NNPF) for a particular picture is indicated by the Neural Network Post-processing Filter Activation (NNPFA) SEI message.

[0264] To use this SEI message, the following variables must be defined.

[0265] Input picture width and height in Luma sample units (labeled as CroppedWidth and CroppedHeight, respectively).

[0266] CroppedYPic[idx], an array of luminance samples, and CroppedCbPic[idx] and CroppedCrPic[idx] (if any), chroma sample arrays of input pictures whose index idx used as input to NNPF is in the range from 0 to numInputPics - 1 (inclusive).

[0267] Bit depth for the luminance sample array of the input picture BitDepthY.

[0268] BitDepthC for the chroma sample array (if any) of the input picture.

[0269] Chroma format indicator displayed as ChromaFormatIdc.

[0270] If nnpfc_auxiliary_inp_idc is 1, the filtering strength control value array StrengthControlVal[idx] contains real numbers in the range of 0 to 1 (inclusive) for input pictures where index idx is in the range of 0 to numInputPics - 1 (inclusive).

[0271] The input picture with index 0 corresponds to the picture in which the NNPF defined in this NNPFC SEI message is activated by the NNPFA SEI message. Input pictures with index i in the range from 1 to numInputPics-1 take precedence over the input picture with index i-1 in the output order.

[0272] The SubWidthC and SubHeightC variables are derived from ChromaFormatIdc.

[0273] Two or more NNPFC SEI messages may exist for the same picture. If two or more NNPFC SEI messages with different nnpfc_id values ​​exist or are enabled for the same picture, the nnpfc_purpose and nnpfc_mode_idc values ​​of those messages may be the same or different.

[0274] nnpfc_purpose represents the purpose of the NNPF specified in Table 6 (Definition of nnpfc_purpose). Here, if (nnpfc_purpose & bitMask) is not 0, it indicates that the NNPF has a purpose associated with the bitMask value in Table 6. If nnpfc_purpose is greater than 0 and (nnpfc_purpose & bitMask) is 0, the purpose associated with the bitMask value cannot be applied to the NNPF. If nnpfc_purpose is 0, the NNPF can be used as determined by the application.

[0275] The value of nnpfc_purpose is in the range of 0 to 63 in bitstreams conforming to this version of this document. Values ​​for nnpfc_purpose from 64 to 65,535 (inclusive) are reserved for future use by ITU-T | ISO / IEC and do not exist in bitstreams conforming to this version of this document. Decoders conforming to this version of this document ignore NNPFC SEI messages with nnpfc_purpose in the range of 64 to 65,535.

[0276] [Table 6]

[0277]

[0278] The variables chromaUpsamplingFlag, resolutionResamplingFlag, pictureRateUpsamplingFlag, bitDepthUpsamplingFlag, and colourizationFlag, which respectively specify whether nnpfc_purpose includes chroma upsampling, resolution resampling, picture rate upsampling, bit depth upsampling, and colorization in the purpose of NNPF, are derived as follows.

[0279] chromaUpsamplingFlag = ((nnpfc_purpose & 0x02) > 0)? 1:0

[0280] resolutionResamplingFlag = ( ( nnpfc_purpose & 0x04 ) > 0 ) ? 1:0

[0281] pictureRateUpsamplingFlag = ((nnpfc_purpose & 0x08) > 0)? 1:0 (76)

[0282] bitDepthUpsamplingFlag = ( ( nnpfc_purpose & 0x10 ) > 0 ) ? 1:0

[0283] colourizationFlag = ( ( nnpfc_purpose & 0x20 ) > 0 ) ? 1:0

[0284] If the reserved value of nnpfc_purpose is used in the future by ITU-T | ISO / IEC, the syntax of this SEI message may be expanded with syntax elements depending on whether nnpfc_purpose is the same as that value.

[0285] If ChromaFormatIdc is 3, chromaUpsamplingFlag becomes 0.

[0286] If ChromaFormatIdc or chromaUpsamplingFlag is not 0, colourizationFlag becomes 0.

[0287] If the input picture with pictureRateUpsamplingFlag 1 and index 0 is associated with a frame packing array SEI message with fp_arrangement_type 5, then all input pictures are associated with a frame packing array SEI message with fp_arrangement_type 5 and the same fp_current_frame_is_frame0_flag value.

[0288] nnpfc_id contains an identification number that can be used to identify an NNPF. The nnpfc_id value is in the range from 0 to 232-2 (including 0). Among the nnpfc_id values, values ​​from 256 to 511 (including 0) and from 231 to 232-2 (including 0) are reserved for future use by ITU-T | ISO / IEC. Decoders complying with this version of this document ignore the NNPFC SEI message if they find an nnpfc_id value in the range from 256 to 511 (including 0) or from 231 to 232-2 (including 0).

[0289] If the NNPFC SEI message is the first NNPFC SEI message with a specific nnpfc_id value within the current CLVS in the decoding order, the following applies.

[0290] This SEI message represents the default NNPF.

[0291] This SEI message is applied to all subsequent decoded pictures of the current layer in output order, from the currently decoded picture to the end of the current CLVS.

[0292] If nnpfc_base_flag is 1, it indicates that the SEI message specifies the default NNPF. If nnpf_base_flag is 0, it indicates that the SEI message specifies an update based on the default NNPF.

[0293] The following constraints apply to the nnpfc_base_flag value.

[0294] If the NNPFC SEI message is the first NNPFC SEI message in decoding order with a specific nnpfc_id value within the current CLVS, the nnpfc_base_flag value is equal to 1.

[0295] If the NNPFC SEI message nnpfcB is not the first NNPFC SEI message in decoding order with a specific nnpfc_id value within the current CLVS, and the nnpfc_base_flag value is 1, the NNPFC SEI message is a repetition of the first NNPFC SEI message nnpfcA with the same nnpfc_id value in decoding order. That is, the payload content of nnpfcB is identical to the payload content of nnpfcA.

[0296] If nnpfc_base_flag is 0, the following applies.

[0297] This SEI message defines updates based on the previous default NNPF with the same nnpfc_id value in decoding order. Updates are not cumulative, and each update is applied to the default NNPF. The default NNPF is the NNPF specified in the first NNPFC SEI message in decoding order and has a specific nnpfc_id value within the current CLVS. The NNPF defined in this SEI message is obtained by applying the updates defined in this SEI message based on the default NNPF with the same nnpfc_id value.

[0298] This SEI message is about the currently decoded picture and all subsequent decoded pictures of the current layer (based on output order), up to the end of the current CLVS or up to the decoded picture following the currently decoded picture in output order within the current CLVS, and is associated with the subsequent NNPFC SEI message based on decoding order, where nnpfc_base_flag is 0 and there is a specific nnpfc_id value within the current CLVS (whichever is earlier).

[0299] If nnpfc_mode_idc is 0, it indicates that this SEI message contains an ISO / IEC 15938-17 bitstream specifying the default NNPF (if nnpfc_base_flag is 1) or is updated based on the default NNPF with the same nnpfc_id value (if nnpfc_base_flag is 0).

[0300] If nnpfc_base_flag is 1, and nnpfc_mode_idc is 1, it indicates that the base NNPF associated with the nnpfc_id value is a neural network in the format identified by a URI represented by nnpfc_uri and a tag URI identified by nnpfc_tag_uri.

[0301] When nnpfc_base_flag is 0, and nnpfc_mode_idc is 1, it indicates that updates to the base NNPF with the same nnpfc_id value are defined by the URI represented by nnpfc_uri and have a format identified by the tag URI nnpfc_tag_uri.

[0302] The nnpfc_mode_idc value is in the range from 0 to 1 in bitstreams compliant with this version of this document. Values ​​of nnpfc_mode_idc from 2 to 255 (inclusive) are reserved for future use by ITU-T | ISO / IEC and do not exist in bitstreams compliant with this version of this document. Decoders compliant with this version of this document ignore NNPFC SEI messages where nnpfc_mode_idc is in the range from 2 to 255. If the nnpfc_mode_idc value is greater than 255, it does not exist in bitstreams compliant with this version of this document and is not reserved for future use.

[0303] nnpfc_reserved_zero_bit_a is equal to 0 in bitstreams following this version of this document. Decoders ignore NNPFC SEI messages where nnpfc_reserved_zero_bit_a is not 0.

[0304] nnpfc_tag_uri contains a tag URI with the syntax and semantics specified in IETF RFC 4151 and identifies the format and related information of a neural network used as an update to a base NNPF or a base NNPF having the same nnpfc_id value specified in nnpfc_uri.

[0305] Using nnpfc_tag_uri allows you to uniquely identify the format of neural network data specified in nnrpf_uri without a central registry.

[0306] If nnpfc_tag_uri is "tag:iso.org,2023:15938-17", it indicates that the neural network data identified by nnpfc_uri complies with ISO / IEC 15938-17.

[0307] nnpfc_uri contains a URI with the syntax and semantics specified in IETF Internet Standard 66, and identifies the neural network used as the base NNPF or the neural network used as an update to the base NNPF with the same nnpfc_id value.

[0308] If nnpfc_property_present_flag is 1, it indicates that syntax elements related to filter purpose, input format, output format, and complexity exist. If nnpfc_property_present_flag is 0, it indicates that syntax elements related to filter purpose, input format, output format, and complexity do not exist.

[0309] If nnpfc_base_flag is 1, then nnpfc_property_present_flag also becomes 1.

[0310] If nnpfc_property_present_flag is 0, the value of all syntax elements that may exist only when nnpfc_property_present_flag is 1 is inferred to be the same as the corresponding syntax element of the NNPFC SEI message containing the underlying NNPF that this SEI message provides updates for.

[0311] If the NNPFC SEI message nnpfcCurr is not the first NNPFC SEI message in decoding order with a specific nnpfc_id value within the current CLVS, and does not overlap with the first NNPFC SEI message with that specific nnpfc_id (i.e., when the nnpfc_base_flag value is 0), or when the nnpfc_property_present_flag value is 1, the following constraints apply.

[0312] The nnpfc_purpose value of an NNPFC SEI message is the same as the nnpfc_purpose value of the first NNPFC SEI message in decoding order that has that specific nnpfc_id value within the current CLVS.

[0313] In NNPFC SEI messages, the syntax element values ​​after nnpfc_property_present_flag and before nnpfc_complexity_info_present_flag in the decoding order are the same as the corresponding syntax element values ​​of the first NNPFC SEI message with the corresponding nnpfc_id value within the current CLVS.

[0314] In the first NNPFC SEI message with the corresponding nnpfc_id value within the current CLVS (indicated as nnpfcBase below), nnpfc_complexity_info_present_flag must be 0 or both 1 in the decoding order, and all of the following apply.

[0315] The nnpfc_parameter_type_idc of nnpfcCurr is the same as the nnpfc_parameter_type_idc of nnpfcBase.

[0316] If nnpfc_log2_parameter_bit_length_minus3 of nnpfcCurr exists, it is less than or equal to nnpfc_log2_parameter_bit_length_minus3 of nnpfcBase.

[0317] If nnpfc_num_parameters_idc of nnpfcBase is 0, then nnpfc_num_parameters_idc of nnpfcCurr also becomes 0.

[0318] Otherwise (if nnpfc_num_parameters_idc of nnpfcBase is greater than 0), nnpfc_num_parameters_idc of nnpfcCurr is greater than 0 and less than or equal to nnpfc_num_parameters_idc of nnpfcBase.

[0319] If nnpfc_num_kmac_operations_idc of nnpfcBase is 0, then nnpfc_num_kmac_operations_idc of nnpfcCurr also becomes 0.

[0320] Otherwise (if nnpfc_num_kmac_operations_idc of nnpfcBase is greater than 0), nnpfc_num_kmac_operations_idc of nnpfcCurr is greater than 0 and less than or equal to nnpfc_num_kmac_operations_idc of nnpfcBase.

[0321] If the nnpfc_total_kilobyte_size of nnpfcBase is 0, the nnpfc_total_kilobyte_size of nnpfcCurr also becomes 0.

[0322] Otherwise (if nnpfc_total_kilobyte_size of nnpfcBase is greater than 0), nnpfc_total_kilobyte_size of nnpfcCurr is greater than 0 and less than or equal to nnpfc_total_kilobyte_size of nnpfcBase.

[0323] nnpfc_num_input_pics_minus1 + 1 represents the number of pictures used as input to the NNPF. The value of nnpfc_num_input_pics_minus1 ranges from 0 to 63. If pictureRateUpsamplingFlag is 1, the value of nnpfc_num_input_pics_minus1 is greater than 0.

[0324] The variable numInputPics, which specifies the number of pictures used as inputs for NNPF, is derived as follows.

[0325] numInputPics = nnpfc_num_input_pics_minus1 + 1 (77)

[0326] If nnpfc_input_pic_output_flag[ i ] is 1, NNPF indicates that the corresponding output picture is generated for the i-th input picture. If nnpfc_input_pic_output_flag[ i ] is 0, NNPF indicates that the corresponding output picture is not generated for the i-th input picture. If nnpfc_num_input_pics_minus1 is 0, nnpfc_input_pic_output_flag

[0000] is inferred to be 1. If pictureRateUpsamplingFlag is 0 and nnpfc_num_input_pics_minus1 is greater than 0, nnpfc_input_pic_output_flag[ i ] is equal to 1 for at least one value of i in the range from 0 to nnpfc_num_input_pics_minus1.

[0327] If nnpfc_absent_input_pic_zero_flag is 1, it indicates that NNPF should represent input pictures not in the bitstream as a sample array with a sample value of 0. If nnpfc_absent_input_pic_flag is 0, it indicates that NNPF should represent input pictures not in the bitstream as the input picture closest to the output order in the bitstream.

[0328] nnpfc_out_sub_c_flag indicates the values ​​of the outSubWidthC and outSubHeightC variables when chromaUpsamplingFlag is 1. When nnpfc_out_sub_c_flag is 1, it indicates that outSubWidthC is 1 and outSubHeightC is 1. When nnpfc_out_sub_c_flag is 0, it indicates that outSubWidthC is 2 and outSubHeightC is 1. If ChromaFormatIdc is 2 and nnpfc_out_sub_c_flag is present, the value of nnpfc_out_sub_c_flag becomes 1.

[0329] nnpfc_out_colour_format_idc specifies the color format of the NNPF output when colourizationFlag is 1, and consequently represents the values ​​of the outSubWidthC and outSubHeightC variables. If nnpfc_out_colour_format_idc is 1, it indicates that the NNPF output color format is 4:2:0 and both outSubWidthC and outSubHeightC are 2. If nnpfc_out_colour_format_idc is 2, it indicates that the NNPF output color format is 4:2:2 and outSubWidthC is 2 and outSubHeightC is 1. If nnpfc_out_colour_format_idc is 3, it indicates that the NNPF output color format is 4:4:4 and both outSubWidthC and outSubHeightC are 1. The value of nnpfc_out_colour_format_idc is not 0.

[0330] If both chromaUpsamplingFlag and colourizationFlag are 0, outSubWidthC and outSubHeightC are inferred as follows: SubWidthC and SubHeightC.

[0331] nnpfc_pic_width_num_minus1 + 1 and nnpfc_pic_width_denom_minus1 + 1 represent the numerator and denominator for the resampling ratio of the NNPF output picture width relative to CroppedWidth, respectively. The value of (nnpfc_pic_width_num_minus1 + 1) χ (nnpfc_pic_width_denom_minus1 + 1) is in the range of 1 χ 16 to 16. If nnpfc_pic_width_num_minus1 and nnpfc_pic_width_denom_minus1 are absent, the values ​​of nnpfc_pic_width_num_minus1 and nnpfc_pic_width_denom_minus1 are both inferred to be 0.

[0332] The variable nnpfcOutputPicWidth, which represents the width of the luminance sample array of the picture resulting from applying the NNPF identified by nnpfc_id to the input picture, is derived as follows.

[0333] nnpfcOutputPicWidth = Ceil(CroppedWidth * (78)

[0334] ( nnpfc_pic_width_num_minus1 + 1 ) χ ( nnpfc_pic_width_denom_minus1 + 1 ) )

[0335] For bitstream compatibility, the value of nnpfcOutputPicWidth % outSubWidthC is equal to 0.

[0336] nnpfc_pic_height_num_minus1 + 1 and nnpfc_pic_height_denom_minus1 + 1 represent the numerator and denominator for the resampling ratio of the NNPF output picture height relative to CroppedHeight, respectively. The values ​​of ( nnpfc_pic_height_num_minus1 + 1 ) χ ( nnpfc_pic_height_denom_minus1 + 1 ) range from 1 χ 16 to 16. If nnpfc_pic_height_num_minus1 and nnpfc_pic_height_denom_minus1 are absent, the values ​​of nnpfc_pic_height_num_minus1 and nnpfc_pic_height_denom_minus1 are both inferred to be 0.

[0337] The variable nnpfcOutputPicHeight, which represents the height of the luminance sample array of the picture resulting from applying the NNPF identified by nnpfc_id to the input picture, is derived as follows.

[0338] nnpfcOutputPicHeight = Ceil( CroppedHeight * (79)

[0339] ( nnpfc_pic_height_num_minus1 + 1 ) χ ( nnpfc_pic_height_denom_minus1 + 1 ) )

[0340] For bitstream compatibility, the value of nnpfcOutputPicHeight % outSubHeightC is equal to 0.

[0341] If nnpfc_pic_width_num_minus1, nnpfc_pic_width_denom_minus1, nnpfc_pic_height_num_minus1, nnpfc_pic_height_denom_minus1 exist, one or more of the following are true.

[0342] The value of nnpfcOutputPicWidth is not equal to CroppedWidth.

[0343] The value of nnpfcOutputPicHeight is not equal to CroppedHeight.

[0344] nnpfc_interpolated_pics[i] represents the number of interpolated pictures generated by NNPF between the i-th picture used as input to NNPF and the (i + 1)-th picture. The value of nnpfc_interpolated_pics[i] ranges from 0 to 63. The value of nnpfc_interpolated_pics[i] is greater than 0 for at least one of the i values ​​in the range from 0 to nnpfc_num_input_pics_minus1 - 1.

[0345] The variable NumInpPicsInOutputTensor, which specifies the number of pictures in the output tensor of NNPF that have the corresponding input picture, InpIdx[ idx ], which specifies the input picture index of the idx-th picture in the output tensor of NNPF that has the corresponding input picture, and numOutputPics, which specifies the total number of pictures in the output tensor of NNPF, are derived as follows.

[0346] for( i = 0, numOutputPics = 0; i < numInputPics; i++ )

[0347] if(nnpfc_input_pic_output_flag[i]) {

[0348] InpIdx[ numOutputPics ] = i

[0349] numOutputPics++

[0350] }

[0351] NumInpPicsInOutputTensor = numOutputPics

[0352] if(pictureRateUpsamplingFlag)

[0353] for( i = 0; i <= numInputPics - 2; i++ )

[0354] numOutputPics += nnpfc_interpolated_pics[ i ]

[0355] If nnpfc_component_last_flag is 1, it indicates that the last dimension of the input tensor inputTensor for the NNPF and the output tensor outputTensor generated by the NNPF are used for the current channel. If nnpfc_component_last_flag is 0, it indicates that the third dimension of the input tensor inputTensor for the NNPF and the output tensor outputTensor generated by the NNPF are used for the current channel.

[0356] The first dimension of the input and output tensors is used for the batch index, which is the approach used in some neural network frameworks. The formula in the semantics of this SEI message uses a batch size with a batch index of 0, but determining the batch size used as input for neural network inference depends on the post-processing implementation.

[0357] For example, when nnpfc_inp_order_idc is 3 and nnpfc_auxiliary_inp_idc is 1, the input tensor has a total of 7 channels, including 4 luminance matrices, 2 chroma matrices, and 1 auxiliary input matrix. In this case, the DeriveInputTensors() process derives the 7 channels of the input tensor one by one, and when a specific channel among these is processed, that channel is referred to as the current channel during the process.

[0358] nnpfc_inp_format_idc indicates how to convert the sample values ​​of the input picture into input values ​​for the NNPF. If nnpfc_inp_format_idc is 0, the input values ​​for the NNPF are real numbers, and the InpY() and InpC() functions are expressed as follows.

[0359] InpY(x) = x χ ( ( 1 << BitDepthY ) - 1 )

[0360] InpC(x)=x χ ( ( 1 << BitDepthC ) - 1 )

[0361] When nnpfc_inp_format_idc is 1, the input value of NNPF is an unsigned integer, and the InpY() and InpC() functions are expressed as follows.

[0362] shiftY = BitDepthY - inpTensorBitDepthY

[0363] if( inpTensorBitDepthY >= BitDepthY)

[0364] InpY(x) = x << ( inpTensorBitDepthY - BitDepthY )

[0365] otherwise

[0366] InpY(x) = Clip3(0, (1 << inpTensorBitDepthY ) - 1, (x + (1 << (shiftY - 1 ) ) ) >> shiftY )

[0367] shiftC = BitDepthC - inpTensorBitDepthC

[0368] If inpTensorBitDepthC >= BitDepthC

[0369] InpC(x) = x << ( inpTensorBitDepthC - BitDepthC )

[0370] otherwise

[0371] InpC(x) = Clip3(0, (1 << inpTensorBitDepthC ) - 1, (x + (1 << (shiftC - ) ) ) >> shiftC )

[0372] The variable inpTensorBitDepthY is derived from the syntax element nnpfc_inp_tensor_luma_bitdepth_minus8 specified below. The variable inpTensorBitDepthC is derived from the syntax element nnpfc_inp_tensor_chroma_bitdepth_minus8 specified below.

[0373] If the value of nnpfc_inp_format_idc is greater than 1, it is reserved for future specifications by ITU-T | ISO / IEC and does not exist in bitstreams compliant with this version of this document. Decoders compliant with this version of this document ignore NNPFC SEI messages containing the reserved value of nnpfc_inp_format_idc.

[0374] If the value of nnpfc_auxiliary_inp_idc is greater than 0, it indicates that there is auxiliary input data in the input tensor of NNPF. If nnpfc_auxiliary_inp_idc is 0, it indicates that there is no auxiliary input data in the input tensor. If nnpfc_auxiliary_inp_idc is 1, it indicates that auxiliary input data is derived as specified in the formula (inpTensorBitDepthY = nnpfc_inp_tensor_luma_bitdepth_minus8 + 8).

[0375] The value of nnpfc_auxiliary_inp_idc is in the range from 0 to 1 in bitstreams compliant with this version of this document. Values ​​of nnpfc_auxiliary_inp_idc from 2 to 255 (inclusive) are reserved for future use by ITU-T | ISO / IEC and do not exist in bitstreams compliant with this version of this document. Decoders compliant with this version of this document ignore NNPFC SEI messages where nnpfc_auxiliary_inp_idc is in the range from 2 to 255. If the value of nnpfc_auxiliary_inp_idc exceeds 255, it does not exist in bitstreams compliant with this version of this document and is not reserved for future use.

[0376] nnpfc_inp_order_idc represents a method of forming an input tensor for NNPF by sorting an array of samples from the input picture.

[0377] The nnpfc_inp_order_idc value is in the range of 0 to 3 in bitstreams compliant with this version of this document. The range of nnpfc_inp_order_idc values ​​from 4 to 255 is reserved for future use by ITU-T | ISO / IEC and does not exist in bitstreams compliant with this version of this document. Decoders compliant with this version of this document ignore NNPFC SEI messages where the nnpfc_inp_order_idc value is in the range of 4 to 255. If the nnpfc_inp_order_idc value exceeds 255, the value does not exist in bitstreams compliant with this version of this document and is not reserved for future use.

[0378] If ChromaFormatIdc is not 1, nnpfc_inp_order_idc is not 3.

[0379] If ChromaFormatIdc is 0, nnpfc_inp_order_idc is not 0.

[0380] If chromaUpsamplingFlag is 1, nnpfc_inp_order_idc is not 0.

[0381] Table 7 (Description of nnpfc_inp_order_idc values) provides a description of the nnpfc_inp_order_idc values.

[0382] [Table 7]

[0383]

[0384] FIG. 25 illustrates the process of inducing a luma channel in a luma component according to the embodiments.

[0385] Figure 25 is an example of deriving 4 luminance channels (right) from the luminance component when nnpfc_inp_order_idc is 3.

[0386] nnpfc_inp_tensor_luma_bitdepth_minus8 + 8 represents the bit depth of the luminance sample values ​​in the input integer tensor. The value of inpTensorBitDepthY is derived as follows.

[0387] inpTensorBitDepthY = nnpfc_inp_tensor_luma_bitdepth_minus8 + 8 (85)

[0388] For bitstream compatibility, the value of nnpfc_inp_tensor_luma_bitdepth_minus8 is in the range of 0 to 24 (inclusive).

[0389] nnpfc_inp_tensor_chroma_bitdepth_minus8 + 8 represents the bit depth of the chroma sample values ​​in the input integer tensor. The value of inpTensorBitDepthC is derived as follows.

[0390] inpTensorBitDepthC = nnpfc_inp_tensor_chroma_bitdepth_minus8 + 8

[0391] For bitstream compatibility, the value of nnpfc_inp_tensor_chroma_bitdepth_minus8 is in the range from 0 to 24.

[0392] When nnpfc_auxiliary_inp_idc is 1, the variable strengthControlScaledVal is derived as follows.

[0393] for( i = 0; i < numInputPics; i++ )

[0394] if(nnpfc_inp_format_idc = = 1)

[0395] if( nnpfc_inp_order_idc = = 0 | | nnpfc_inp_order_idc = = 2 | |

[0396] nnpfc_inp_order_idc = = 3 )

[0397] strengthControlScaledVal[ i ] =

[0398] Floor ( StrengthControlVal[ i ] * ( ( 1 << inpTensorBitDepthY ) - 1 ) )

[0399] else if(nnpfc_inp_order_idc = = 1)

[0400] strengthControlScaledVal[ i ] =

[0401] Floor ( StrengthControlVal[ i ] * ( ( 1 << inpTensorBitDepthC ) - 1 ) )

[0402] otherwise

[0403] strengthControlScaledVal[i] = StrengthControlVal[i]

[0404] A patch is a rectangular array of samples extracted from a component of a picture (e.g., a luma or chroma component).

[0405] The DeriveInputTensors() process, which derives the input tensor inputTensor for given vertical sample coordinates cTop and horizontal sample coordinates cLeft, represents the top-left sample position of the sample patch contained in the input tensor and is defined as follows.

[0406] for( i = 0; i < numInputPics; i++ ) {

[0407] if(nnpfc_inp_order_idc = = 0)

[0408] for( yP = -nnpfc_overlap; yP < inpPatchHeight + nnpfc_overlap; yP++)

[0409] for( xP = -nnpfc_overlap; xP < inpPatchWidth + nnpfc_overlap; xP++ ) {

[0410] inpVal = InpY( InpSampleVal( cTop + yP, cLeft + xP, CroppedHeight,

[0411] CroppedWidth, CroppedYPic[ i ], 0 ) )

[0412] yPovlp = yP + nnpfc_overlap

[0413] xPovlp = xP + nnpfc_overlap

[0414] if( !nnpfc_component_last_flag )

[0415] inputTensor

[0000] [ i ]

[0000] [ yPovlp ][ xPovlp ] = inpVal

[0416] else

[0417] inputTensor

[0000] [ i ][ yPovlp ][ xPovlp ]

[0000] = inpVal

[0418] if( nnpfc_auxiliary_inp_idc = = 1 )

[0419] if( !nnpfc_component_last_flag )

[0420] inputTensor

[0000] [ i ]

[0001] [ yPovlp ][ xPovlp ] = strengthControlScaledVal[ i ]

[0421] else

[0422] inputTensor

[0000] [ i ][ yPovlp ][ xPovlp ]

[0001] = strengthControlScaledVal[ i ]

[0423] }

[0424] else if( nnpfc_inp_order_idc = = 1 )

[0425] for( yP = -nnpfc_overlap; yP < inpPatchHeight + nnpfc_overlap; yP++)

[0426] for( xP = -nnpfc_overlap; xP < inpPatchWidth + nnpfc_overlap; xP++ ) {

[0427] inpCbVal = InpC( InpSampleVal( cTop + yP, cLeft + xP, CroppedHeight / SubHeightC,

[0428] CroppedWidth / SubWidthC, CroppedCbPic[ i ], 1 ) )

[0429] inpCrVal = InpC( InpSampleVal( cTop + yP, cLeft + xP, CroppedHeight / SubHeightC,

[0430] CroppedWidth / SubWidthC, CroppedCrPic[ i ], 2 ) )

[0431] yPovlp = yP + nnpfc_overlap

[0432] xPovlp = xP + nnpfc_overlap

[0433] if( !nnpfc_component_last_flag ) {

[0434] inputTensor

[0000] [ i ]

[0000] [ yPovlp ][ xPovlp ] = inpCbVal

[0435] inputTensor

[0000] [ i ]

[0001] [ yPovlp ][ xPovlp ] = inpCrVal

[0436] } else {

[0437] inputTensor

[0000] [ i ][ yPovlp ][ xPovlp ]

[0000] = inpCbVal

[0438] inputTensor

[0000] [ i ][ yPovlp ][ xPovlp ]

[0001] = inpCrVal

[0439] }

[0440] if( nnpfc_auxiliary_inp_idc = = 1 )

[0441] if( !nnpfc_component_last_flag )

[0442] inputTensor

[0000] [ i ]

[0002] [ yPovlp ][ xPovlp ] = strengthControlScaledVal[ i ]

[0443] else

[0444] inputTensor

[0000] [ i ][ yPovlp ][ xPovlp ]

[0002] = strengthControlScaledVal[ i ]

[0445] }

[0446] else if( nnpfc_inp_order_idc = = 2 )

[0447] for( yP = -nnpfc_overlap; yP < inpPatchHeight + nnpfc_overlap; yP++)

[0448] for( xP = -nnpfc_overlap; xP < inpPatchWidth + nnpfc_overlap; xP++ ) {

[0449] yY = cTop + yP

[0450] xY = cLeft + xP

[0451] yC = yY / SubHeightC

[0452] xC = xY / SubWidthC

[0453] inpYVal = InpY( InpSampleVal( yY, xY, CroppedHeight,

[0454] CroppedWidth, CroppedYPic[ i ], 0 ) )

[0455] inpCbVal = InpC( InpSampleVal( yC, xC, CroppedHeight / SubHeightC,

[0456] CroppedWidth / SubWidthC, CroppedCbPic[ i ], 1 ) )

[0457] inpCrVal = InpC( InpSampleVal( yC, xC, CroppedHeight / SubHeightC,

[0458] CroppedWidth / SubWidthC, CroppedCrPic[ i ], 2 ) )

[0459] yPovlp = yP + nnpfc_overlap

[0460] xPovlp = xP + nnpfc_overlap

[0461] if( !nnpfc_component_last_flag ) {

[0462] inputTensor

[0000] [ i ]

[0000] [ yPovlp ][ xPovlp ] = inpYVal

[0463]

[0464] *

[0465] * inputTensor

[0000] [ i ]

[0001] [ yPovlp ][ xPovlp ] = inpCbVal

[0466] inputTensor

[0000] [ i ]

[0002] [ yPovlp ][ xPovlp ] = inpCrVal

[0467] } else {

[0468] inputTensor

[0000] [ i ][ yPovlp ][ xPovlp ]

[0000] = inpYVal

[0469]

[0470] *

[0471] * inputTensor

[0000] [ i ][ yPovlp ][ xPovlp ]

[0001] = inpCbVal

[0472] inputTensor

[0000] [ i ][ yPovlp ][ xPovlp ]

[0002] = inpCrVal

[0473] }

[0474] if( nnpfc_auxiliary_inp_idc = = 1 )

[0475] if( !nnpfc_component_last_flag )

[0476] inputTensor

[0000] [ i ]

[0003] [ yPovlp ][ xPovlp ] = strengthControlScaledVal[ i ]

[0477] else

[0478] inputTensor

[0000] [ i ][ yPovlp ][ xPovlp ]

[0003] = strengthControlScaledVal[ i ]

[0479] }

[0480] else if( nnpfc_inp_order_idc = = 3 )

[0481] for( yP = -nnpfc_overlap; yP < inpPatchHeight + nnpfc_overlap; yP++)

[0482] for( xP = -nnpfc_overlap; xP < inpPatchWidth + nnpfc_overlap; xP++ ) {

[0483] yTL = cTop + yP * 2

[0484] xTL = cLeft + xP * 2

[0485] yBR = yTL + 1

[0486] xBR = xTL + 1

[0487] yC = cTop / 2 + yP

[0488] xC = cLeft / 2 + xP

[0489] inpTLVal = InpY( InpSampleVal( yTL, xTL, CroppedHeight,

[0490] CroppedWidth, CroppedYPic[ i ], 0 ) )

[0491] inpTRVal = InpY( InpSampleVal( yTL, xBR, CroppedHeight,

[0492] CroppedWidth, CroppedYPic[ i ], 0 ) )

[0493] inpBLVal = InpY( InpSampleVal( yBR, xTL, CroppedHeight,

[0494] CroppedWidth, CroppedYPic[ i ], 0 ) )

[0495] inpBRVal = InpY( InpSampleVal( yBR, xBR, CroppedHeight,

[0496] CroppedWidth, CroppedYPic[ i ], 0 ) )

[0497] inpCbVal = InpC( InpSampleVal( yC, xC, CroppedHeight / 2,

[0498] CroppedWidth / 2, CroppedCbPic[ i ], 1 ) )

[0499] inpCrVal = InpC( InpSampleVal( yC, xC, CroppedHeight / 2,

[0500] CroppedWidth / 2, CroppedCrPic[ i ], 2 ) )

[0501] yPovlp = yP + nnpfc_overlap

[0502] xPovlp = xP + nnpfc_overlap

[0503] if( !nnpfc_component_last_flag ) {

[0504] inputTensor

[0000] [ i ]

[0000] [ yPovlp ][ xPovlp ] = inpTLVal

[0505] inputTensor

[0000] [ i ]

[0001] [ yPovlp ][ xPovlp ] = inpTRVal

[0506] inputTensor

[0000] [ i ]

[0002] [ yPovlp ][ xPovlp ] = inpBLVal

[0507] inputTensor

[0000] [ i ]

[0003] [ yPovlp ][ xPovlp ] = inpBRVal

[0508] inputTensor

[0000] [ i ]

[0004] [ yPovlp ][ xPovlp ] = inpCbVal

[0509] inputTensor

[0000] [ i ]

[0005] [ yPovlp ][ xPovlp ] = inpCrVal

[0510] } else {

[0511] inputTensor

[0000] [ i ][ yPovlp ][ xPovlp ]

[0000] = inpTLVal

[0512] inputTensor

[0000] [ i ][ yPovlp ][ xPovlp ]

[0001] = inpTRVal

[0513] inputTensor

[0000] [ i ][ yPovlp ][ xPovlp ]

[0002] = inpBLVal

[0514] inputTensor

[0000] [ i ][ yPovlp ][ xPovlp ]

[0003] = inpBRVal

[0515] inputTensor

[0000] [ i ][ yPovlp ][ xPovlp ]

[0004] = inpCbVal

[0516] inputTensor

[0000] [ i ][ yPovlp ][ xPovlp ]

[0005] = inpCrVal

[0517] }

[0518] if( nnpfc_auxiliary_inp_idc = = 1 )

[0519] if( !nnpfc_component_last_flag )

[0520] inputTensor

[0000] [ i ]

[0006] [ yPovlp ][ xPovlp ] = strengthControlScaledVal[ i ]

[0521] else

[0522] inputTensor

[0000] [ i ][ yPovlp ][ xPovlp ]

[0006] = strengthControlScaledVal[ i ]

[0523] }

[0524] }

[0525] If nnpfc_out_format_idc is 0, the sample values ​​output by NNPF are real numbers, and the range of values ​​from 0 to 1 is linearly mapped to the range of unsigned integer values ​​from 0 to (1 << bitDepth) - 1 for the desired bit depth bitDepth for subsequent post-processing or display.

[0526] If nnpfc_out_format_idc is 1, it indicates that the luminance sample value output by NNPF is an unsigned integer from 0 to ( 1 << outTensorBitDepthY ) - 1, and the chroma sample value output by NNPF is an unsigned integer from 0 to ( 1 << outTensorBitDepthC ) - 1.

[0527] nnpfc_out_format_idc values ​​greater than 1 are reserved for future specifications by ITU-T | ISO / IEC and should not be included in bitstreams compliant with this version of this document. Decoders compliant with this version of this document ignore NNPFC SEI messages containing reserved nnpfc_out_format_idc values.

[0528] nnpfc_out_order_idc indicates the output order of samples generated by NNPF.

[0529] The nnpfc_out_order_idc value is in the range of 0 to 3 in bitstreams compliant with this version of this document. Values ​​of nnpfc_out_order_idc from 4 to 255 (inclusive) are reserved for future use by ITU-T | ISO / IEC and do not exist in bitstreams compliant with this version of this document. Decoders compliant with this version of this document ignore NNPFC SEI messages where nnpfc_out_order_idc is in the range of 4 to 255. nnpfc_out_order_idc values ​​greater than 255 do not exist in bitstreams compliant with this version of this document and are not reserved for future use.

[0530] If chromaUpsamplingFlag is 1, nnpfc_out_order_idc cannot be 0 or 3.

[0531] If colourizationFlag is 1, nnpfc_out_order_idc cannot be 0.

[0532] Table 8 (Description of nnpfc_out_order_idc values) provides a description of the nnpfc_out_order_idc values.

[0533] [Table 8]

[0534]

[0535] nnpfc_out_tensor_luma_bitdepth_minus8 + 8 represents the bit depth of the lumina sample values ​​in the output integer tensor. The value of nnpfc_out_tensor_luma_bitdepth_minus8 ranges from 0 to 24. The value of outTensorBitDepthY is derived as follows.

[0536] outTensorBitDepthY = nnpfc_out_tensor_luma_bitdepth_minus8 + 8

[0537] nnpfc_out_tensor_chroma_bitdepth_minus8 + 8 represents the bit depth of the chroma sample values ​​in the output integer tensor. The value of nnpfc_out_tensor_chroma_bitdepth_minus8 ranges from 0 to 24. The value of outTensorBitDepthC is derived as follows.

[0538] outTensorBitDepthC = nnpfc_out_tensor_chroma_bitdepth_minus8 + 8

[0539] If bitDepthUpsamplingFlag is 1, the value of nnpfc_out_format_idc must be 1 and satisfy one or more of the following conditions.

[0540] nnpfc_out_tensor_luma_bitdepth_minus8 exists and outTensorBitDepthY is greater than BitDepthY.

[0541] nnpfc_out_tensor_chroma_bitdepth_minus8 exists and outTensorBitDepthC is greater than BitDepthC.

[0542] If nnpfc_inp_tensor_luma_bitdepth_minus8, nnpfc_inp_tensor_chroma_bitdepth_minus8, nnpfc_out_tensor_luma_bitdepth_minus8, and nnpfc_out_tensor_chroma_bitdepth_minus8 exist and outTensorBitDepthY is greater than inpTensorBitDepthY, then outTensorBitDepthC cannot be less than inpTensorBitDepthC. If nnpfc_inp_tensor_luma_bitdepth_minus8, nnpfc_inp_tensor_chroma_bitdepth_minus8, nnpfc_out_tensor_luma_bitdepth_minus8, and nnpfc_out_tensor_chroma_bitdepth_minus8 exist and outTensorBitDepthC is greater than inpTensorBitDepthC, then outTensorBitDepthY cannot be less than inpTensorBitDepthY.

[0543] The StoreOutputTensors() process, which derives sample values ​​of the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor outputTensor for a given vertical sample coordinate cTop and a horizontal sample coordinate cLeft specifying the top-left sample position of the sample patch included in the input tensor, is expressed as follows.

[0544] for( i = 0; i < numOutputPics; i++ ) {

[0545] if(nnpfc_out_order_idc = = 0)

[0546] for(yP = 0; yP < outPatchHeight; yP++)

[0547] for( xP = 0; xP < outPatchWidth; xP++ ) {

[0548] yY = cTop * outPatchHeight / inpPatchHeight + yP

[0549] xY = cLeft * outPatchWidth / inpPatchWidth + xP

[0550] if ( yY < nnpfcOutputPicHeight && xY < nnpfcOutputPicWidth )

[0551] if( !nnpfc_component_last_flag )

[0552] FilteredYPic[ i ][ xY ][yY ] = outputTensor

[0000] [ i ]

[0000] [ yP ][ xP ]

[0553] else

[0554] FilteredYPic[ i ][ xY ][ yY ] = outputTensor

[0000] [ i ][ yP ][ xP ]

[0000] }

[0555] else if( nnpfc_out_order_idc = = 1 ) (91)

[0556] for( yP = 0; yP < outPatchCHeight; yP++)

[0557] for( xP = 0; xP < outPatchCWidth; xP++ ) {

[0558] xSrc = cLeft * horCScaling + xP

[0559] ySrc = cTop * verCScaling + yP

[0560] if ( ySrc < nnpfcOutputPicHeight / outSubHeightC &&

[0561] xSrc < nnpfcOutputPicWidth / outSubWidthC )

[0562] if( !nnpfc_component_last_flag ) {

[0563] FilteredCbPic[ i ][ xSrc ][ ySrc ] = outputTensor

[0000] [ i ]

[0000] [ yP ][ xP ]

[0564] FilteredCrPic[ i ][ xSrc ][ ySrc ] = outputTensor

[0000] [ i ]

[0001] [ yP ][ xP ]

[0565] } else {

[0566] FilteredCbPic[ i ][ xSrc ][ ySrc ] = outputTensor

[0000] [ i ][ yP ][ xP ]

[0000]

[0567] FilteredCrPic[ i ][ xSrc ][ ySrc ] = outputTensor

[0000] [ i ][ yP ][ xP ]

[0001]

[0568] }

[0569] }

[0570] else if( nnpfc_out_order_idc = = 2 )

[0571] for( yP = 0; yP < outPatchHeight; yP++)

[0572] for( xP = 0; xP < outPatchWidth; xP++ ) {

[0573] yY = cTop * outPatchHeight / inpPatchHeight + yP

[0574] xY = cLeft * outPatchWidth / inpPatchWidth + xP

[0575] yC = yY / outSubHeightC

[0576] xC = xY / outSubWidthC

[0577] yPc = ( yP / outSubHeightC ) * outSubHeightC

[0578] xPc = ( xP / outSubWidthC ) * outSubWidthC

[0579] if ( yY < nnpfcOutputPicHeight && xY < nnpfcOutputPicWidth )

[0580] if( !nnpfc_component_last_flag ) {

[0581] FilteredYPic[ i ][ xY ][ yY ] = outputTensor

[0000] [ i ]

[0000] [ yP ][ xP ]

[0582] FilteredCbPic[ i ][ xC ][ yC ] = outputTensor

[0000] [ i ]

[0001] [ yPc ][ xPc ]

[0583] FilteredCrPic[ i ][ xC ][ yC ] = outputTensor

[0000] [ i ]

[0002] [ yPc ][ xPc ]

[0584] } else {

[0585] FilteredYPic[ i ][ xY ][ yY ] = outputTensor

[0000] [ i ][ yP ][ xP ]

[0000]

[0586] FilteredCbPic[ i ][ xC ][ yC ] = outputTensor

[0000] [ i ][ yPc ][ xPc ]

[0001]

[0587] FilteredCrPic[ i ][ xC ][ yC ] = outputTensor

[0000] [ i ][ yPc ][ xPc ]

[0002]

[0588] }

[0589] }

[0590] else if( nnpfc_out_order_idc = = 3 )

[0591] for( yP = 0; yP < outPatchHeight; yP++ )

[0592] for( xP = 0; xP < outPatchWidth; xP++ ) {

[0593] ySrc = cTop / 2 * outPatchHeight / inpPatchHeight + yP

[0594] xSrc = cLeft / 2 * outPatchWidth / inpPatchWidth + xP

[0595] if ( ySrc < nnpfcOutputPicHeight / 2 &&

[0596] xSrc < nnpfcOutputPicWidth / 2 )

[0597] if( !nnpfc_component_last_flag ) {

[0598] FilteredYPic[ i ][ xSrc * 2 ][ ySrc * 2 ] = outputTensor

[0000] [ i ]

[0000] [ yP ][ xP ]

[0599] FilteredYPic[ i ][ xSrc * 2 + 1 ][ ySrc * 2 ] = outputTensor

[0000] [ i ]

[0001] [ yP ][ xP ]

[0600] FilteredYPic[ i ][ xSrc * 2 ][ ySrc * 2 + 1 ] = outputTensor

[0000] [ i ]

[0002] [ yP ][ xP ]

[0601] FilteredYPic[ i ][ xSrc * 2 + 1][ ySrc * 2 + 1 ] = outputTensor

[0000] [ i ]

[0003] [ yP ][ xP ]

[0602] FilteredCbPic[ i ][ xSrc ][ ySrc ] = outputTensor

[0000] [ i ]

[0004] [ yP ][ xP ]

[0603] FilteredCrPic[ i ][ xSrc ][ ySrc ] = outputTensor

[0000] [ i ]

[0005] [ yP ][ xP ]

[0604] } else {

[0605] FilteredYPic[ i ][ xSrc * 2 ][ ySrc * 2 ] = outputTensor

[0000] [ i ][ yP ][ xP ]

[0000]

[0606] FilteredYPic[ i ][ xSrc * 2 + 1 ][ ySrc * 2 ] = outputTensor

[0000] [ i ][ yP ][ xP ]

[0001]

[0607] FilteredYPic[ i ][ xSrc * 2 ][ ySrc * 2 + 1 ] = outputTensor

[0000] [ i ][ yP ][ xP ]

[0002]

[0608] FilteredYPic[ i ][ xSrc * 2 + 1][ ySrc * 2 + 1 ] = outputTensor

[0000] [ i ][ yP ][ xP ]

[0003]

[0609] FilteredCbPic[ i ][ xSrc ][ ySrc ] = outputTensor

[0000] [ i ][ yP ][ xP ]

[0004]

[0610] FilteredCrPic[ i ][ xSrc ][ ySrc ] = outputTensor

[0000] [ i ][ yP ][ xP ]

[0005]

[0611] }

[0612] }

[0613] }

[0614] If nnpfc_separate_colour_description_present_flag is 1, it indicates that a unique combination of color primary, transfer property, matrix factor, scaling, and offset values ​​applied in relation to the matrix factor for the picture generated by NNPF is specified in the SEI message syntax structure. If nnpfc_separate_colour_description_present_flag is 0, it indicates that the combination of color primary, transfer property, matrix factor, scaling, and offset values ​​applied in relation to the matrix factor for the picture generated by NNPF is the same as that specified in the VUI parameters of CLVS.

[0615] nnpfc_colour_primaries has the same meaning as the vui_colour_primaries syntax element, but with the following differences.

[0616] nnpfc_colour_primaries represents the color primary colors of the picture generated by applying the NNPF specified in the SEI message, rather than the color primary colors used in CLVS.

[0617] If nnpfc_colour_primaries is not present in the NNPFC SEI message, the value of nnpfc_colour_primaries is inferred to be the same as vui_colour_primaries.

[0618] nnpfc_transfer_characteristics has the same meaning as specified for the vui_transfer_characteristics syntax element, except for the following.

[0619] nnpfc_transfer_characteristics represents the transfer characteristics of the picture generated by applying the NNPF specified in the SEI message, rather than the transfer characteristics used in CLVS.

[0620] If nnpfc_transfer_characteristics is not present in the NNPFC SEI message, the value of nnpfc_transfer_characteristics is inferred to be the same as vui_transfer_characteristics.

[0621] nnpfc_matrix_coeffs describes the equations used to derive luminance and saturation signals from green, blue, red, or the primary colors Y, Z, and X. The meaning of this function applies to the picture generated by applying the NNPF specified in this SEI message, as specified in the MatrixCoefficients of Rec. ITU-T H.273 | ISO / IEC 23091-2, where BitDepthY and BitDepthC are equal to outTensorBitDepthY and outTensorBitDepthC, respectively.

[0622] If nnpfc_matrix_coeffs is not in the NNPFC SEI message, the value of nnpfc_matrix_coeffs is inferred to be the same as vui_matrix_coeffs.

[0623] nnpfc_matrix_coeffs cannot be 0 except when both of the following two conditions are true.

[0624] nnpfc_out_tensor_chroma_bitdepth_minus8 is equal to nnpfc_out_tensor_luma_bitdepth_minus8.

[0625] nnpfc_out_order_idc is 2, outSubHeightC is 1, and outSubWidthC is 1.

[0626] nnpfc_matrix_coeffs cannot be 8 unless one of the following conditions is true.

[0627] nnpfc_out_tensor_chroma_bitdepth_minus8 is equal to nnpfc_out_tensor_luma_bitdepth_minus8.

[0628] nnpfc_out_tensor_chroma_bitdepth_minus8 is equal to nnpfc_out_tensor_luma_bitdepth_minus8 + 1, nnpfc_out_order_idc is equal to 2, outSubHeightC is equal to 1, and outSubWidthC is equal to 1.

[0629] nnpfc_full_range_flag represents the scaling and offset values ​​applied in relation to the matrix coefficients specified in nnpfc_matrix_coeffs. The meaning of this value is the same as that specified in the VideoFullRangeFlag parameter of Rec. ITU-T H.273 | ISO / IEC 23091-2. If nnpfc_full_range_flag is not present, the value is inferred to be 0.

[0630] If the value of nnpfc_chroma_loc_info_present_flag is 1, it indicates that the nnpfc_chroma_sample_loc_type_frame syntax element is present in the NNPFC SEI message. If the value of nnpfc_chroma_loc_info_present_flag is 0, it indicates that the nnpfc_chroma_sample_loc_type_frame syntax element is not present in the NNPFC SEI message. If colourizationFlag is 0 or nnpfc_out_colour_format_idc is not 1, the value of nnpfc_chroma_loc_info_present_flag is equal to 0.

[0631] If nnpfc_chroma_sample_loc_type_frame is not 6 and nnpfc_out_colour_format_idc is 1, it indicates the chroma sample location of the output picture. If nnpfc_chroma_sample_loc_type_frame is 6 and nnpfc_out_colour_format_idc is 1, it indicates that the chroma sample location is unknown, unspecified, or specified in another way not specified in this document. The value of nnpfc_chroma_sample_loc_type_frame ranges from 0 to 6.

[0632] nnpfc_overlap indicates the overlap in the number of horizontal and vertical samples of adjacent input tensors in NNPF. The nnpfc_overlap value ranges from 0 to 16,383.

[0633] If nnpfc_constant_patch_size_flag is 1, it indicates that NNPF exactly accepts the patch sizes specified in nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1 as input. If nnpfc_constant_patch_size_flag is 0, it indicates that NNPF accepts any patch size as input with a width of inpPatchWidth and a height of inpPatchHeight, wherein the width of the extended patch (i.e., the area overlapping with the patch) is equal to inpPatchWidth + 2 * nnpfc_overlap and this width is a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 * nnpfc_overlap, and the height of the extended patch is equal to inpPatchHeight + 2 * nnpfc_overlap and this width is a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 * nnpfc_overlap.

[0634] The value of nnpfc_patch_width_minus1 plus 1 represents the horizontal sample size of the patch size required for the NNPF input when nnpfc_constant_patch_size_flag is 1. The value of nnpfc_patch_width_minus1 ranges from 0 to Min(32,766, CroppedWidth - 1).

[0635] The value of nnpfc_patch_height_minus1 plus 1 represents the number of vertical samples of the patch size required for the NNPF input when nnpfc_constant_patch_size_flag is 1. The value of nnpfc_patch_height_minus1 ranges from 0 to Min(32,766, CroppedHeight - 1).

[0636] nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 * nnpfc_overlap represents the common divisor of all allowed values ​​for the extended patch width required for the NNPF input when nnpfc_constant_patch_size_flag is 0. The value of nnpfc_extended_patch_width_cd_delta_minus1 is in the range from 0 to Min(32,766, CroppedWidth - 1).

[0637] nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 * nnpfc_overlap represents the common divisor of all allowable values ​​of the extended patch height required for the NNPF input when nnpfc_constant_patch_size_flag is 0. The value of nnpfc_extended_patch_height_cd_delta_minus1 is in the range from 0 to Min(32,766, CroppedHeight - 1).

[0638] Set the inpPatchWidth and inpPatchHeight variables to the patch size width and patch size height, respectively.

[0639] If nnpfc_constant_patch_size_flag is 0, the following applies.

[0640] The inpPatchWidth and inpPatchHeight values ​​are provided through external means not specified in this document or are set by the postprocessor itself.

[0641] The value of inpPatchWidth + 2 * nnpfc_overlap must be a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 * nnpfc_overlap, and inpPatchWidth is less than or equal to CroppedWidth. The value of inpPatchHeight + 2 * nnpfc_overlap must be a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 * nnpfc_overlap, and inpPatchHeight is less than or equal to CroppedHeight.

[0642] Otherwise (when nnpfc_constant_patch_size_flag is 1), the inpPatchWidth value is set to nnpfc_patch_width_minus1 + 1 and the inpPatchHeight value is set to nnpfc_patch_height_minus1 + 1.

[0643] The variables outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, and outPatchCHeight are derived as follows.

[0644] outPatchWidth = (nnpfcOutputPicWidth * inpPatchWidth) / CroppedWidth

[0645] outPatchHeight = (nnpfcOutputPicHeight * inpPatchHeight) / CroppedHeight

[0646] horCScaling = SubWidthC / outSubWidthC

[0647] verCScaling = SubHeightC / outSubHeightC

[0648] outPatchCWidth = outPatchWidth * horCScaling

[0649] outPatchCHeight = outPatchHeight * verCScaling

[0650] For bitstream conformance, outPatchWidth * CroppedWidth is equal to nnpfcOutputPicWidth * inpPatchWidth, and outPatchHeight * CroppedHeight is equal to nnpfcOutputPicHeight * inpPatchHeight.

[0651] nnpfc_padding_type represents the padding process when referring to sample locations outside the input picture boundaries, as described in Table 9 (Informative description of nnpfc_padding_type values). The values ​​of nnpfc_padding_type range from 0 to 4 in bitstreams compliant with this version of this document. The range of nnpfc_padding_type values ​​from 5 to 15 is reserved for future use by ITU-T | ISO / IEC and does not exist in bitstreams compliant with this version of this document. Decoders compliant with this revision of this document ignore NNPFC SEI messages where the nnpfc_padding_type value is between 5 and 15 (inclusive). nnpfc_padding_type values ​​greater than 15 do not exist in bitstreams compliant with this revision of this document and are not reserved for future use.

[0652] [Table 9]

[0653]

[0654] nnpfc_luma_padding_val represents the lumina value to be used for padding when nnpfc_padding_type is 4. The value of nnpfc_luma_padding_val is in the range from 0 to (1 << BitDepthY) - 1.

[0655] nnpfc_cb_padding_val represents the Cb value to be used for padding when nnpfc_padding_type is 4. The value of nnpfc_cb_padding_val is in the range from 0 to (1 << BitDepthC) - 1.

[0656] nnpfc_cr_padding_val represents the Cr value to be used for padding when nnpfc_padding_type is 4. The nnpfc_cr_padding_val value is in the range of 0 to (1 << BitDepthC) - 1.

[0657] The InpSampleVal(y, x, picHeight, picWidth, croppedPic, cIdx) function takes the vertical sample position y, horizontal sample position x, picture height picHeight, picture width picWidth, sample array croppedPic, and component index cIdx (0 for Luma, 1 for Cb, 2 for Cr) as input and returns the sampleVal value derived as follows.

[0658] For the input to the InpSampleVal() function, vertical positions are listed before horizontal positions for compatibility with the input tensor rules of some inference engines.

[0659] if(nnpfc_padding_type = = 0)

[0660] if( y < 0 | | x < 0 | | y >= picHeight | | x >= picWidth )

[0661] sampleVal = 0

[0662] else

[0663] sampleVal = croppedPic[x][y] (98)

[0664] else if(nnpfc_padding_type = = 1)

[0665] sampleVal = croppedPic[ Clip3( 0, picWidth - 1, x ) ][ Clip3( 0, picHeight - 1, y ) ]

[0666] else if(nnpfc_padding_type = = 2)

[0667] sampleVal = croppedPic[ Reflect( picWidth - 1, x ) ][ Reflect( picHeight - 1, y ) ]

[0668] else if(nnpfc_padding_type = = 3)

[0669] if( y >= 0 && y < picHeight )

[0670] sampleVal = croppedPic[ Wrap( picWidth - 1, x ) ][ y ]

[0671] else if(nnpfc_padding_type = = 4)

[0672] if( y < 0 | | x < 0 | | y >= picHeight | | x >= picWidth )

[0673] sampleVal = ( cIdx = = 0 nnpfc_luma_padding_val :

[0674] ( cIdx = = 1 ? nnpfc_cb_padding_val : nnpfc_cr_padding_val ) )

[0675] else

[0676] sampleVal = croppedPic[x][y]

[0677] NNPF PostProcessingFilter() is the target NNPF derived from the semantics of the NNPFA SEI message. The following example process can be used with NNPF PostProcessingFilter() to generate a patch-filtered and / or interpolated image. This image contains the Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, respectively, as indicated in nnpfc_out_order_idc.

[0678] if( nnpfc_inp_order_idc = = 0 | | nnpfc_inp_order_idc = = 2 )

[0679] for( cTop = 0; cTop < CroppedHeight; cTop += inpPatchHeight )

[0680] for( cLeft = 0; cLeft < CroppedWidth; cLeft += inpPatchWidth ) {

[0681] DeriveInputTensors( )

[0682] outputTensor = PostProcessingFilter( inputTensor )

[0683] StoreOutputTensors( )

[0684] }

[0685] else if( nnpfc_inp_order_idc = = 1 )

[0686] for( cTop = 0; cTop < CroppedHeight / SubHeightC; cTop += inpPatchHeight )

[0687] for( cLeft = 0; cLeft < CroppedWidth / SubWidthC; cLeft += inpPatchWidth ) {

[0688] DeriveInputTensors( )

[0689] outputTensor = PostProcessingFilter( inputTensor )

[0690] StoreOutputTensors( )

[0691] }

[0692] else if( nnpfc_inp_order_idc = = 3 )

[0693] for( cTop = 0; cTop < CroppedHeight; cTop += inpPatchHeight * 2 )

[0694] for( cLeft = 0; cLeft < CroppedWidth; cLeft += inpPatchWidth * 2 ) {

[0695] DeriveInputTensors()

[0696] outputTensor = PostProcessingFilter( inputTensor )

[0697] StoreOutputTensors()

[0698] }

[0699] If present, the NNPF generated image with index i includes the sample arrays FilteredYPic[ i ], FilteredCbPic[ i ], and FilteredCrPic[ i ] derived by the above formula (cLeft = 0; cLeft < CroppedWidth / SubWidthC; cLeft += inpPatchWidth). The NNPF generated image does not contain overlapping regions.

[0700] The NNPF process consists of the process defined in the above formula (cLeft = 0; cLeft < CroppedWidth / SubWidthC; cLeft += inpPatchWidth) and then the process of outputting the NNPF-generated images in index order. Here, all NNPF-generated images interpolated by NNPF are output, and the NNPF-generated images corresponding to the input images for NNPF are output as specified in the semantics of the NNPFA SEI message.

[0701] If nnpfc_complexity_info_present_flag is 1, it indicates that there is one or more syntax elements representing the complexity of the NNPF associated with nnpfc_id. If nnpfc_complexity_info_present_flag is 0, it specifies that there are no syntax elements representing the complexity of the NNPF associated with nnpfc_id.

[0702] If nnpfc_parameter_type_idc is 0, it indicates that the neural network uses only integer parameters. If nnpfc_parameter_type_flag is 1, it indicates that the neural network can use floating-point or integer parameters. If nnpfc_parameter_type_idc is 2, it indicates that the neural network uses only binary parameters. If nnpfc_parameter_type_idc is 3, it is reserved for future use by ITU-T | ISO / IEC and is not included in bitstreams compliant with this version of this document. Decoders compliant with this version of this document ignore NNPFC SEI messages where nnpfc_parameter_type_idc is 3.

[0703] If nnpfc_log2_parameter_bit_length_minus3 is 0, 1, 2, or 3, it indicates that the neural network does not use parameters with bit lengths greater than 8, 16, 32, and 64, respectively. If nnpfc_parameter_type_idc is present and nnpfc_log2_parameter_bit_length_minus3 is absent, the neural network does not use parameters with a bit length greater than 1.

[0704] nnpfc_num_parameters_idc represents the maximum number of neural network parameters for the NNPF in powers of 2048. If nnpfc_num_parameters_idc is 0, it indicates that the maximum number of neural network parameters is unknown. The value of nnpfc_num_parameters_idc ranges from 0 to 52 (inclusive). nnpfc_num_parameters_idc values ​​greater than 52 are reserved for future use by ITU-T | ISO / IEC and do not exist in bitstreams compliant with this version of this document. Decoders compliant with this version of this document ignore NNPFC SEI messages where nnpfc_num_parameters_idc is greater than 52.

[0705] If the value of nnpfc_num_parameters_idc is greater than 0, the variable maxNumParameters is derived as follows.

[0706] maxNumParameters = ( 2 048 << nnpfc_num_parameters_idc ) - 1

[0707] The requirement for bitstream conformance is that the number of neural network parameters of NNPF must be less than or equal to maxNumParameters.

[0708] If nnpfc_num_kmac_operations_idc is greater than 0, it indicates that the maximum number of multiplicative-accumulator operations per NNPF sample is less than or equal to nnpfc_num_kmac_operations_idc * 1000. If nnpfc_num_kmac_operations_idc is 0, it indicates that the maximum number of multiplicative-accumulator operations of the network is unknown. The value of nnpfc_num_kmac_operations_idc is in the range from 0 to 232-2(2).

[0709] If nnpfc_total_kilobyte_size is greater than 0, it indicates the total size (KB) required to store the uncompressed parameters of the neural network. The total size (bits) is a number greater than or equal to the sum of the bits used to store each parameter. nnpfc_total_kilobyte_size is the total size (bits) divided by 8,000 and rounded. If nnpfc_total_kilobyte_size is 0, it indicates that the total size required to store the parameters of the neural network is unknown. The value of nnpfc_total_kilobyte_size is in the range from 0 to 2³²-2(2).

[0710] If nnpfc_metadata_extension_num_bits is 0, it indicates that there is no nnpfc_reserved_metadata_extension. If nnpfc_metadata_extension_num_bits is greater than 0, it indicates the length (in bits) of the nnpfc_reserved_metadata_extension. In this version of this document, nnpfc_metadata_extension_num_bits is 0. Values ​​for nnpfc_metadata_extension_num_bits in the range of 1 to 2,048 (inclusive) are reserved for future use by ITU-T | ISO / IEC and do not exist in bitstreams compliant with this version of this document. Decoders compliant with this version of this document accept all nnpfc_metadata_extension_num_bits values ​​in the range of 0 to 2,048 (inclusive). If the value of nnpfc_metadata_extension_num_bits is greater than 2,048, it does not exist in the bitstream following this version of this document and is not reserved for future use.

[0711] The nnpfc_reserved_metadata_extension value does not exist in bitstreams following this version of this document. However, decoders following this version of this document ignore the existence and value of nnpfc_reserved_metadata_extension. If nnpfc_reserved_metadata_extension exists, the length (in bits) of nnpfc_metadata_extension is equal to nnpfc_metadata_extension_num_bits.

[0712] nnpfc_reserved_zero_bit_b is equal to 0 in bitstreams following this version of this document. Decoders ignore NNPFC SEI messages where nnpfc_reserved_zero_bit_b is not 0.

[0713] nnpfc_payload_byte[ i ] contains the i-th byte of a bitstream compliant with ISO / IEC 15938-17. The byte sequence for all current values ​​of i, nnpfc_payload_byte[ i ], is a complete bitstream compliant with ISO / IEC 15938-17.

[0714] The encoding method and apparatus according to the embodiments may include and perform the source device of FIG. 1, the encoder of FIG. 2, the transformer of the encoder of FIG. 15, the LFNST of FIG. 16, the CABAC encoding of FIG. 17, the entropy encoding of FIG. 18, the picture encoding of FIG. 21, the coding layer of FIG. 22, the generation of SEI messages of FIG. 23 to 30a and FIG. 30b, the method of FIG. 31, etc.

[0715] The decoding method and apparatus according to the embodiments may include and perform the receiving device of FIG. 2, the decoder of FIG. 3, the inverse transformer of the decoder of FIG. 15, the LFNST of FIG. 16, the entropy decoding of FIG. 19, the picture decoding of FIG. 20, the coding layer of FIG. 22, the acquisition of SEI messages from FIG. 23 to FIG. 30a and FIG. 30b, the method of FIG. 32, etc.

[0716] The embodiments include a method and apparatus for signaling an update flag for each display rectangle in a display rectangle SEI message for an encoded video bitstream.

[0717] The present disclosure describes a method for signaling update flags for each display rectangle within a display rectangles SEI message in a picture unit for an encoded video bitstream. The described method is based on Versatile Video Coding (VVC) and Versatile supplemental enhancement information messages (VSEI) for encoded video bitstreams, but may be applicable to other video encoding technologies.

[0718] This disclosure relates to: Versatile Video Coding (VVC). The latest VVC specification can be found at: https: / / jvet-experts.org / doc_end_user / documents / 20_Teleconference / wg11 / JVET-T2001-v2.zip. Versatile supplemental enhancement information messages for coded video bitstreams (VSEI). The latest VSEI specification can be found at: https: / / www.itu.int / rec / T-REC-H.274. Additional SEI messages for VSEI (Draft 3). The latest draft can be found at: https: / / jvet-experts.org / doc_end_user / documents / 31_Geneva / wg11 / JVET-AE2006-v2.zip. Technologies under consideration for future extensions of VSEI (draft 7). The latest draft can be found at: https: / / jvet-experts.org / doc_end_user / documents / 37_Geneva / wg11 / JVET-AK2032-v1.zip. Additional SEI messages for VSEI (Draft 5). The latest draft can be found at: https: / / jvet-experts.org / doc_end_user / documents / 37_Geneva / wg11 / JVET-AK2006-v2.zip.

[0719] FIG. 26 shows the neural-network post-filter activation SEI message syntax according to the embodiments.

[0720] The Neural Network Post-processing Filter Enable (NNPFA) SEI message can enable or disable the possible use of the target Neural Network Post-processing Filter (NNPF) identified by nnpfa_target_id and nnpfa_target_base_flag for post-processing filtering of a series of pictures.

[0721] The target NNPF can be derived as follows.

[0722] If nnpfa_target_base_flag is equal to 1, the target NNPF may be a base NNPF where nnpfc_id is equal to nnpfa_target_id.

[0723] Otherwise, that is, when nnpfa_target_base_flag is equal to 0, the target NNPF precedes the first VCL NAL unit of the current picture in the decoding order and is not a repetition of the NNPFC SEI message containing the base NNPF, and may be the NNPF specified by the last NNPFC SEI message where nnpfc_id is equal to nnpfa_target_id.

[0724] Multiple NNPFA SEI messages may exist for the same picture. For example, if NNPFs are intended for different purposes or for filtering different color components, multiple NNPFA SEI messages may exist for the same picture.

[0725] To use SEI messages, the definition of the CandInPicList variable may be required.

[0726] CandInPicList may include a list of pictures in output order to which input pictures for the target NNPF are selected or padded.

[0727] Target NNPF can be used for post-processing filtering for each picture in CandInPicList, which is a corresponding picture or an associated insertion picture for the picture where the NNPFA SEI message persists.

[0728] The target ID (nnpfa_target_id) may indicate the nnpfc_id of a target NNPF that is associated with the current picture and is specified by one or more NNPFC SEI messages where nnpfc_id is equal to nnpfa_target_id. The value of nnpfa_target_id may be in the range of 0 to 2^32 - 2.

[0729] An NNPFA SEI message with a specific nnpfa_target_id value may not exist in the current PU, but may exist in the current PU if one or both of the following conditions are true.

[0730] In the current CLVS, an NNPFC SEI message with nnpfc_id equal to a specific nnpfa_target_id value may exist in a PU that precedes the current PU in the decoding order.

[0731] Currently, within the PU, there may be NNPFC SEI messages where nnpfc_id is the same as a specific nnpfa_target_id value.

[0732] If PU includes both an NNPFC SEI message with a specific nnpfc_id value and an NNPFA SEI message where nnpfa_target_id is the same as a specific nnpfc_id value, the NNPFC SEI message may precede the NNPFA SEI message in the decoding order.

[0733] If the cancellation flag (nnpfa_cancel_flag) is equal to 1, it may indicate that the persistence of a target NNPF set by any previous NNPFA SEI message having the same nnpfa_target_id as the current SEI message is canceled. That is, the target NNPF may no longer be used unless it is activated by another NNPFA SEI message having the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag being equal to 0. If nnpfa_cancel_flag is equal to 0, nnpfa_persistence_flag, nnpfa_target_base_flag, nnpfa_no_prev_clvs_flag, nnpfa_no_foll_clvs_flag (where nnpfa_persistence_flag is equal to 1), and nnpfa_num_output_entries may follow.

[0734] The persistence flag (nnpfa_persistence_flag) can specify the persistence of the target NNPF for the current layer.

[0735] If nnpfa_persistence_flag is equal to 0, it can be specified that NNPFA SEI messages persist only for the current picture.

[0736] If nnpfa_persistence_flag is equal to 1, it can be specified that NNPFA SEI messages persist for the current picture and all subsequent pictures of the current layer in output order, and may persist until one or more of the following conditions are true.

[0737] New CLVS of the current layer can be started.

[0738] The bitstream can be terminated.

[0739] In the current layer, a picture associated with an NNPFA SEI message having the same nnpfa_target_id as the current SEI message may be output, and that picture may follow the current picture in the output order.

[0740] Target Pictures (nnpfcTargetPictures) can be defined as a set of pictures associated with an NNPFC SEI message corresponding to a target NNPF, and nnpfaTargetPictures can be defined as a set of pictures associated with an NNPFA SEI message. As a requirement for bitstream conformity, any picture included in nnpfaTargetPictures may also be included in nnpfcTargetPictures.

[0741] If the target base flag (nnpfa_target_base_flag) is 1, the target NNPF can be identified as a base NNPF where nnpfc_id is equal to nnpfa_target_id. If nnpfa_target_base_flag is 0, the target NNPF can be identified as an NNPF that precedes the first VCL NAL unit of the current picture in the decoding order, is not a repetition of the NNPFC SEI message containing the base NNPF, and is identified by the last NNPFC SEI message where nnpfc_id is equal to nnpfa_target_id.

[0742] If nnpfa_target_base_flag in an NNPFA SEI message is equal to 0, at least one NNPFC SEI message having nnpfc_id equal to nnpfa_target_id and nnpfc_base_flag equal to 0 may precede the NNPFA SEI message in decoding order.

[0743] If nnpfa_no_prev_clvs_flag is equal to 1, it can be specified that the input pictures for NNPF do not originate from previous CLVS. If nnpfa_no_prev_clvs_flag is equal to 0, it can be specified that the input pictures for NNPF may or may not originate from previous CLVS.

[0744] If nnpfa_no_foll_clvs_flag is equal to 1, it can be specified that when an NNPFA SEI message is persisted to the last PU of CLVS in output order, the NNPFA SEI message is treated as if it were persisted to the last PU of the current layer in output order within the bitstream. If an NNPFA SEI message is not persisted to the last PU of CLVS in output order or if nnpfa_no_foll_clvs_flag is equal to 0, the value of nnpfa_no_foll_clvs_flag may not have a specific effect.

[0745] The number of output entries (nnpfa_num_output_entries) can specify the number of nnpfa_output_flag[i] syntax elements present in the NNPFA SEI message. The value of nnpfa_num_output_entries can be in the range from 0 to NumInpPicsInOutputTensor. If PictureRateUpsamplingFlag is equal to 0, TemporalExtrapolationFlag is equal to 0, and nnpfa_num_output_entries is equal to NumInpPicsInOutputTensor, nnpfa_output_flag[i] can be equal to 1 for at least one value of i within the range from 0 to nnpfa_num_output_entries - 1.

[0746] If the output flag (nnpfa_output_flag[i]) is 1, it can be specified that the NNPF generated picture corresponding to the input picture with index InpIdx[i] is output by the NNPF process activated by the corresponding NNPFA SEI message, according to the NNPF process defined in the semantics of the NNPFC SEI message. If the nnpfa_output_flag[i] is 0, it can be specified that the NNPF generated picture corresponding to the input picture with index InpIdx[i] is not output by the NNPF process activated by the corresponding NNPFA SEI message. If nnpfa_num_output_entries is smaller than NumInpPicsInOutputTensor, nnpfa_output_flag[i] can be derived to be equal to 1 for each i value within the range of nnpfa_num_output_entries to NumInpPicsInOutputTensor - 1.

[0747] If the prompt update flag (nnpfa_prompt_update_flag) is equal to 1, it can be determined that an nnpfa_prompt syntax element exists and that an nnpfa_alignment_zero_bit syntax element may exist. If nnpfa_prompt_update_flag is equal to 0, it can be determined that neither an nnpfa_prompt syntax element nor an nnpfa_alignment_zero_bit syntax element exists. If they do not exist, the value of nnpfa_prompt_update_flag can be derived to equal to 0.

[0748] If the inband prompt flag (nnpfc_inband_prompt_flag) is equal to 0, the value of the existing nnpfa_prompt_update_flag may be equal to 0.

[0749] nnpfa_alignment_zero_bit can be equal to 0.

[0750] The prompt (nnpfa_prompt) can specify a text string prompt used as input to the target NNPF. If nnpfa_prompt_update_flag is equal to 1, nnpfa_prompt may not be a null string.

[0751] If the seed update flag (nnpfa_seed_update_flag) is equal to 1, it can be determined that an nnpfa_seed syntax element exists. If nnpfa_seed_update_flag is equal to 0, it can be determined that an nnpfa_seed syntax element does not exist. If it does not exist, the value of nnpfa_seed_update_flag can be derived to be equal to 0.

[0752] If (nnpfc_auxiliary_inp_idc & 4) is equal to 0 or nnpfc_inband_seed_flag exists and is equal to 0, the value of nnpfa_seed_update_flag may be equal to 0.

[0753] The seed (nnpfa_seed) can specify the seed value used as input for the target NNPF.

[0754] The number of input picture shifts (nnpfa_num_input_pic_shift) can specify the number of times input pictures are shifted in a list of candidate input pictures to obtain the final input pictures for the target NNPF. If none exist, the value of nnpfa_num_input_pic_shift can be derived as 0. The value of nnpfa_num_input_pic_shift can be in the range of 0 to 63.

[0755] FIGS. 27a, FIGS. 27b, FIGS. 27c, and FIGS. 27d illustrate constituent rectangles SEI message syntax according to embodiments.

[0756] Constituent rectangles SEI messages enable the composition of multiple rectangles in an encoded picture within one or more layers and can provide information about the rectangles including ID, type, text description, location, and size.

[0757] To use the Constituent rectangles SEI message, the definition of the following variables may be required.

[0758] PicWidthInLumaSamples[lId] and PicHeightInLumaSamples[lId] can be defined as the picture width and picture height in luma sample units for lId, respectively.

[0759] MaxPicWidth[lId] and MaxPicHeight[lId] can be defined as the maximum picture width and maximum picture height, respectively, in luma sample units for lId.

[0760] A chroma format specifier, denoted by ChromaFormatIdc, can be defined.

[0761] A bit depth for samples of a luminance component, denoted by BitDepthY, and a bit depth for samples of two associated chroma components, denoted by BitDepthC, can be defined if ChromaFormatIdc is not equal to 0.

[0762] The number of subpictures represented by NumSubpics[lId] can be defined.

[0763] Arrays of the top-left X and Y positions of subpictures, denoted as SubPicTopLeftX[lId][i] and SubPicTopLeftY[lId][i] respectively, can be defined.

[0764] Arrays of widths and heights of subpictures, denoted as SubPicTopLeftX[lId][i] and SubPicTopLeftY[lId][i] respectively, can be defined.

[0765] If a constituent rectangles SEI message exists in any picture unit other than the first picture unit of CLVS in the decoding order, a constituent rectangles SEI message with the same payload content may exist in the first picture unit of CLVS in the decoding order.

[0766] If cr_enhanced_chroma_format_enabled_flag is equal to 1, cr_group_enhanced_chroma_format_idc[i] may indicate that a syntax element exists. If cr_enhanced_chroma_format_enabled_flag is equal to 0, cr_group_enhanced_chroma_format_idc[i] may indicate that a syntax element does not exist.

[0767] The value obtained by adding 1 to the number of layers (cr_num_layers_minus1) can specify the number of layers in which constituent rectangles are described in the SEI message.

[0768] Layer ID (cr_layer_id[lIdx]) can identify the layer identifier of the lIdx-th layer.

[0769] The variable lId can be set to be equal to cr_layer_id[lIdx].

[0770] The value obtained by adding 1 to the number of rectangles in the layer (cr_num_rects_in_layer_minus1[lIdx]) can specify the number of constituent rectangles in the lIdx-th layer for which information is signaled in the SEI message.

[0771] The variable CrNumRects can be derived as follows.

[0772] CrNumRects = 0

[0773] for( lIdx = 0; lIdx <= cr_num_layers_minus1; lIdx++ )

[0774] CrNumRects += cr_num_rects_in_layer[lIdx] + 1.

[0775] If the rectangle ID enable flag (cr_rect_id_enabled_flag) is equal to 1, it can be determined that the cr_rect_id_present_flag[i] syntax element exists in the SEI message. If the cr_rect_id_present_flag is equal to 0, it can be determined that the cr_rect_id_present_flag[i] syntax element does not exist in the SEI message.

[0776] If the associated rectangle ID enable flag (cr_associated_rect_id_enabled_flag) is equal to 1, it can be determined that the cr_associated_rect_id_present_flag[i] syntax elements may exist in the SEI message. If the cr_associated_rect_id_enabled_flag is equal to 0, it can be determined that the cr_associated_rect_id_present_flag[i] syntax elements do not exist in the SEI message.

[0777] The number of groups (cr_num_groups) can specify the maximum number of square groups described in the SEI message.

[0778] The value obtained by adding 1 to the rectangle ID length (cr_rect_id_len_minus1) can specify the length of the cr_rect_id[i] and cr_associated_rect_id[i] syntax elements.

[0779] If the subpicture index enable flag (cr_subpic_idx_enabled_flag) is equal to 1, it can be determined that the cr_subpic_idx_len_minus1 syntax element exists in the SEI message. If the cr_subpic_idx_enabled_flag is equal to 0, it can be determined that the cr_subpic_idx_len_minus1 syntax element does not exist in the SEI message.

[0780] The value obtained by adding 1 to the subpicture index length (cr_subpic_idx_len_minus1) can specify the bitwise length of cr_subpic_idx[lIdx][crIdx].

[0781] If the rectangle type enable flag (cr_rect_type_enabled_flag) is 1, it can be determined that the cr_rect_type_present_flag[i] syntax element is present in the SEI message. If the cr_rect_type_enabled_flag is 0, it can be determined that the cr_rect_type_present_flag[i] syntax element is not present in the SEI message.

[0782] If cr_group_enhanced_chroma_format_idc[i] is equal to 2, it may indicate that the i-th target picture of 4:4:4 chroma format is formed from three squares of the i-th square group. If cr_group_enhanced_chroma_format_idc[i] is equal to 1, it may indicate that the i-th target picture of 4:2:2 chroma format is formed from three squares of the i-th square group. If cr_group_chroma_format_flag[i] is equal to 0, it may indicate that the 4:4:4 or 4:2:2 target picture is not formed from the squares of the i-th square group.

[0783] If the color description present flag (cr_colour_description_present_flag[i]) is 1, it may indicate that the i-th rectangle group contains color description information derived from the VUI or explicitly defined. If cr_colour_description_present_flag[i] is 0, it may indicate that color description information is not specified for the i-th rectangle group.

[0784] If cr_group_enhanced_chroma_format_idc[i] is 1 or greater and cr_colour_description_present_flag[i] is 0, then VUI information exists within the layer containing this CR SEI message, and that VUI may precede this CR SEI message in the decoding order.

[0785] If the VUI layer ID present flag (cr_vui_layer_id_present_flag[i]) is 1, it indicates that cr_vui_layer_id[i] exists for the i-th rectangle group. If cr_vui_layer_id_present_flag[i] is 0, cr_vui_layer_id[i] does not exist for the i-th rectangle group, and the syntax elements cr_colour_primaries[i], cr_transfer_characteristics[i], cr_matrix_coefficients[i], and cr_full_range_flag[i] can be explicitly specified.

[0786] The VUI layer ID (cr_vui_layer_id[i]) can identify the layer identifier of a layer within the current bitstream that contains VUI information for the i-th rectangle group.

[0787] If cr_vui_layer_id[i] exists, VUI information exists within the layer indicated by cr_vui_layer_id[i], and that VUI may precede this CR SEI message in the decoding order.

[0788] Color primary (cr_colour_primaries[i]) has the same semantics as specified for the vui_colour_primaries syntax element, but can be applied to the i-th target picture. If cr_group_444_flag[i] is equal to 1 and cr_colour_description_present_flag[i] is equal to 0, the value of cr_colour_primaries[i] can be derived to be equal to the value of vui_colour_primaries of the VUI existing within the layer containing this SEI.

[0789] The transfer characteristic (cr_transfer_characteristics[i]) has the same semantics as specified for the vui_transfer_characteristics syntax element of the VUI existing within the layer containing this SEI, and can be applied to the i-th target picture. If cr_group_444_flag[i] is equal to 1 and cr_colour_description_present_flag[i] is equal to 0, the value of cr_transfer_characteristics[i] can be derived to be equal to the value of vui_transfer_characteristics of the VUI existing within the layer containing this SEI.

[0790] The matrix coefficients (cr_matrix_coeffs[i]) have the same semantics as specified in Section 7.3 for the vui_matrix_coeffs syntax element of the VUI existing within the layer containing this SEI, and can be applied to the i-th target picture. If cr_group_444_flag[i] is equal to 1 and cr_colour_description_present_flag[i] is equal to 0, the value of cr_matrix_coeffs[i] can be derived to be equal to the value of vui_matrix_coeffs of the VUI existing within the layer containing this SEI.

[0791] The full range flag (cr_full_range_flag[i]) has the same semantics as specified in Section 7.3 for the vui_full_range_flag syntax element of the VUI existing within the layer containing this SEI, and can be applied to the i-th target picture. If cr_group_444_flag[i] is equal to 1 and cr_colour_description_present_flag[i] is equal to 0, the value of cr_full_range_flag[i] can be derived to be equal to the value of the vui_full_range_flag of the VUI existing within the layer containing this SEI.

[0792] If the rectangle type description enable flag (cr_rect_type_descriptions_enabled_flag) is 1, it can be determined that the cr_rect_type_description_present_flag[i] syntax element is present in the SEI message. If the cr_rect_type_descriptions_enabled_flag is 0, it can be determined that the cr_rect_type_description_present_flag[i] syntax element is not present in the SEI message.

[0793] If the subpicture partitioning flag (cr_subpics_partitioning_flag[lIdx]) is 1, it may indicate that the subpicture partitioning parameters within the SPS associated with the lIdx-th layer are used to determine the size and position of the constituent rectangle. If cr_subpics_partitioning_flag[lIdx] is 0, it may indicate that the determination of the size and position of the constituent rectangle is not based on the subpicture partitioning parameters within the SPS.

[0794] If the subpicture partitioning flag (cr_subpics_partitioning_flag) is equal to 1, NumSubpics may be greater than or equal to cr_num_rects_in_layer_minus1[lIdx] + 1.

[0795] If the rectangle same size flag (cr_rect_same_size_flag[lIdx]) is 1, it indicates that all constituent rectangles within the encoded picture of the lIdx-th layer are of the same size and arranged in a grid pattern. If cr_rect_same_size_flag[lIdx] is 0, it indicates that the sizes of the constituent rectangles may differ from each other.

[0796] When the rectangle same size flag (cr_rect_same_size_flag) is equal to 1, cr_num_cols_minus1[lIdx] + 1 and cr_num_rows_minus1[lIdx] + 1 can specify the number of columns and rows of the constituent rectangle grid within the encoded picture of the lIdx-th layer, respectively. The values ​​of cr_num_cols_minus1[lIdx] and cr_num_rows_minus1[lIdx] can each be in the range of 0 to 4095.

[0797] The variable crNumCols[lIdx] can be set to be equal to cr_num_cols_minus1[lIdx] + 1.

[0798] The variable crNumRows[lIdx] can be set to be equal to cr_num_rows_minus1[lIdx] + 1.

[0799] If cr_rect_same_size_flag[lIdx] exists and is equal to 1, crNumCols[lIdx] * crNumRows[lIdx] may be equal to cr_num_rects_in_layer_minus1[lIdx] + 1.

[0800] When cr_rect_same_size_flag is equal to 1, cr_guardband_hor_size_minus1[lIdx] + 1 can specify the size of the horizontal guardband between rectangles in the encoded picture of the lIdx-th layer in luminance samples.

[0801] The variable GuardbandHor[lIdx] can be set to be equal to cr_guardbands_present_flag[lIdx] ? cr_guardband_hor_size_minus1[lIdx] + 1 : 0.

[0802] GuardbandHor[lIdx] % SubWidthC can be equal to 0.

[0803] If cr_rect_same_size_flag is equal to 1, cr_guardband_ver_size_minus1[lIdx] + 1 can specify the size of the vertical guardband between rectangles in luminous samples.

[0804] The variable GuardbandVer[lIdx] can be set to be equal to cr_guardbands_present_flag[lIdx] ? cr_guardband_ver_size_minus1[lIdx] + 1 : 0.

[0805] GuardbandVer[lIdx] % SubHeightC can be equal to 0.

[0806] The unit size (cr_log2_unit_size[lIdx]) can specify the unit size used for variable calculations on the constituent rectangle parameters of the lIdx-th layer.

[0807] The variable crUnitSize[lIdx] can be set to equal 1 << cr_log2_unit_size[lIdx].

[0808] The value of the rectangle size length (cr_rect_size_len_minus1[lIdx]) plus 1 can specify the lengths of the syntax elements cr_rect_top_left_in_units_x[lIdx][i], cr_rect_top_left_in_units_y[lIdx][i], cr_rect_width_in_units_minus1[lIdx][i], and cr_rect_height_in_units_minus1[lIdx][i].

[0809] If the rectangle type present flag (cr_rect_type_present_flag[lIdx][i]) is equal to 1, it can be determined that the cr_rect_type_idc[lIdx][i] syntax element exists in the SEI message for the lIdx-th layer. If the cr_rect_type_present_flag[lIdx][i] is equal to 0, it can be determined that the cr_rect_type_idc[lIdx][i] syntax element does not exist in the SEI message for the lIdx-th layer.

[0810] The rectangle type indicator (cr_rect_type_idc[lIdx][i]) can indicate the constituent picture type of the i-th rectangle for the lIdx-th layer from Table 10 (mapping between cr_rect_type_idc[lIdx][i] and the type of the constituent rectangle). If it does not exist and i is equal to 0, the value of cr_rect_type_idc[lIdx][i] can be derived as equal to 0. If it does not exist and i is greater than 0, cr_rect_type_idc[lIdx][i] can be derived as equal to cr_rect_type_idc[lIdx][i - 1].

[0811] [Table 10]

[0812]

[0813] For convenience of notation and terminology in this specification, variables and terms associated with color components may be referred to as Luma (or L or Y) and Chroma, regardless of the actual color representation method used, and two Chroma arrays may be referred to as Cb and Cr. The actual color representation method used may be indicated in this SEI message or VUI SEI message.

[0814] If the rectangle ID present flag (cr_rect_id_present_flag[lIdx][i]) is equal to 1, it can be determined that the cr_rect_id[lIdx][i] construct element exists within the SEI message for the lIdx-th layer. If the cr_rect_id_present_flag[lIdx][i] is equal to 0, it can be determined that the cr_rect_id[i] construct element does not exist within the SEI message for the lIdx-th layer.

[0815] Rectangle ID (cr_rect_id[lIdx][i]) may indicate the ID of the i-th rectangle of the lIdx-th layer. The length of the syntax element may be cr_rect_id_len_minus1 + 1 bit. If it does not exist and i is equal to 0, the value of cr_rect_id[lIdx][i] may be derived as 0. If it does not exist and i is greater than 0, the value of cr_rect_id[lIdx][i] may be derived as cr_rect_id[lIdx][i - 1] + 1.

[0816] If j is not equal to k, cr_rect_id[lIdx][j] may not be equal to cr_rect_id[lIdx][k].

[0817] If the associated rectangle ID present flag (cr_associated_rect_id_present_flag[lIdx][i]) is equal to 1, it can be determined that the cr_associated_rect_id[lIdx][i] syntax element exists within the SEI message for the lIdx-th layer. If the cr_associated_rect_id_present_flag[lIdx][i] is equal to 0, it can be determined that the cr_associated_rect_id[lIdx][i] syntax element does not exist within the SEI message for the lIdx-th layer.

[0818] The associated rectangle ID (cr_associated_rect_id[lIdx][i]) may indicate the ID of the primary rectangle associated with the i-th rectangle of the lIdx-th layer. The length of the syntax element may be cr_rect_id_len_minus1 + 1 bit.

[0819] A primary rectangle can be a constituent rectangle with cr_rect_type_idc equal to 0.

[0820] If present, the value of cr_associated_rect_id[i] may be equal to cr_rect_id[lIdx][j] for any value of j in the range from 0 to CrNumRects - 1, and cr_rect_type_idc[lIdx][j] may not be equal to 255.

[0821] If it does not exist, the value of cr_associated_rect_id[i] may be undefined.

[0822] The rectangle group ID (cr_rect_group_id[lIdx][i]) may indicate the group ID of the i-th rectangle of the lIdx-th layer. The length of the syntax element may be Ceil(Log2(cr_num_groups)) bits. If it does not exist and cr_rect_type_idc[lIdx][i] is not equal to 255, the value of cr_rect_group_id[lIdx][i] may be derived as 0.

[0823] If the rectangle type indicator (cr_rect_type_idc[lIdx][i]) is within the range of 4 .. 6, cr_group_enhanced_chroma_format_idc[cr_rect_group_id[i]] may be greater than 0.

[0824] If cr_group_enhanced_chroma_format_idc[cr_rect_group_id[i]] is greater than 0, for j within the range 0 .. cr_num_groups - 1, all of the following may apply.

[0825] Within the range of 0 .. CrNumRects - 1, there may be exactly one y value such that cr_rect_group_id[lIdx][y] is equal to j and cr_rect_type_idc[lIdx][y] is equal to 0 or 4.

[0826] Within the range of 0 .. CrNumRects - 1, there may be exactly one value of u such that cr_rect_group_id[lIdx][u] is equal to j and cr_rect_type_idc[lIdx][u] is equal to 5.

[0827] Within the range of 0 .. CrNumRects - 1, there may be exactly one value of v such that cr_rect_group_id[lIdx][v] is equal to j and cr_rect_type_idc[lIdx][v] is equal to 6.

[0828] If cr_group_enhanced_chroma_format_idc[i] is equal to 2, the following may apply.

[0829] The values ​​of crRectWidth[y], crRectWidth[u], and crRectWidth[v] can be the same.

[0830] The values ​​of crRectHeight[y], crRectHeight[u], and crRectHeight[v] can be equal to each other.

[0831] Otherwise, if cr_group_enhanced_chroma_format_idc[i] is equal to 1, the following may apply.

[0832] The values ​​of crRectWidth[u] and crRectWidth[v] can be equal to crRectWidth[y] / 2.

[0833] The values ​​of crRectHeight[u] and crRectHeight[v] can be equal to crRectHeight[y] / 2.

[0834] The j-th target picture can be created with a width equal to crRectWidth[y], a height equal to crRectHeight[y], a SubWidthC equal to 1, a SubHeightC equal to 1, and a ChromaFormatIdc equal to 3, and the luminance samples can be set to be equal to the luminance samples of the y-th rectangle, the Cb samples can be set to be equal to the luminance samples of the u-th rectangle, and the Cr samples can be set to be equal to the luminance samples of the v-th rectangle.

[0835] If cr_rect_type_description_present_flag[lIdx][i] is equal to 1, it can be determined that the cr_rect_type_description[lIdx][i] syntax element exists within the SEI message for the lIdx-th layer. If cr_rect_type_description_present_flag[lIdx][i] is equal to 0, it can be determined that the cr_rect_type_description[lIdx][i] syntax element does not exist within the SEI message for the lIdx-th layer. If it does not exist, the value of cr_rect_type_description_present_flag[lIdx][i] can be derived to be equal to 0.

[0836] cr_subpic_idx[lIdx][crIdx], if present, can specify the index of the subpicture of the lIdx-th layer associated with the i-th constituent rectangle unit. The number of bits for signaling cr_subpic_idx[lIdx][crIdx] can be cr_subpic_idx_len_minus1 + 1. If not present, the value of cr_subpic_idx[lIdx][crIdx] can be derived as crIdx.

[0837] cr_rect_top_left_in_units_x[lIdx][crIdx] and cr_rect_top_left_in_units_y[lIdx][crIdx], if present, may indicate the horizontal and vertical positions of the top-left position of the crIdx-th constituent rectangle within the lIdx-th layer picture, respectively, in units. The length of the syntax elements may be cr_rect_size_len_minus1[lIdx] + 1.

[0838] The variables SubWidthC and SubHeightC can be derived from ChromaFormatIdc.

[0839] The value of the rectangle width unit (cr_rect_width_in_units_minus1[lIdx][crIdx]) plus 1 and the value of the cr_rect_height_in_units_minus1[lIdx][crIdx] plus 1, if present, may indicate the width and height of the crIdx-th constituent rectangle in the lIdx-th layer picture, respectively, in units. The length of the syntax elements may be cr_rect_size_len_minus1 + 1.

[0840] The variables crRectTopLeftX[lIdx][crIdx] and crRectTopLeftY[lIdx][crIdx], representing the x and y positions respectively, and the variables crRectWidth[lIdx][crIdx] and crRectHeight[lIdx][crIdx], representing the width and height of the i-th constituent rectangle, i.e., the crIdx-th constituent rectangle within the lIdx-th layer picture respectively, can be derived as follows.

[0841] If the subpicture partitioning flag (cr_subpics_partitioning_flag) is equal to 0 and the cr_rect_same_size_flag is equal to 0, the following may apply.

[0842] The variable crRectTopLeftX[lIdx][crIdx] can be set to be equal to cr_rect_top_left_in_units_x[lIdx][crIdx] * crUnitSize[lIdx].

[0843] The variable crRectTopLeftY[lIdx][crIdx] can be set to be equal to cr_rect_top_left_in_units_y[lIdx][crIdx] * crUnitSize[lIdx].

[0844] The variable crRectWidth[lIdx][crIdx] can be set to be equal to (cr_rect_width_in_units_minus1[lIdx][crIdx] + 1) * crUnitSize[lIdx].

[0845] The variable crRectHeight[lIdx][crIdx] can be set to be equal to (cr_rect_height_in_units_minus1[lIdx][crIdx] + 1) * crUnitSize[lIdx].

[0846] Otherwise, if cr_subpics_partitioning_flag is equal to 1, the following may apply.

[0847] The variable crRectTopLeftX[lIdx][crIdx] can be set to be equal to SubPicTopLeftX[lIdx][cr_subpic_idx[lIdx][crIdx]].

[0848] The variable crRectTopLeftY[lIdx][crIdx] can be set to be equal to SubPicTopLeftY[lIdx][cr_subpic_idx[lIdx][crIdx]].

[0849] The variable crRectWidth[lIdx][crIdx] can be set to be equal to SubPicWidth[lIdx][cr_subpic_idx[lIdx][crIdx]].

[0850] The variable crRectHeight[lIdx][crIdx] can be set to be equal to SubPicHeight[lIdx][cr_subpic_idx[lIdx][crIdx]].

[0851] Otherwise, that is, if cr_rect_same_size_flag is equal to 1, the following may apply.

[0852] The variable currRow[lIdx] can be set to be equal to i % crNumCols[lIdx].

[0853] The variable currCol[lIdx] can be set to be equal to i / crNumCols[lIdx].

[0854] The variable crRectWidth[lIdx][crIdx] can be set to be equal to (MaxPicWidth[lId] - GuardbandHor[lIdx] * (crNumCols[lIdx] - 1)) / crNumCols[lIdx].

[0855] The variable crRectHeight[lIdx][crIdx] can be set to equal (MaxPicHeight[lId] - GuardbandVer[lIdx] * (crNumRows[lIdx] - 1)) / crNumRows[lIdx].

[0856] The variable crRectTopLeftX[lIdx][crIdx] can be set to be equal to currCol[lIdx] * (crRectWidth[lIdx] + GuardbandHor[lIdx]).

[0857] The variable crRectTopLeftY[lIdx][crIdx] can be set to be equal to currRow[lIdx] * (crRectHeight[lIdx] + GuardbandVer[lIdx]).

[0858] If PicWidthInLumaSamples[lId] is not equal to MaxPicWidth[lId], the following may apply.

[0859] crRectTopLeftX[lIdx][crIdx] can be set to be equal to (crRectTopLeftX[lIdx][crIdx] * PicWidthInLumaSamples[lId] + MaxPicWidth[lId] / 2) / MaxPicWidth[lId].

[0860] crRectWidth[lIdx][crIdx] can be set to be equal to (crRectWidth[lIdx][crIdx] * PicWidthInLumaSamples[lId] + MaxPicWidth[lId] / 2) / MaxPicWidth[lId].

[0861] If PicHeightInLumaSamples[lId] is not equal to MaxPicHeight[lId], the following may apply.

[0862] crRectTopLeftY[lIdx][crIdx] can be set to be equal to (crRectTopLeftY[lIdx][crIdx] * PicHeightInLumaSamples[lId] + MaxPicHeight[lId] / 2) / MaxPicHeight[lId].

[0863] crRectHeight[lIdx][crIdx] can be set to be equal to (crRectHeight[lIdx][crIdx] * PicHeightInLumaSamples[lId] + MaxPicHeight[lId] / 2) / MaxPicHeight[lId].

[0864] As a requirement for bitstream conformity, for each sample position (x, y) in the encoded picture, there may be at most one rectangle j to which all of the following conditions apply.

[0865] x may be within the range of (crRectTopLeftX[lIdx][j] .. crRectTopLeftX[lIdx][j] + crRectWidth[lIdx][j] - 1).

[0866] y may be within the range of (crRectTopLeftY[lIdx][j] .. crRectTopLeftY[lIdx][j] + crRectHeight[lIdx][i] - 1).

[0867] As a requirement for bitstream conformity, crRectWidth[lIdx][crIdx] and crRectHeight[lIdx][crIdx] may be greater than 0.

[0868] As a requirement for bitstream conformity, crRectTopLeftX[lIdx][crIdx] + crRectWidth[lIdx][crIdx] may be less than or equal to MaxPicWidth[lId], and crRectTopLeftY[lIdx][crIdx] + crRectHeight[lIdx][crIdx] may be less than or equal to MaxPicHeight[lId].

[0869] As a requirement for bitstream conformance, crRectTopLeftX[lIdx][crIdx] % SubWidthC may be equal to 0, crRectTopLeftY[lIdx][crIdx] % SubHeightC may be equal to 0, crRectWidth[lIdx][crIdx] % SubWidthC may be equal to 0, and crRectHeight[lIdx][crIdx] % SubHeightC may be equal to 0.

[0870] cr_bit_equal_to_zero can be equal to 0.

[0871] cr_rect_type_description[lIdx][i] can specify the text description of the constituent rectangle. The length of the syntax element can be 4097 bytes or less, excluding the null termination byte.

[0872] FIGS. 28a and FIGS. 28 illustrate the display rectangles SEI message syntax according to the embodiments.

[0873] The Display rectangles SEI message may describe one or more display rectangles that can be selected to be used to form a recommended display picture from a cropped decoding picture.

[0874] To use the Display rectangles SEI message, the definition of the following variables may be required.

[0875] The width and height of the cropped decoding output picture in luma samples, referred to as CroppedWidth and CroppedHeight respectively, can be defined.

[0876] A chroma format specifier referred to as ChromaFormatIdc can be defined.

[0877] A cropped decoding picture array DecodedPicture[cIdx][x][y] can be defined, having cIdx = 0..((ChromaFormatIdc == 0) ? 0 : 2), x = 0..((cIdx == 0) ? CroppedWidth : CroppedWidth / SubWidthC - 1), and y = 0..((cIdx == 0) CroppedHeight : CroppedHeight / SubHeightC - 1).

[0878] A bit depth for samples of a luminance component referred to as BitDepthY, and a bit depth for samples of two associated chroma components referred to as BitDepthC when ChromaFormatIdc is not equal to 0 may be defined.

[0879] The variables SubWidthC and SubHeightC can be derived from ChromaFormatIdc.

[0880] If dr_cancel_flag is 1, the SEI message may indicate that the persistence of any previous DR SEI message is canceled. If dr_cancel_flag is 0, the display rectangles information may be followed.

[0881] The DR SEI message can be applied to the current picture and can be persisted to all subsequent pictures in output order until one or more of the following conditions become true.

[0882] New CLVS can be started.

[0883] The bitstream can be terminated.

[0884] A picture having a DR SEI message can be output from the current layer, and that picture can follow the current picture in the output order.

[0885] If dr_pos_updates_flag is 0, it may indicate that the SEI message provides descriptions of one or more display rectangles used to form display pictures from a cropped decoding picture. If dr_pos_updates_flag is 1, it may indicate that the SEI message includes updates for one or more display rectangles. The dr_pos_updates_flag value of the first display rectangles SEI in CLVS may be 0.

[0886] dr_log2_unit_size can specify the size of the display rectangles and the granularity of their positions.

[0887] The variable DrUnitSize can be derived as follows.

[0888] DrUnitSize = 1 << dr_log2_unit_size

[0889] As a requirement for bitstream conformity, DrUnitSize % SubWidthC and DrUnitSize % SubHeightC may be equal to 0.

[0890] If dr_pos_updates_flag is equal to 0, the variable DrUnitSizeInit can be set to equal DrUnitSize.

[0891] dr_rect_size_len_minus1 + 1 can specify the lengths of the syntax elements dr_rect_top_left_x[i], dr_rect_top_left_y[i], dr_rect_width[i], dr_rect_height[i], dr_fill_top_left_x[i], dr_fill_top_left_y[i], dr_fill_width[i], and dr_fill_height[i].

[0892] dr_num_rect_minus1 + 1 can specify the number of display rectangles described in the SEI message.

[0893] dr_display_aspect_ratio[i] can indicate the display aspect ratio of the i-th display rectangle as specified by Table 11 (mapping between dr_display_aspect_ratio[i] and DrDarWidth and DrDarHeight).

[0894] [Table 11]

[0895]

[0896] Rectangle width (dr_rect_width[i]) can specify the width of the i-th rectangle in units. If it does not exist, the value of dr_rect_width[i] can be derived from the value of the corresponding syntax element in the previous DR SEI message in terms of output order.

[0897] The variable DrRectWidth can be derived as follows.

[0898] DrRectWidth[i] = DrUnitSizeInit * dr_rect_width[i]

[0899] dr_rect_height[i], if present, can specify the height of the i-th rectangle based on the DrUnitSizeInit unit size. If it does not exist and dr_display_aspect_ratio[i] is equal to 0, the value of dr_rect_height[i] can be derived to be equal to the value of the corresponding syntax element in the previous DR SEI message in terms of output order.

[0900] The variable DrRectHeight can be derived as follows.

[0901] if (dr_display_aspect_ratio[i] == 0)

[0902] DrRectHeight[i] = DrUnitSizeInit * dr_rect_height[i]

[0903] else

[0904] DrRectHeight[i] = (DrRectWidth[i] * DrDarHeight[i]) / DrDarWidth[i]

[0905] If prs_fill_method_present_flag[i] is equal to 1, it can be determined that the dr_fill_method_idc[i] syntax element exists. If prs_fill_method_present_flag[i] is equal to 0, it can be determined that the dr_fill_method_idc[i] syntax element does not exist.

[0906] dr_fill_method_idc[i], if present, may indicate the filling method applied to the i-th rectangle as specified by Table 12 (mapping of dr_fill_method_idc[i]). If not present, the value of dr_fill_method_idc[i] may be derived to be 0. In a bitstream conforming to this version of the specification, the value of dr_fill_method_idc[i] may be in the range of 0 to 2. Values ​​of dr_fill_method_idc[i] from 3 to 8 may be reserved for future use by ITU-T | ISO / IEC. Even though this version of the specification requires that the value of dr_fill_method_idc[i] be in the range of 0 to 2, decoders may allow values ​​of dr_fill_method_idc[i] in the range of 3 to 7 to appear within the syntax, in which case the filling method may not be defined.

[0907] [Table 12]

[0908]

[0909] A value of 1 for the three color component flag (dr_three_colour_comp_flag[ i ]) indicates that the dr_target_init_sample[ i ] syntax element is signaled for the three color components of the i-th square. A value of 0 for dr_three_colour_comp_flag indicates that the dr_fill_value[ i ] syntax element is signaled for the one color component.

[0910] The value obtained by adding 8 to the fill bit depth (dr_fill_bit_depth_minus8[ i ]) represents the bit depth of the fill value for the i-th display rectangle. dr_fill_bit_depth_minus8[ i ] must be a value within the range of 0 to 8, that is, a range that includes both ends.

[0911] The variable initBitDepth[ i ] is set to be equal to dr_fill_bit_depth_minus8[ i ] + 8.

[0912] The fill value (dr_fill_value[ i ][ j ]) represents the fill value for the j-th color component of the i-th rectangle. The length of the corresponding syntax element is initBitDepth[ i ] bits.

[0913] If not present, the value of dr_fill_value[0] is derived to a value equal to 0, and the values ​​of dr_fill_value[1] and dr_fill_value[2] are derived to a value equal to 1 << (BitDepth - 1).

[0914] Fill blur sigma(dr_fill_blur_sigma[ i ]) represents the sigma parameter of the Gaussian blur applied to the fill of the i-th display rectangle, if present.

[0915] The fill content width (dr_fill_content_width[ i ]) represents the width of the content area used to fill the i-th display rectangle in DrUnitSizeInit units. If it does not exist, the value of dr_fill_content_width[ i ] is set to the value of the corresponding syntax element of the previous DR SEI message in the output order.

[0916] The variable DrFillWidth[ i ] is derived as follows.

[0917] DrFillWidth[i] = DrUnitSizeInit * dr_fill_content_width[i]

[0918] The variable DrFillHeight[ i ] is derived as follows.

[0919] DrFillHeight[ i ] = ( DrFillWidth[ i ] * DrRectWidth[ i ] ) / DrRectHeight[ i ]

[0920] Rectangle top left x (dr_rect_top_left_x[ i ]) and rectangle top left y (dr_rect_top_left_y[ i ]) represent the top-left x and y coordinates of the i-th rectangle, respectively, in DrUnitSize units.

[0921] The rectangle top-left x sign (dr_rect_top_left_x_sign[ i ]) and the rectangle top-left y sign (dr_rect_top_left_y_sign[ i ]) represent the signs to be applied to dr_rect_top_left_x[ i ] and dr_rect_top_left_y[ i ], respectively.

[0922] Fill content area top left x (dr_fill_content_region_top_left_x[ i ]) and fill content area top left y (dr_fill_content_region_top_left_y[ i ]) represent the top-left x and y coordinates of the area used to fill the i-th display rectangle, respectively, in DrUnitSize units.

[0923] Extrapolate(OutPic, InPic, x0, y0, w0, h0) is defined as a function that extracts a rectangle with position (x0, y0) and luma size w0Хh0 within the InPic input picture and extrapolates it to the OutPic output picture. If OutPic has a larger dimension than the extracted area within InPic, an upscaling operation is performed.

[0924] Blur(OutPic, InPic, x0, y0, w0, h0, sigma0) is defined as a function that extracts a rectangle with position (x0, y0) and luma size w0Xh0 within the InPic input picture, scales it to the resolution of the OutPic output picture, and applies a Gaussian blur with standard deviation sigma0. If OutPic has a larger dimension than the extracted area within InPic, an upscaling operation is performed.

[0925] The selected display rectangle index selIdx can be determined by external means by comparing the aspect ratio and display resolution of the intended display with the values ​​of the dr_display_aspect_ratio[ i ], dr_rect_width[ i ], and dr_rect_height[ i ] syntax elements for each display rectangle.

[0926] When the selIdx value is selected, it is recommended to form a picture for display DrDisplay[ selIdx ] having a luma resolution DrRectWidth[ selIdx ] Х DrRectHeight[ selIdx ] and a chroma format ChromaFormatIdc as follows.

[0927] If dr_fill_method_idc[ selIdx ] is equal to 0,

[0928] For y = 0 to less than DrRectHeight[ selIdx ], and for x = 0 to less than DrRectWidth[ selIdx ],

[0929] DrDisplay[ selIdx ]

[0000] [ x ][ y ] is set to the same value as dr_fill_value[ selIdx ]

[0000] .

[0930] Also, for cIdx = 1 up to ((ChromaFormatIdc == 0) ? 1 : 3) less than,

[0931] For y = 0 to less than DrRectHeight[ selIdx ] / SubHeightC, and for x = 0 to less than DrRectWidth[ selIdx ] / SubWidthC,

[0932] DrDisplay[ selIdx ][ cIdx ][ x ][ y ] is set to the same value as dr_fill_value[ selIdx ][ cIdx ].

[0933] If dr_fill_method_idc[ selIdx ] is equal to 1,

[0934] Extrapolate(DrDisplay[ selIdx ], DecodedPicture, dr_fill_top_left_x[ selIdx ], dr_fill_top_left_y[ selIdx ], DrFillWidth[ selIdx ], DrFillHeight[ selIdx ]) is performed.

[0935] If dr_fill_method_idc[ selIdx ] is equal to 2,

[0936] Blur(DrDisplay[ selIdx ], DecodedPicture, dr_fill_top_left_x[ selIdx ], dr_fill_top_left_y[ selIdx ], DrFillWidth[ selIdx ], DrFillHeight[ selIdx ], dr_fill_blur_sigma[ selIdx ]) is performed.

[0937] DrRectTopLeftY[ selIdx ] is set to be equal to ( dr_rect_top_left_y_sign[ selIdx ] == 0 ? 1 : -1 ) * dr_rect_top_left_y[ selIdx ].

[0938] DrRectTopLeftX[ selIdx ] is set to be equal to ( dr_rect_top_left_x_sign[ selIdx ] == 0 ? 1 : -1 ) * dr_rect_top_left_x[ selIdx ].

[0939] If starting at y0 = 0, y = DrRectTopLeftY[ selIdx ] and iterating while y0 is less than DrRectHeight[ selIdx ] while increasing y0 and y respectively,

[0940] yInCodedPic is set to a value such as ( y >= 0 ) && ( y < DrRectHeight[ selIdx ] )

[0941] And starting at x0 = 0, x = DrRectTopLeftX[ selIdx ], and iterating while x0 is less than DrRectWidth[ selIdx ] while increasing x0 and x respectively,

[0942] xInCodedPic is set to a value such as ( x >= 0 ) && ( x < DrRectWidth[ selIdx ] )

[0943] In the case of yInCodedPic && xInCodedPic,

[0944] DrDisplay[ selIdx ]

[0000] [ x0 ][ y0 ] is set to the same value as DecodedPicture

[0000] [ x ][ y ].

[0945] Also, for cIdx = 1 up to (ChromaFormatIdc == 0 ? 1 : 3) less than,

[0946] If starting at y0 = 0, y = DrRectTopLeftY[ selIdx ] / SubHeightC and iterating while y0 is less than DrRectHeight[ selIdx ] / SubHeightC, y0 and y are incremented respectively

[0947] yInCodedPic is set to a value such as ( y >= 0 ) && ( y < DrRectHeight[ selIdx ] / SubHeightC )

[0948] And starting from x0 = 0, x = DrRectTopLeftX[ selIdx ] / SubWidthC, and iterating while x0 is less than DrRectWidth[ selIdx ] / SubWidthC, x0 and x respectively,

[0949] xInCodedPic is set to a value such as ( x >= 0 ) && ( x < DrRectWidth[ selIdx ] / SubWidthC )

[0950] In the case of yInCodedPic && xInCodedPic,

[0951] DrDisplay[ selIdx ][ cIdx ][ x0 ][ y0 ] is set to the same value as DecodedPicture[ cIdx ][ x ][ y ].

[0952] if( dr_fill_method_idc[ selIdx ] = = 0 ) {

[0953] for( y = 0; y < DrRectHeight [ selIdx ]; y++ )

[0954] for( x = 0; x < DrRectWidth [ selIdx ]; x++ )

[0955] DrDisplay [ selIdx ]

[0000] [ x ][ y ] = dr_fill_value[ selIdx ]

[0000]

[0956]

[0957] for( cIdx = 1; cIdx < ((ChromaFormatIdc = = 0 ) ? 1 : 3 ); cIdx++ ) {

[0958] for( y = 0; y < DrRectHeight [ selIdx ] / SubHeightC; y++ )

[0959] for( x = 0; x < DrRectWidth [ selIdx ] / SubWidthC; x++ )

[0960] DrDisplay [ selIdx ][ cIdx ][ x ][ y ] = dr_fill_value[ selIdx ][ cIdx ]

[0961] } else if( dr_fill_method_idc[ selIdx ] = = 1 ) {

[0962] Extrapolate(DrDisplay [ selIdx ], DecodedPicture, dr_fill_top_left_x[ selIdx ], dr_fill_top_left_y[ selIdx ],

[0963] DrFillWidth[ selIdx ], DrFillHeight[ selIdx ] )

[0964] } else if( dr_fill_method_idc[ selIdx ] = = 2 ) {

[0965] Blur(DrDisplay [ selIdx ], DecodedPicture, dr_fill_top_left_x[ selIdx ], dr_fill_top_left_y[ selIdx ],

[0966] DrFillWidth[ selIdx ], DrFillHeight[ selIdx ], dr_fill_blur_sigma[ selIdx ] )

[0967] }

[0968] DrRectTopLeftY [ selIdx ] = ( dr_rect_top_left_y_sign[ selIdx ] = = 0 ? 1 : -1 ) * dr_rect_top_left_y [ selIdx ]

[0969] DrRectTopLeftX [ selIdx ] = ( dr_rect_top_left_x_sign[ selIdx ] = = 0 ? 1 : -1 ) * dr_rect_top_left_x [ selIdx ]

[0970] for( y0 = 0, y = DrRectTopLeftY [ selIdx ]; y0 < DrRectHeight [ selIdx ]; y0++, y++ ) {

[0971] yInCodedPic = ( y > = 0 ) && ( y < DrRectHeight [ selIdx ] )

[0972] for( x0 = 0, x = DrRectTopLeftX [ selIdx ]; x0 < DrRectWidth [ selIdx ]; x0++, x++ ) {

[0973] xInCodedPic = ( x > = 0 ) && ( x < DrRectWidth [ selIdx ] )

[0974] if ( yInCodedPic & & xInCodedPic )

[0975] DrDisplay [ selIdx ]

[0000] [ x0 ][ y0 ] = DecodedPicture

[0000] [ x ][ y ]

[0976] }

[0977] for( cIdx = 1; cIdx < (ChromaFormatIdc = = 0 ? 1 : 3); cIdx++ ) {

[0978] for( y0 = 0, y = DrRectTopLeftY [ selIdx ] ] / SubHeightC; y0 < DrRectHeight [ selIdx ] ] / SubHeightC;

[0979] y0++, y++ ) {

[0980] yInCodedPic = ( y > = 0 ) & & ( y < DrRectHeight [ selIdx ] / SubHeightC )

[0981] for( x0 = 0, x = DrRectTopLeftX [ selIdx ] / SubWidthC; x0 < DrRectWidth [ selIdx ] / SubWidthC;

[0982] x0++, x++ ) {

[0983] xInCodedPic = ( x > = 0 ) & & ( x < DrRectWidth [ selIdx ] / SubWidthC )

[0984] if ( yInCodedPic & & xInCodedPic )

[0985] DrDisplay [ selIdx ][ cIdx ][ x0 ][ y0 ] = DecodedPicture [ cIdx ][ x ][ y ]

[0986] }

[0987] }

[0988] }

[0989] }

[0990] The problems that the embodiments aim to solve are as follows.

[0991] Display rectangles SEI messages are currently included in TuC (JVET-AK2032) for future extensions of VSEI. The current design of display rectangles SEI messages allows DR SEI messages to contain only updates to the previous DR SEI message. To this end, a flag (i.e., dr_pos_updates_flag) exists to indicate whether the SEI message contains updates to one or more display rectangles.

[0992] It is argued that the current update mechanism can still require updates for all display rectangles within an SEI message even when only one or a few display rectangles need to be updated. It may be desirable to have a flag for each display rectangle indicating whether that display rectangle is updated.

[0993] The description according to the embodiments is as follows.

[0994] The embodiments provide solutions to the aforementioned problems. Each item of the embodiment may be applied individually or in combination.

[0995] Signals a syntax element (i.e., dr_rect_pos_update_flag[ i ]) to specify whether the display rectangle position information of the i-th rectangle is updated.

[0996] Alternatively, if dr_pos_updates_flag is a value equal to 1, signal a syntax element (i.e., dr_rect_pos_update_flag[ i ]) to specify whether the display rectangle position information of the i-th rectangle is updated.

[0997] Example 1:

[0998] FIG. 29 shows the display rectangles SEI message syntax according to the embodiments.

[0999] In the display rectangles SEI message syntax, for each value of i from 0 to dr_num_rect_minus1, dr_rect_pos_update_flag[ i ] may be signaled. dr_rect_pos_update_flag[ i ] may indicate whether the display rectangle position information of the i-th rectangle is updated.

[1000] If the rectangle position update flag (dr_rect_pos_update_flag[ i ]) exists, position information for the i-th rectangle may be signaled. For example, dr_rect_top_left_x[ i ], dr_rect_top_left_x_sign[ i ], dr_rect_top_left_y[ i ], and dr_rect_top_left_y_sign[ i ] may be signaled. dr_rect_top_left_x[ i ] and dr_rect_top_left_y[ i ] may represent the top-left x and y coordinates of the i-th rectangle, respectively, and dr_rect_top_left_x_sign[ i ] and dr_rect_top_left_y_sign[ i ] may represent the signs applied to the x and y coordinates, respectively.

[1001] If dr_fill_method_idc[ i ] is equal to 2, dr_fill_content_region_top_left_x[ i ] and dr_fill_content_region_top_left_y[ i ] may be additionally signaled as information related to the top-left position of the content area used to fill the i-th rectangle.

[1002] DR SEI Message Semantics (Display rectangles SEI message semantics)

[1003] A value of 1 for the rectangle position update flag (dr_rect_pos_update_flag[ i ]) indicates that the display rectangle position information of the i-th rectangle is updated. A value of 0 for dr_rect_pos_update_flag[ i ] indicates that there is no change in the display rectangle position information of the i-th rectangle. If dr_pos_updates_flag is 0, the value of dr_rect_pos_update_flag[ i ] must be 1.

[1004] The top-left x (dr_rect_top_left_x[ i ]) and top-left y (dr_rect_top_left_y[ i ]) represent the top-left x and y coordinates of the i-th rectangle, respectively, in units of DrUnitSize. If they do not exist, the values ​​of dr_rect_top_left_x[ i ] and dr_rect_top_left_y[ i ] are set to the values ​​of the corresponding syntax elements of the previous DR SEI message in the output order, respectively.

[1005] The top-left x sign (dr_rect_top_left_x_sign[ i ]) and the top-left y sign (dr_rect_top_left_y_sign[ i ]) represent the signs to be applied to dr_rect_top_left_x[ i ] and dr_rect_top_left_y[ i ], respectively. If they do not exist, the values ​​of dr_rect_top_left_x_sign[ i ] and dr_rect_top_left_y_sign[ i ] are set to the values ​​of the corresponding syntax elements of the previous DR SEI message in the output order, respectively.

[1006] Example 2:

[1007] FIGS. 30a and FIG. 30b show display rectangles SEI message syntax according to embodiments.

[1008] According to the embodiments, the display rectangle SEI message may signal information for a plurality of rectangles via dr_num_rect_minus1, and information for each rectangle may be transmitted through an iteration statement from i=0 to i<=dr_num_rect_minus1. According to the embodiments, when dr_pos_updates_flag is true, dr_rect_pos_update_flag[i] may be signaled in the form of u(1) for each i-th rectangle. Additionally, when dr_rect_pos_update_flag[i] is true, dr_rect_top_left_x[i], dr_rect_top_left_x_sign[i], dr_rect_top_left_y[i], and dr_rect_top_left_y_sign[i] may be signaled in relation to the top-left position information of the i-th rectangle. Here, dr_rect_top_left_x[i] and dr_rect_top_left_y[i] can each be encoded in the form of u(v), and dr_rect_top_left_x_sign[i] and dr_rect_top_left_y_sign[i] can each be encoded in the form of u(1). Meanwhile, according to the embodiments, if the value of dr_fill_method_idc[i] for the i-th rectangle is equal to 2, dr_fill_content_region_top_left_x[i] and dr_fill_content_region_top_left_y[i] related to the location information of the fill content region can be additionally signaled, and these syntax elements can each be encoded in the form of u(v).Accordingly, the display rectangle SEI message according to the embodiments selectively transmits coordinate information depending on whether the position of each rectangle is updated, and additionally transmits the top-left position information of the content area corresponding to a specific fill method, thereby enabling the receiving device to more accurately identify the display rectangle and the related content area.

[1009] A rectangle position update flag (dr_rect_pos_update_flag[i]) of 1 indicates that the display rectangle position information of the i-th rectangle is updated, and a dr_rect_pos_update_flag[i] of 0 indicates that there is no change in the display rectangle position information of the i-th rectangle. If dr_rect_pos_update_flag[i] does not exist, the value of dr_rect_pos_update_flag[i] is derived to 1.

[1010] The top-left x (dr_rect_top_left_x[i]) and top-left y (dr_rect_top_left_y[i]) represent the top-left x and y coordinates of the i-th rectangle, respectively, in units of DrUnitSize. If dr_rect_top_left_x[i] and dr_rect_top_left_y[i] do not exist, their values ​​are set to the values ​​of the corresponding syntax elements within the previous DR SEI message in the output order. dr_rect_top_left_x_sign[i] and dr_rect_top_left_y_sign[i] represent the signs to be applied to dr_rect_top_left_x[i] and dr_rect_top_left_y[i], respectively.

[1011] If the top-left x sign (dr_rect_top_left_x_sign[i]) and the top-left y sign (dr_rect_top_left_y_sign[i]) do not exist, the values ​​of dr_rect_top_left_x_sign[i] and dr_rect_top_left_y_sign[i] are each set to the values ​​of the corresponding syntax elements in the previous DR SEI message in the output order.

[1012] FIG. 31 shows a encoding method according to embodiments.

[1013] An encoding method according to embodiments (source device of FIG. 1, encoder of FIG. 2, transformer of encoder of FIG. 15, LFNST of FIG. 16, CABAC encoding of FIG. 17, entropy encoding of FIG. 18, picture encoding of FIG. 21, coding layer of FIG. 22, SEI message generation of FIG. 23 to 30a and FIG. 30b, method of FIG. 31, etc.) may include the step of generating at least one SEI message (S3100); the step of encoding a picture (S3110); and / or the step of generating a bitstream including at least one SEI message and said picture (S3120); etc.

[1014] At least one SEI message generated by the step (S310) of generating at least one SEI message includes a display rectangles SEI message, and the display rectangles SEI message may include a rectangle position update flag (dr_rect_pos_update_flag) indicating whether the display rectangle position information of the i-th rectangle is updated.

[1015] The display rectangles SEI message includes a rectangle area update flag (dr_pos_updates_flag) indicating whether the SEI message includes updates for one or more display rectangles, and a flag (dr_rect_pos_update_flag[ i ]) indicating whether the display rectangle position information of the i-th rectangle is updated based on the first value (1) of the rectangle area update flag (dr_pos_updates_flag) may be included in the display rectangles SEI message.

[1016] The method of FIG. 31 can be performed by a device. A device according to embodiments includes a memory; and at least one processor connected to the memory; and the at least one processor may be configured to: generate at least one SEI message; encode a picture; and generate a bitstream comprising at least one SEI message and a picture.

[1017] The embodiments further include a computer-readable storage medium for storing a bitstream generated by the method according to FIG. 31.

[1018] The embodiments further include a method comprising the steps of acquiring a bitstream, generating at least one SEI message based on the bitstream; encoding a picture based on the bitstream including at least one SEI message and the picture; and transmitting data including the bitstream.

[1019] FIG. 32 illustrates a decoding method according to embodiments.

[1020] A decoding method according to embodiments (a receiving device of FIG. 1, a decoder of FIG. 3, an inverse transformer of the decoder of FIG. 15, an LFNST of FIG. 16, an entropy decoding of FIG. 19, a picture decoding of FIG. 20, a coding layer of FIG. 22, acquisition of SEI messages from FIG. 23 to FIG. 30a and FIG. 30b, method of FIG. 32, etc.) may include the step of acquiring a bitstream (S3200); the step of acquiring at least one SEI message from the bitstream (S3210); and / or the step of decoding a picture within the bitstream (S3220); etc.

[1021] The method in Fig. 32 can follow the reverse process of the method in Fig. 31.

[1022] Referring to Example 1, at least one SEI message includes a display rectangles SEI message, and the display rectangles SEI message may include a rectangle position update flag (dr_rect_pos_update_flag) indicating whether the display rectangle position information of the i-th rectangle is updated.

[1023] Referring to Example 2, at least one SEI message includes a display rectangles SEI message, and the display rectangles SEI message includes a rectangle area update flag (dr_pos_updates_flag) indicating whether the SEI message includes updates for one or more display rectangles, and a flag (dr_rect_pos_update_flag[ i ]) indicating whether the display rectangle position information of the i-th rectangle is updated based on a first value (1) of the rectangle area update flag (dr_pos_updates_flag) may be included in the display rectangles SEI message.

[1024] Referring to Example 2, the step of obtaining at least one SEI message (S3210) includes: the step of obtaining a display rectangles SEI message, wherein the step of obtaining the display rectangles SEI message includes: for a plurality of rectangles associated with the display rectangle, if dr_pos_updates_flag is 1, obtaining dr_rect_pos_update_flag[i] for the i-th rectangle, and if dr_rect_pos_update_flag[i] is true, obtaining dr_rect_top_left_x[i], dr_rect_top_left_x_sign[i], dr_rect_top_left_y[i] and dr_rect_top_left_y_sign[i] for the i-th rectangle, and if dr_fill_method_idc[i] is equal to 2, obtaining dr_fill_content_region_top_left_x[i] and for the i-th rectangle It may include a step of obtaining dr_fill_content_region_top_left_y[i].

[1025] Additionally, when dr_pos_updates_flag according to the embodiments is 0, the bitstream according to the embodiments does not contain the dr_rect_pos_update_flag[i] syntax element, but the value of dr_rect_pos_update_flag[i] can be derived to 1. This means that when dr_pos_updates_flag according to the embodiments is 0, the bitstream according to the embodiments does not contain the dr_rect_pos_update_flag[i] syntax element, but the value of dr_rect_pos_update_flag[i] can be inferred to be 1.

[1026] Referring to Example 2, if the value of dr_pos_updates_flag is 0, it indicates that the display rectangle SEI message provides a description of one or more display rectangles used to form a display picture from a cropped decoded picture; if the value of dr_pos_updates_flag is 1, it indicates that the display rectangle SEI message includes an update for one or more display rectangles; if the value of dr_rect_pos_update_flag[i] is 1, it indicates that the display rectangle position information of the i-th rectangle is updated; and if the value of dr_rect_pos_update_flag[i] is 0, it indicates that the display rectangle position information of the i-th rectangle is not changed.

[1027] Referring to Example 2, dr_rect_top_left_x[i] and dr_rect_top_left_y[i] represent the top-left x and y coordinates of the i-th rectangle within the display rectangle unit size (DrUnitSize), and if dr_rect_top_left_x[i] and dr_rect_top_left_y[i] exist, dr_rect_top_left_x[i] and dr_rect_top_left_y[i] may represent the top-left x and y coordinates of the i-th rectangle, respectively.

[1028] If dr_rect_top_left_x[i] and dr_rect_top_left_y[i] do not exist, the values ​​of dr_rect_top_left_x[i] and dr_rect_top_left_y[i] can be derived from the values ​​of syntax elements related to the previous display rectangle SEI message in the output order.

[1029] Referring to Example 2, if dr_rect_top_left_x_sign[i] and dr_rect_top_left_y_sign[i] exist, the signs applied to dr_rect_top_left_x[i] and dr_rect_top_left_y[i] are respectively indicated, and if dr_rect_top_left_x_sign[i] and dr_rect_top_left_y_sign[i] do not exist, the values ​​of dr_rect_top_left_x_sign[i] and dr_rect_top_left_y_sign[i] can be derived from the values ​​of the syntax elements related to the previous display rectangle SEI message in the output order.

[1030] Referring to Example 2, if the value of dr_fill_method_idc[i] is 0, the filling method applied to the i-th rectangle is a constant value-based filling method; if the value of dr_fill_method_idc[i] is 1, the filling method is a spatial extrapolation-based filling method; and if the value of dr_fill_method_idc[i] is 2, the filling method is a replication-based filling method with scaling and blurring. The step of obtaining at least one SEI message comprises: obtaining a filling method identifier dr_fill_method_idc[i] for the i-th display rectangle; Based on the fact that dr_fill_method_idc[i] is 2, the method includes the step of obtaining dr_fill_content_region_top_left_x[i] and dr_fill_content_region_top_left_y[i], which respectively represent the top-left x-coordinate and top-left y-coordinate of the area used for filling the i-th display rectangle, and dr_fill_content_region_top_left_x[i] and dr_fill_content_region_top_left_y[i] can respectively represent the top-left x-coordinate and top-left y-coordinate of the area used for filling the i-th display rectangle within the unit size of the display rectangle (DrUnitSize).

[1031] The method of FIG. 32 can be performed by a device. A device according to embodiments includes a memory; and at least one processor connected to the memory; and the at least one processor may be configured to: acquire a bitstream; acquire at least one SEI message from the bitstream; and decode a picture within the bitstream.

[1032] The method and apparatus according to the embodiments provide the following technical effects.

[1033] According to the embodiments, by individually providing information indicating whether an update is required for each display rectangle (DR), relevant location information, size information, or other parameter information may not be repeatedly transmitted for DRs that have not changed. Accordingly, update information can be selectively transmitted only for DRs that have actually changed, thereby reducing unnecessary signaling overhead and improving the encoding efficiency of the bitstream. In addition, since the receiving side can distinguish and process the changed DR and the maintained DR based on the update information, the interpretation and restoration of DR-related information can be performed more clearly and efficiently.

[1034] The embodiments have been described in terms of methods and / or devices, and the description of the methods and the description of the devices may be applied complementarily.

[1035] Although the drawings have been described separately for the convenience of explanation, it is also possible to design a new embodiment by combining the embodiments described in each drawing. Furthermore, designing a computer-readable recording medium containing a program for executing the previously described embodiments, as required by a person skilled in the art, falls within the scope of the embodiments. The apparatus and method according to the embodiments are not limited to the configuration and method of the embodiments described above; rather, the embodiments may be configured by selectively combining all or part of each embodiment to allow for various modifications. Although preferred embodiments have been illustrated and described, the embodiments are not limited to the specific embodiments described above. It is not only possible for a person skilled in the art to make various modifications without departing from the essence of the embodiments claimed in the claims, but such modifications should not be understood individually from the technical concept or perspective of the embodiments.

[1036] Various components of the device of the embodiments may be implemented by hardware, software, firmware, or a combination thereof. Various components of the embodiments may be implemented as a single chip, for example, a single hardware circuit. Depending on the embodiments, the components according to the embodiments may each be implemented as separate chips. Depending on the embodiments, at least one of the components of the device according to the embodiments may be composed of one or more processors capable of executing one or more programs, and one or more programs may include instructions for performing or executing any one or more of the operations / methods according to the embodiments. Executable instructions for performing the methods / operations of the device according to the embodiments may be stored in non-transient CRMs or other computer program products configured to be executed by one or more processors, or may be stored in transient CRMs or other computer program products configured to be executed by one or more processors. Additionally, memory according to the embodiments may be used as a concept that includes not only volatile memory (e.g., RAM, etc.) but also non-volatile memory, flash memory, PROM, etc. In addition, it may also include implementation in the form of carrier waves, such as transmission over the Internet. Furthermore, processor-readable recording media are distributed across networked computer systems, allowing processor-readable code to be stored and executed in a distributed manner.

[1037] In this document, “ / ” and “,” are interpreted as “and / or.” For example, “A / B” is interpreted as “A and / or B,” and “A, B” is interpreted as “A and / or B.” Additionally, “A / B / C” means “at least one of A, B and / or C.” Also, “A, B, C” means “at least one of A, B and / or C.” Additionally, in this document, “or” is interpreted as “and / or.” For example, “A or B” may mean 1) “A” alone, 2) “B” alone, or 3) “A and B.” In other words, “or” in this document may mean “additionally or alternatively.”

[1038] Terms such as "first," "second," etc., may be used to describe various components of the embodiments. However, the interpretation of the various components according to the embodiments should not be limited by these terms. These terms are merely used to distinguish one component from another. For example, the first user input signal may be referred to as the second user input signal. Similarly, the second user input signal may be referred to as the first user input signal. The use of these terms should be interpreted as not departing from the scope of the various embodiments. Although the first user input signal and the second user input signal are both user input signals, they do not imply the same user input signals unless clearly indicated in the context.

[1039] The terms used to describe the embodiments are intended for the purpose of describing specific embodiments and are not intended to limit the embodiments. As used in the description of the embodiments and in the claims, the singular is intended to include the plural unless explicitly indicated in the context. Expressions of and / or are used to mean including all possible combinations between the terms. Expressions of include describe the presence of features, numbers, steps, elements, and / or components and do not imply the exclusion of additional features, numbers, steps, elements, and / or components. Conditional expressions such as "if" or "when" used to describe the embodiments are not limited to being optional. It is intended to be interpreted as "when a specific condition is satisfied," "when a related action is performed in response to a specific condition," or "when a related definition is interpreted."

[1040] Additionally, operations according to the embodiments described herein may be performed by a transmitting and receiving device including memory and / or a processor, depending on the embodiments. The memory may store programs for processing / controlling operations according to the embodiments, and the processor may control various operations described in this document. The processor may be referred to as a controller, etc. Operations in the embodiments may be performed by firmware, software, and / or a combination thereof, and the firmware, software, and / or a combination thereof may be stored in the processor or in memory.

[1041] Meanwhile, the operation according to the embodiments described above may be performed by a transmitting device and / or a receiving device according to the embodiments. The transmitting and receiving device may include a transmitting and receiving unit for transmitting and receiving media data, a memory for storing instructions (program code, algorithm, flowchart and / or data) for a process according to the embodiments, and a processor for controlling the operations of the transmitting and receiving devices.

[1042] The processor may be referred to as a controller, etc., and may correspond, for example, to hardware, software, and / or a combination thereof. The operation according to the embodiments described above may be performed by the processor. Additionally, the processor may be implemented as an encoder / decoder, etc., for the operation of the embodiments described above.

[1043] As described above, the relevant details have been explained in the best mode for carrying out the embodiments.

[1044] As described above, the embodiments may be applied wholly or partially to an image encoding method, an image encoding device, an image decoding method, an image decoding device, and a system.

[1045] Those skilled in the art may make various changes or modifications to the embodiments within the scope of the embodiments.

[1046] The embodiments may include modifications / variations, and such modifications / variations do not exceed the scope of the claims and their equivalents.

Claims

1. Step of acquiring the bitstream; A step of obtaining at least one SEI message from the bitstream; and A step of decoding a picture within the bitstream; comprising method.

2. In Paragraph 1, The above at least one SEI message includes a display rectangles SEI message, and The above display rectangles SEI message is: including a rectangle position update flag (dr_rect_pos_update_flag) indicating whether the display rectangle position information of the i-th rectangle is updated, method.

3. In Paragraph 1, The above at least one SEI message includes a display rectangles SEI message, and The above display rectangles SEI message is: It includes a rectangle area update flag (dr_pos_updates_flag) indicating whether the SEI message includes updates for one or more display rectangles, and Based on the first value (1) of the above rectangular area update flag (dr_pos_updates_flag), A flag (dr_rect_pos_update_flag[ i ]) indicating whether the display rectangle position information of the i-th rectangle is updated is included in the display rectangles SEI message, method.

4. In Paragraph 1, The step of acquiring at least one SEI message above is: The method includes the step of obtaining a display rectangles SEI message, The step of obtaining the above display rectangles SEI message is: For multiple rectangles related to the display rectangle, If dr_pos_updates_flag is 1, obtain dr_rect_pos_update_flag[i] for the i-th rectangle, and If the above dr_rect_pos_update_flag[i] is true, obtain dr_rect_top_left_x[i], dr_rect_top_left_x_sign[i], dr_rect_top_left_y[i], and dr_rect_top_left_y_sign[i] for the i-th rectangle, and If dr_fill_method_idc[i] is equal to 2, the step of obtaining dr_fill_content_region_top_left_x[i] and dr_fill_content_region_top_left_y[i] for the i-th rectangle is included, Based on the case where the above dr_pos_updates_flag is 0, the value of the above dr_rect_pos_update_flag[i] is induced to be 1, method.

5. In Paragraph 4, If the value of the above dr_pos_updates_flag is 0, it indicates that the display rectangle SEI message provides a description of one or more display rectangles used to form a display picture from a cropped decoded picture, and If the value of the above dr_pos_updates_flag is 1, it indicates that the display rectangle SEI message includes updates for one or more display rectangles, and If the value of dr_rect_pos_update_flag[i] above is 1, it indicates that the display rectangle position information of the i-th rectangle is updated, and If the value of dr_rect_pos_update_flag[i] is 0, it indicates that the display rectangle position information of the i-th rectangle is not changed, method.

6. In Paragraph 4, The above dr_rect_top_left_x[i] and the above dr_rect_top_left_y[i] represent the top-left x and y coordinates of the i-th rectangle within the display rectangle unit size (DrUnitSize), and If the above dr_rect_top_left_x[i] and the above dr_rect_top_left_y[i] exist, the above dr_rect_top_left_x[i] and the above dr_rect_top_left_y[i] represent the top-left x and y coordinates of the i-th rectangle, respectively, and If the above dr_rect_top_left_x[i] and the above dr_rect_top_left_y[i] do not exist, the values ​​of the above dr_rect_top_left_x[i] and the above dr_rect_top_left_y[i] are respectively derived from the values ​​of the syntax elements related to the previous display rectangle SEI message in the output order, method.

7. In Paragraph 4, If the above dr_rect_top_left_x_sign[i] and the above dr_rect_top_left_y_sign[i] exist, the signs applied to the above dr_rect_top_left_x[i] and the above dr_rect_top_left_y[i] are respectively indicated, and If the above dr_rect_top_left_x_sign[i] and the above dr_rect_top_left_y_sign[i] do not exist, the values ​​of the above dr_rect_top_left_x_sign[i] and the above dr_rect_top_left_y_sign[i] are respectively derived from the values ​​of the syntax elements related to the previous display rectangle SEI message in the output order, method.

8. In Paragraph 4, If the value of dr_fill_method_idc[i] above is 0, the filling method applied to the i-th rectangle is a constant value-based filling method, and If the value of dr_fill_method_idc[i] is 1, the filling method is a spatial extrapolation-based filling method, and When the value of dr_fill_method_idc[i] is 2, the filling method is a replication-based filling method using scaling and blurring, and The step of obtaining at least one SEI message above comprises: obtaining a filling method identifier dr_fill_method_idc[i] for the i-th display rectangle; and Based on the fact that dr_fill_method_idc[i] is 2, the method includes the step of obtaining dr_fill_content_region_top_left_x[i] and dr_fill_content_region_top_left_y[i], which respectively represent the top-left x-coordinate and top-left y-coordinate of the area used for filling the i-th display rectangle, and The above dr_fill_content_region_top_left_x[i] and the above dr_fill_content_region_top_left_y[i] respectively represent the top-left x-coordinate and top-left y-coordinate of the area used for filling the i-th display rectangle within the display rectangle unit size (DrUnitSize). method.

9. Memory; and At least one processor connected to the memory; comprising, wherein the at least one processor: Acquire bitstream; Acquire at least one SEI message from the bitstream above; and Configured to decode a picture within the above bitstream, device.

10. A step of generating at least one SEI message; Step of encoding a picture; and A step of generating a bitstream including at least one SEI message and the picture; comprising method.

11. In Paragraph 10, The above at least one SEI message includes a display rectangles SEI message, and The above display rectangles SEI message is: including a rectangle position update flag (dr_rect_pos_update_flag) indicating whether the display rectangle position information of the i-th rectangle is updated, method.

12. In Paragraph 10, The above at least one SEI message includes a display rectangles SEI message, and The above display rectangles SEI message is: It includes a rectangle area update flag (dr_pos_updates_flag) indicating whether the SEI message includes updates for one or more display rectangles, and Based on the first value (1) of the above rectangular area update flag (dr_pos_updates_flag), A flag (dr_rect_pos_update_flag[ i ]) indicating whether the display rectangle position information of the i-th rectangle is updated is included in the display rectangles SEI message, method.

13. Memory; and At least one processor connected to the memory; comprising, wherein the at least one processor: Generate at least one SEI message; Encoding the picture; and Configured to generate a bitstream including at least one SEI message and the picture. device.

14. A computer-readable storage medium for storing a bitstream generated by the method according to paragraph 10.

15. Step of acquiring the bitstream, The bitstream is generated based on the steps of: generating at least one SEI message; encoding a picture; and generating a bitstream comprising the at least one SEI message and the picture; and A method comprising the step of transmitting data including the bitstream above.