Method and computer-readable storage medium

The video encoding/decoding method addresses high-resolution video transmission costs by optimizing neural-network post-filters, enhancing coding efficiency and reducing unnecessary constraints.

WO2026155524A1PCT designated stage Publication Date: 2026-07-23LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2026-01-13
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

The increasing demand for high-resolution, high-quality video has led to higher transmission and storage costs due to the increase in transmitted information or bits, necessitating high-efficiency video compression technology.

Method used

A video encoding/decoding method that alleviates unnecessary constraints on neural-network post-filter output by determining and encoding neural network post-processing filters based on SEI messages, including NNPFC and NNPFA, to improve coding efficiency.

Benefits of technology

The method enhances coding efficiency by relaxing constraints on neural-network post-filter outputs, providing improved video encoding/decoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2026000753_23072026_PF_FP_ABST
    Figure KR2026000753_23072026_PF_FP_ABST
Patent Text Reader

Abstract

A method according to one aspect comprises the steps of: acquiring, from a bitstream, a neural-network post-filter characteristics (NNPFC) SEI message and a neural-network post-filter activation (NNPFA) SEI message; determining, on the basis of the NNPFC SEI message, at least one neural network that can be used as a neural network post-processing filter; and determining, on the basis of the NNPFA SEI message, whether to activate the neural network post-processing filter, wherein: the NNPFA SEI message includes NNPFA output information indicating whether a picture, which corresponds to an input picture and is generated on the basis of the neural network post-processing filter, is output; and the NNPFA output information can indicate, on the basis of a condition including that the purpose of the neural network post-processing filter does not correspond to temporal extrapolation, that the picture generated on the basis of the neural network post-processing filter is output.
Need to check novelty before this filing date? Find Prior Art

Description

Method and computer-readable storage medium

[0001] The present disclosure relates to a method for encoding and decoding image information, a method for transmitting a bitstream, and a computer-readable storage medium for storing the same non-transiently.

[0002] Recently, the demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition), has been increasing across various fields. As video data becomes higher in resolution and quality, the relative amount of information or bits transmitted increases compared to conventional video data. This increase in transmitted information or bits leads to higher transmission and storage costs.

[0003] Accordingly, high-efficiency video compression technology is required to effectively transmit, store, and play back high-resolution, high-quality video information.

[0004] The present disclosure aims to provide a video encoding / decoding method and apparatus with improved coding efficiency.

[0005] The present disclosure aims to provide a video encoding / decoding method and apparatus with improved coding efficiency by alleviating unnecessary constraints on NNPF output.

[0006] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure belongs from the description below.

[0007] A method according to one aspect comprises: acquiring an NNPFC (neural-network post-filter characteristics) SEI message and an NNPFA (neural-network post-filter activation) SEI message from a bitstream; determining at least one neural network that can be used as a neural network post-processing filter based on the NNPFC SEI message; and determining whether the neural network post-processing filter is activated based on the NNPFA SEI message, wherein the NNPFA SEI message includes NNPFA output information indicating whether a picture generated based on the neural network post-processing filter, corresponding to an input picture, is output, and the NNPFA output information may indicate that a picture generated based on the neural network post-processing filter is output based on a condition including that the purpose of the neural network post-processing filter does not correspond to temporal extrapolation.

[0008] The satisfaction of the above condition can be determined based on the value of a temporal extrapolation flag indicating whether the purpose of the neural network post-processing filter corresponds to the temporal extrapolation.

[0009] The above NNPFC SEI message includes NNPFC objective information indicating the objective of the neural network post-processing filter, and the value of the temporal extrapolation flag can be derived based on the NNPFC objective information.

[0010] The above condition can be satisfied based on the fact that the purpose of the neural network post-processing filter does not correspond to picture rate upsampling and temporal extrapolation, and that the number of NNPFA output information present in the NNPFA SEI message is equal to the number of pictures having corresponding input pictures among the pictures included in the NNPF output tensor.

[0011] A method according to one aspect comprises: determining at least one neural network that can be used as a neural network post-processing filter; determining whether the neural network post-processing filter is activated; and encoding image information including an NNPFC (neural-network post-filter characteristics) SEI message containing information regarding at least one neural network that can be used as a neural network post-processing filter and an NNPFA (neural-network post-filter activation) SEI message indicating whether the neural network post-processing filter is activated.

[0012] The above NNPFA SEI message includes NNPFA output information indicating whether a picture generated based on the neural network post-processing filter corresponding to the input picture is output, and the NNPFA output information may indicate that a picture generated based on the neural network post-processing filter is output based on a condition including that the purpose of the neural network post-processing filter does not correspond to temporal extrapolation.

[0013] The satisfaction of the above condition can be determined based on the value of a temporal extrapolation flag indicating whether the purpose of the neural network post-processing filter corresponds to the temporal extrapolation.

[0014] The above NNPFC SEI message includes NNPFC objective information indicating the objective of the neural network post-processing filter, and the value of the temporal extrapolation flag can be derived based on the NNPFC objective information.

[0015] The above condition can be satisfied based on the fact that the purpose of the neural network post-processing filter does not correspond to picture rate upsampling and temporal extrapolation, and that the number of NNPFA output information present in the NNPFA SEI message is equal to the number of pictures having corresponding input pictures among the pictures included in the NNPF output tensor.

[0016] A computer-readable storage medium according to one aspect can store a bitstream generated by the method non-transiently.

[0017] A method according to one aspect comprises: a step of generating a bitstream; and a step of transmitting data including said bitstream; wherein the step of generating the bitstream comprises: a step of determining at least one neural network that can be used as a neural network post-processing filter; a step of determining whether said neural network post-processing filter is activated; and a step of encoding image information including an NNPFC (neural-network post-filter characteristics) SEI message containing information regarding at least one neural network that can be used as a neural network post-processing filter and an NNPFA (neural-network post-filter activation) SEI message indicating whether said neural network post-processing filter is activated, said NNPFA SEI message includes NNPFA output information indicating whether a picture generated based on said neural network post-processing filter, corresponding to an input picture, is output, said NNPFA output information may indicate that a picture generated based on said neural network post-processing filter is output based on a condition including that the purpose of said neural network post-processing filter does not correspond to temporal extrapolation.

[0018] The features briefly summarized above regarding the present disclosure are merely exemplary aspects of the detailed description of the present disclosure that follows and do not limit the scope of the present disclosure.

[0019] According to the present disclosure, a video encoding / decoding method and apparatus with improved coding efficiency may be provided.

[0020] In addition, according to the present disclosure, coding efficiency can be improved by relaxing unnecessary constraints on the NNPF output.

[0021] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure pertains from the description below.

[0022] FIG. 1 is a schematic diagram illustrating a video coding system to which an embodiment according to the present disclosure can be applied.

[0023] FIG. 2 is a schematic diagram illustrating an encoding device to which an embodiment according to the present disclosure can be applied.

[0024] FIG. 3 is a schematic diagram illustrating a decoding device to which an embodiment according to the present disclosure can be applied.

[0025] FIG. 4 shows an example of a video / image decoding method to which an embodiment of the present disclosure can be applied.

[0026] FIG. 5 shows an example of a video / image encoding method to which an embodiment of the present disclosure can be applied.

[0027] Figure 6 illustrates an exemplary hierarchical structure for a coded video / image.

[0028] Figure 7 is a diagram illustrating an interleaved method for inducing a luma channel.

[0029] FIG. 8 is a diagram illustrating a method for encoding image information according to one embodiment.

[0030] FIG. 9 is a diagram illustrating a method for decoding image information according to one embodiment.

[0031] FIG. 10 is a drawing illustrating an exemplary content streaming system to which an embodiment according to the present disclosure can be applied.

[0032] Hereinafter, embodiments of the present disclosure are described in detail with reference to the attached drawings so that those skilled in the art can easily implement them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein.

[0033] In describing the embodiments of the present disclosure, detailed descriptions of known configurations or functions are omitted if it is determined that such descriptions could obscure the essence of the present disclosure. Additionally, parts of the drawings unrelated to the description of the present disclosure have been omitted, and similar parts are denoted by similar reference numerals.

[0034] In the present disclosure, when a component is described as being "connected," "combined," or "joined" with another component, this may include not only a direct connection but also an indirect connection in which another component exists in between. Furthermore, when a component is described as "comprising" or "having" another component, this means that, unless specifically stated otherwise, it does not exclude the other component but may include an additional component.

[0035] In the present disclosure, terms such as first, second, etc. are used solely for the purpose of distinguishing one component from another and do not limit the order or importance of the components unless specifically stated otherwise. Accordingly, within the scope of the present disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and likewise, a second component in one embodiment may be referred to as a first component in another embodiment.

[0036] In this disclosure, distinct components are intended to clearly describe their respective features and do not imply that the components are separate. That is, multiple components may be integrated to form a single hardware or software unit, or a single component may be distributed to form multiple hardware or software units. Accordingly, such integrated or distributed embodiments are included within the scope of this disclosure, unless otherwise noted.

[0037] In the present disclosure, the components described in various embodiments do not necessarily mean essential components, and some may be optional components. Accordingly, embodiments consisting of a subset of the components described in one embodiment are also included within the scope of the present disclosure. Furthermore, embodiments including additional components in addition to the components described in various embodiments are also included within the scope of the present disclosure.

[0038] The present disclosure relates to the encoding and decoding of images. For example, the methods and embodiments disclosed in this document may be applied to methods disclosed in the VVC (versatile video coding) standard, EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard) or next-generation video / image coding standards (e.g., H.267, H.268 or H.274, etc.).

[0039] The present disclosure presents various embodiments relating to video / image coding, and unless otherwise stated, said embodiments may be performed in combination with one another.

[0040] Unless newly defined in this disclosure, the terms used herein may have the ordinary meanings commonly used in the technical field to which this disclosure belongs.

[0041] In this disclosure, "picture" generally refers to a unit representing a single image of a specific time period, and a slice / tile is a coding unit constituting a part of the picture, and a picture may be composed of one or more slices / tiles. Additionally, a slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more tile groups. A tile group may include one or more tiles. A brick may represent a rectangular area of ​​the rows of CTUs of a tile within a picture. In this document, tile group and slice may be used interchangeably. For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.

[0042] In the present disclosure, "pixel" or "pel" may refer to the smallest unit constituting a picture (or image). Additionally, "sample" may be used as a term corresponding to pixel. A sample may generally represent a pixel or a pixel value, may represent only the pixel / pixel value of the luminance component, or may represent only the pixel / pixel value of the chroma component.

[0043] In this disclosure, "unit" may represent a basic unit of image processing. A unit may include at least one of a specific area of ​​a picture and information related to that area. A unit may include one luminance block and two chroma (e.g., cb, cr) blocks. Depending on the case, the term "unit" may be used interchangeably with terms such as "sample array," "block," or "area." In general, an MxN block may include samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows.

[0044] In the present disclosure, "current block" may mean one of "current coding block," "current coding unit," "block to be encoded," "block to be decoded," or "block to be processed." When prediction is performed, "current block" may mean "current prediction block" or "block to be predicted." When transformation (inverse transformation) / quantization (inverse quantization) is performed, "current block" may mean "current transformation block" or "block to be transformed." When filtering is performed, "current block" may mean "block to be filtered."

[0045] In the present disclosure, "current block" may mean a block comprising both a luminous component block and a chroma component block, or "luma block of the current block," unless explicitly stated as a chroma block. The luminous component block of the current block may be expressed by including an explicit description of a luminous component block, such as "luma block" or "current luminous block." Additionally, the chroma component block of the current block may be expressed by including an explicit description of a chroma component block, such as "chroma block" or "current chroma block."

[0046] In the present disclosure, " / " and "," may be interpreted as "and / or." For example, "A / B" and "A, B" may be interpreted as "A and / or B." Additionally, "A / B / C" and "A, B, C" may mean "at least one of A, B and / or C."

[0047] In the present disclosure, "or" may be interpreted as "and / or". For example, "A or B" may mean 1) "A" only, 2) "B" only, or 3) "A and B". Alternatively, in the present disclosure, "or" may mean "additionally or alternatively".

[0048] FIG. 1 is a schematic diagram illustrating a video / image coding system to which an embodiment according to the present disclosure can be applied.

[0049] Referring to FIG. 1, a video / image coding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image or data in the form of a file or streaming to the receiving device via a digital storage medium or a network.

[0050] The source device may include a video source, an encoding device, and a transmission unit. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.

[0051] A video source may acquire video / images through processes such as video / image capture, synthesis, or generation. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include, for example, a computer, a tablet, and a smartphone, etc., and may generate video / images (electronically). For example, virtual video / images may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process in which related data is generated.

[0052] The encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0053] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device in the form of a file or streaming via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit may include elements for creating a media file through a predetermined file format and elements for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.

[0054] The decoding device can decode video / images by performing a series of procedures such as inverse quantization, inverse transform, and prediction corresponding to the operation of the encoding device.

[0055] The renderer can render the decoded video / image. The rendered video / image can be displayed through the display unit.

[0056] FIG. 2 is a schematic diagram illustrating an encoding device to which an embodiment according to the present disclosure can be applied.

[0057] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a predictor (220), a residual processor (230), an entropy encoder (240), an adder (250), a filter (260), and a memory (270). The predictor (220) may include an inter-predictor (221) and an intra-predictor (222). The residual processor (230) may include a transformer (232), a quantizer (233), a dequantizer (234), and an inverse transformer (235). The residual processor (230) may further include a subtractor (231). The addition unit (250) may be referred to as a reconstructor or a reconstructed block generator. The aforementioned image segmentation unit (210), prediction unit (220), residual processing unit (230), entropy encoding unit (240), addition unit (250), and filtering unit (260) may be configured by one or more hardware components (e.g., an encoder chipset or processor) according to the embodiment. Additionally, the memory (270) may include a Decoded Picture Buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.

[0058] The image segmentation unit (210) can divide an input image (or picture, frame) input to an encoding device (200) into one or more processing units. For example, the processing unit may be called a coding unit (CU). A coding unit may be recursively divided into a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, a single coding unit may be divided into multiple coding units of a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. For example, a quad-tree structure may be applied first, and a binary-tree structure and / or a ternary-tree structure may be applied later. Alternatively, a binary-tree structure may be applied first. A coding procedure according to the present disclosure may be performed based on the final coding unit that is no longer divided. In this case, based on coding efficiency according to image characteristics, the maximum coding unit may be used directly as the final coding unit, or, if necessary, the maximum coding unit may be recursively divided into lower-depth coding units so that a coding unit of the optimal size is used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). The prediction unit and the transformation unit may each be divided or partitioned from the final coding unit.The above prediction unit may be a unit of sample prediction, and the above transformation unit may be a unit that derives transformation coefficients and / or a unit that derives a residual signal from transformation coefficients.

[0059] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the context. In general, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chroma component. A sample may be used to refer to a single picture (or image) as a term corresponding to a pixel or pel.

[0060] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (predicted block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input image signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, as illustrated, the unit that subtracts the prediction signal (predicted block, prediction sample array) from the input image signal (original block, original sample array) within the encoding device (200) may be called a subtraction unit (231). The prediction unit (220) can perform a prediction for a block to be processed (hereinafter, current block) and generate a predicted block (predicted block) containing prediction samples for said current block. The prediction unit (220) can determine whether intra prediction is applied or inter prediction is applied in units of the current block or CU. The prediction unit (220) can generate various information regarding prediction, such as prediction mode information, as described below in the description of each prediction mode, and transmit it to the entropy encoding unit (240). The information regarding prediction can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.

[0061] The intra prediction unit (222) can predict the current block by referring to samples within the current picture. The referenced samples may be located near the current block or away from it, depending on the prediction mode. In intra prediction, the prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a Planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is merely an example, and depending on the settings, more or fewer directional prediction modes may be used. The intra prediction unit (222) may also determine the prediction mode applied to the current block by using the prediction mode applied to the surrounding blocks.

[0062] The inter prediction unit (221) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture containing the reference blocks and the reference picture containing the temporal neighboring blocks may be the same or different from each other. The temporal neighboring blocks may be referred to by names such as collocated reference block, collocated CU (colCU), etc. A reference picture containing the aforementioned temporal surrounding blocks may be called a collocated picture (colPic). For example, the inter prediction unit (221) may construct a list of motion information candidates based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of surrounding blocks as motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0063] The prediction unit (220) may generate a prediction signal based on various prediction methods and / or prediction techniques described below. For example, the prediction unit (220) may apply intra prediction or inter prediction for the prediction of the current block, as well as apply intra prediction and inter prediction simultaneously. A prediction method that applies intra prediction and inter prediction simultaneously for the prediction of the current block may be called combined inter and intra prediction (CIIP). Additionally, the prediction unit (220) may be based on an intra block copy (IBC) prediction mode or a palette mode for the prediction of the block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, for example, screen content coding (SCC). IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. Palette mode can be viewed as an example of intra-coding or intra-prediction. When palette mode is applied, sample values ​​within a picture can be signaled based on information regarding palette tables and palette indices.

[0064] The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or to generate a residual signal. The subtraction unit (231) can generate a residual signal (residual signal, residual block, residual sample array) by subtracting the prediction signal (predicted block, prediction sample array) output from the prediction unit (220) from the input image signal (original block, original sample array). The generated residual signal can be transmitted to the conversion unit (232).

[0065] The transformation unit (232) can generate transform coefficients by applying a transformation technique to a residual signal. For example, the transformation technique may include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transformation obtained based on a prediction signal generated using all previously reconstructed pixels. The transformation process may be applied to a block of pixels of the same size in a square, or to a block of variable size that is not square.

[0066] The quantization unit (233) can quantize the transformation coefficients and transmit them to the entropy encoding unit (240). The entropy encoding unit (240) can encode the quantized signal (information regarding the quantized transformation coefficients) and output it as a bitstream. The information regarding the quantized transformation coefficients may be called residual information. The quantization unit (233) can rearrange the block-shaped quantized transformation coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information regarding the quantized transformation coefficients based on the one-dimensional vector-shaped quantized transformation coefficients.

[0067] The entropy encoding unit (240) can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit (240) may encode information required for video / image restoration (e.g., values ​​of syntax elements) together or separately, in addition to quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may further include general constraint information. The signaling information, transmitted information, and / or syntax elements mentioned in the present disclosure may be included in the video / image information. The video / image information may be encoded through the encoding procedure described above and included in the bitstream.

[0068] The above bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmitting a signal output from the entropy encoding unit (240) and / or a storage unit (not shown) for storing it may be provided as an internal / external element of the encoding device (200), or the transmission unit may be provided as a component of the entropy encoding unit (240).

[0069] The quantized transformation coefficients output from the quantization unit (233) can be used to generate a residual signal. For example, a residual signal (residual block or residual samples) can be restored by applying inverse quantization and inverse transformation to the quantized transformation coefficients through the inverse quantization unit (234) and the inverse transformation unit (235).

[0070] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.

[0071] The adder (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit (221) or the intra prediction unit (222). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The adder (250) may be called a reconstructed unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after undergoing filtering as described below.

[0072] The filtering unit (260) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (260) can generate a modified restored picture by applying various filtering methods to the restored picture, and can store the modified restored picture in memory (270), specifically in the DPB of memory (170). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (240), as described below in the description of each filtering method. The information regarding filtering can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.

[0073] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter-prediction unit (221). Through this, the encoding device (200) can avoid prediction mismatches between the encoding device (200) and the decoding device when inter-prediction is applied, and can also improve encoding efficiency.

[0074] The DPB in memory (270) can store a modified restored picture to be used as a reference picture in the inter prediction unit (221). Memory (270) can store motion information of blocks from which motion information is derived (or encoded) in the current picture and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter prediction unit (221) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory (270) can store restoration samples of restored blocks in the current picture and transmit them to the intra prediction unit (222).

[0075] FIG. 3 is a schematic diagram illustrating a decoding device to which an embodiment according to the present disclosure can be applied.

[0076] As illustrated in FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-predictor (332) and an intra-predictor (331). The residual processor (320) may include a dequantizer (321) and an inverse transformer (322). The aforementioned entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) may be configured by a single hardware component (e.g., a decoder chipset or a processor) according to an embodiment. Additionally, the memory (360) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.

[0077] When a bitstream containing video / image information is input, the decoding device (300) can restore the image by performing a process corresponding to the process performed by the encoding device (200) of FIG. 2. For example, the decoding device (300) can perform decoding using a processing unit applied in the encoding device (200). Thus, the processing unit for decoding may be, for example, a coding unit. The coding unit may be a coding tree unit, or a maximum coding unit may be obtained by dividing it according to a quad tree structure, a binary tree structure, and / or a binary tree structure. And, the restored image signal decoded and output through the decoding device (300) can be played back through a playback device (not shown).

[0078] The decoding device (300) can receive a signal output from the encoding device (200) of FIG. 2 in the form of a bitstream. The received signal can be decoded through an entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information necessary for image restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may further include general constraint information. The decoding device (300) can decode the picture based on the information regarding the parameter sets and / or the general constraint information. The signaling / received information and / or syntax elements described below can be obtained from the bitstream by decoding through the decoding procedure. For example, the entropy decoding unit (310) can decode information within the bitstream based on coding methods such as exponential chord coding, CAVLC, or CABAC, and output values ​​of syntax elements required for image restoration and quantized values ​​of transformation coefficients regarding residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using information on the syntax element to be decoded and decoding information of surrounding and decoding target blocks or information on symbols / bins decoded in the previous step, predict the probability of occurrence of the bin according to the determined context model, and perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin after determining the context model. Among the information decoded in the entropy decoding unit (310), information regarding prediction is provided to the prediction unit (330), and residual values ​​for which entropy decoding was performed in the entropy decoding unit (310), i.e., quantized transformation coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive residual signals (residual blocks, residual samples, residual sample array). Additionally, among the information decoded in the entropy decoding unit (310), information regarding filtering can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (300), or the receiving unit may be a component of the entropy decoding unit (310). Meanwhile, the decoding device according to the present document may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit (310), and the sample decoder may include at least one of the inverse quantization unit (321), inverse transform unit (322), adder (340), filtering unit (350), memory (360), inter prediction unit (332), and intra prediction unit (331).

[0079] In the inverse quantization unit (321), the quantized transformation coefficients can be inversely quantized to output transformation coefficients. The inverse quantization unit (321) can rearrange the quantized transformation coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed in the encoding device (200). The inverse quantization unit (321) can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain transformation coefficients.

[0080] In the inverse conversion unit (322), the conversion coefficients can be inversely converted to obtain a residual signal (residual block, residual sample array).

[0081] The prediction unit (330) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit may apply intra prediction or inter prediction for a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called combined inter and intra prediction (CIIP). Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, for example, screen content coding (SCC). IBC basically performs prediction within the current picture, but it can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. The palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and palette index can be included in the above video / image information and signaled.

[0082] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The description of the intra prediction unit (222) of the encoding device (200) may be applied equally to the intra prediction unit (331). The referenced samples may be located next to the current block or away from it depending on the prediction mode. In intra prediction, the prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit (331) may determine the prediction mode applied to the current block by using the prediction mode applied to the surrounding blocks.

[0083] The inter prediction unit (332) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) may construct a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-prediction can be performed based on various prediction modes (techniques), and information regarding the prediction may include information indicating the mode (technique) of inter-prediction for the current block.

[0084] The adder (340) can generate a restoration signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the prediction signal (predicted block, predicted sample array) output from the prediction unit (330) (including the inter prediction unit (332) and / or intra prediction unit (331)). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the restoration block. The description of the adder (250) can be applied equally to the adder (340). The adder (340) may be called a restoration unit or a restoration block generation unit. The generated restoration signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after undergoing filtering as described below.

[0085] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.

[0086] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can generate a modified restored picture by applying various filtering methods to the restored picture, and can store the modified restored picture in memory (360), specifically in the DPB of memory (360). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0087] The (modified) restored picture stored in the DPB of the memory (360) can be used as a reference picture in the inter-prediction unit (332). The memory (360) can store motion information of blocks from which motion information within the current picture has been derived (or decoded) and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (332) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (360) can store restoration samples of blocks restored within the current picture and transmit them to the intra-prediction unit (331).

[0088] In this specification, the embodiments described in the filtering unit (260), inter prediction unit (221), and intra prediction unit (222) of the encoding device (200) may be applied to the filtering unit (350), inter prediction unit (332), and intra prediction unit (331) of the decoding device (300) in the same or corresponding manner.

[0089] A video / image coding method according to the present disclosure may be performed based on the following partitioning structure. Specifically, the procedures described below, such as prediction, residual processing ((inverse)transform, (inverse)quantization, etc.), syntax element coding, and filtering, may be performed based on CTU and CU (and / or TU, PU) derived based on the partitioning structure. The block partitioning procedure may be performed in the image segmentation unit (210) of the encoding device described above, and the partitioning-related information may be processed (encoded) in the entropy encoding unit (240) and transmitted to the decoding device in the form of a bitstream. The entropy decoding unit (310) of the decoding device may derive the block partitioning structure of the current picture based on the partitioning-related information obtained from the bitstream, and perform a series of procedures for image decoding (e.g., prediction, residual processing, block / picture restoration, in-loop filtering, etc.) based thereon. The CU size and the TU size may be the same, or multiple TUs may exist within the CU area. Meanwhile, the term CU size generally refers to the CB size of the luminous component (sample). The term TU size generally refers to the TB size of the luminous component (sample).

[0090] The chroma component (sample) CB or TB size can be derived based on the luminance component (sample) CB or TB size according to the component ratio based on the color format (chroma format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.) of the picture / image. The TU size can be derived based on maxTbSize. For example, if the CU size is larger than the maxTbSize, multiple TUs (TBs) of the maxTbSize are derived from the CU, and conversion / inverse conversion can be performed in units of the TU (TB). Additionally, for example, when intra prediction is applied, the intra prediction mode / type is derived in units of the CU (or CB), and the procedure for deriving surrounding reference samples and generating prediction samples can be performed in units of the TU (or TB). In this case, one or more TUs (or TBs) may exist within a single CU (or CB) region, and in this case, the multiple TUs (or TBs) may share the same intra prediction mode / type.

[0091] Additionally, in the coding of video / image according to the present disclosure, the image processing unit may have a hierarchical structure. A picture may be divided into one or more tiles, bricks, slices, and / or tile groups. A slice may include one or more bricks. A brick may include one or more CTU rows within the tile. A slice may include an integer number of bricks in the picture. A tile group may include one or more tiles. A tile may include one or more CTUs. The CTU may be divided into one or more CUs. A tile is a rectangular area within a picture that includes CTUs within a specific tile row and a specific tile column. A tile group may include an integer number of tiles according to a tile raster scan within the picture. A slice header may carry information / parameters that can be applied to the corresponding slice (blocks within the slice). If the encoding / decoding device has a multi-core processor, the encoding / decoding procedure for the tile, slice, brick, and / or tile group may be processed in parallel.

[0092] In the present disclosure, slices or tile groups may be used interchangeably. That is, a tile group header may be referred to as a slice header. Here, a slice may have one of the slice types including an intra (I) slice, a predictive (P) slice, and a bi-predictive (B) slice. For blocks within an I slice, only intra prediction may be used for prediction, and no inter prediction may be used. Of course, even in this case, the original sample value may be coded and signaled without prediction. For blocks within a P slice, intra prediction or inter prediction may be used, and if inter prediction is used, only uni prediction may be used. Meanwhile, for blocks within a B slice, intra prediction or inter prediction may be used, and if inter prediction is used, up to bi-prediction may be used.

[0093] In an encoding device, tile / tile group, brick, slice, and maximum and minimum coding unit sizes are determined based on video characteristics (e.g., resolution) or by considering coding efficiency or parallel processing, and information regarding this or information that can derive it may be included in the bitstream.

[0094] The decoding device can obtain information indicating whether the tile / tile group, brick, slice, or CTU within the tile of the current picture has been divided into multiple coding units. Efficiency can be increased by obtaining (transmitting) this information only under specific conditions.

[0095] The slice header (slice header syntax) may include information / parameters that can be commonly applied to the slice. The APS (APS syntax) or PPS (PPS syntax) may include information / parameters that can be commonly applied to one or more pictures. The SPS (SPS syntax) may include information / parameters that can be commonly applied to one or more sequences. The VPS (VPS syntax) may include information / parameters that can be commonly applied to multiple layers. The DPS (DPS syntax) may include information / parameters that can be commonly applied to the entire video. The DPS may include information / parameters related to the concatenation of the CVS (coded video sequence).

[0096] In the present disclosure, the term "higher-level syntax" may include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, and slice header syntax.

[0097] In addition, for example, information regarding the division and configuration of the above tile / tile group / brick / slice can be configured at the encoding stage through the above-mentioned upper-level syntax and transmitted to the decoding device in the form of a bitstream.

[0098] Pictures can be divided into sequences of Coding Tree Units (CTUs). A CTU may correspond to a Coding Tree Block (CTB). Alternatively, a CTU may include a Coding Tree Block of Luma Samples and two Coding Tree Blocks of corresponding Chroma Samples. In other words, for a picture containing three sample arrays, a CTU may include an NxN block of Luma Samples and two corresponding blocks of Chroma Samples.

[0099] The maximum allowable size of a CTU for coding and prediction, etc., may differ from the maximum allowable size of a CTU for transformation. For example, the maximum allowable size of a luminance block within a CTU may be 128x128 (even though the maximum size of luminance ring blocks is 64x64).

[0100] A picture is divided into one or more tile rows and one or more tile columns. A tile is a sequence of CTUs covering a rectangular area of ​​the picture. The CTUs within a tile are scanned in raster scan order within that tile.

[0101] A slice consists of an integer number of complete tiles within a picture or an integer number of consecutive complete CTU rows within a single tile. Two modes are supported for slicing: raster-scan slice mode and rectangular slice mode.

[0102] In raster-scan slice mode, a slice comprises a sequence of complete tiles according to the tile raster scan order of the picture. In rectangular slice mode, a slice comprises a number of complete tiles collectively configured to form a rectangular area of ​​the picture, or a number of consecutive complete CTU rows that collectively form a rectangular area within a single tile. The tiles within a rectangular slice are scanned in the tile raster scan order within the rectangular area corresponding to that slice.

[0103] A subpicture includes one or more slices that collectively cover a rectangular area of ​​a picture. According to one example, it is also possible for a picture to be divided into 28 subpictures of different sizes.

[0104] When a picture is encoded into three separate color planes (where separate_colour_plane_flag is 1), the slice contains only CTUs of a single color component identified by the corresponding value of colour_plane_id, and each array of color components of the picture consists of slices having the same colour_plane_id value.

[0105] Encoded slice NAL units having different colour_plane_id values ​​within a picture can be interleaved with respect to each colour_plane_id value, provided that for each colour_plane_id value, the encoded slice NAL units having that colour_plane_id value are arranged in an order of increasing CTU addresses in the tile scan order for the first CTU of each slice NAL unit.

[0106] Meanwhile, when separate_colour_plane_flag is 0, each CTU of the picture is contained in exactly one slice. When separate_colour_plane_flag is 1, the CTU of each color component is contained in exactly one slice (i.e., information for each CTU of the picture is contained in exactly three slices, and these three slices have different colour_plane_id values).

[0107] Tiles change the order of CTUs within a picture. If a picture is divided into two or more tiles, the order of CTUs becomes the raster-scan order within each tile, which can be exemplified by a case where the picture is divided into two tiles and each tile has 8 CTUs. Note that the CTUs are arranged in raster-scan order within each tile.

[0108] FIG. 4 shows an example of a video / image decoding method to which an embodiment of the present disclosure can be applied.

[0109] In video coding, the pictures constituting the video can be decoded according to a series of decoding orders. The picture order corresponding to the output order of the decoded pictures can be set differently from the decoding order, and based on this, not only forward prediction but also reverse prediction can be performed during inter-prediction.

[0110] In FIG. 4, S400 may be performed in the entropy decoding unit (310) of the aforementioned decoding device (300), S410 may be performed in the prediction unit (330), S420 may be performed in the residual processing unit (320), S430 may be performed in the addition unit (340), and S440 may be performed in the filtering unit (350). S400 may include a decoding procedure according to the present disclosure, S410 may include an inter / intra prediction procedure according to the present disclosure, S420 may include a residual processing procedure according to the present disclosure, S430 may include a block / picture restoration procedure according to the present disclosure, and S440 may include an in-loop filtering procedure according to the present disclosure.

[0111] Referring to FIG. 4, the decoding device acquires image / video information from a bitstream (S400), performs a prediction based on the acquired image / video information (S410), and can restore a picture through residual processing (S420, inverse quantization and inverse transformation of the quantized transformation coefficients) (S430).

[0112] A modified restored picture can be generated by applying an in-loop filtering procedure (S440) to the restored picture generated through the above restoration procedure, and the modified restored picture can be output as a decoded picture and also stored in the buffer or memory of the decoding device to be used as a reference picture in the inter-prediction procedure when decoding the next picture. In some cases, the above in-loop filtering procedure may be omitted, in which case the restored picture can be output as a decoded picture and also stored in the buffer or memory of the decoding device to be used as a reference picture in the inter-prediction procedure when decoding a subsequent picture.

[0113] The in-loop filtering procedure (S440) may include a deblocking filtering procedure, a sample adaptive offset (SAO) procedure, an adaptive loop filter (ALF) procedure, and / or a bilateral filter procedure, and some or all of these may be omitted. Additionally, one or some of the deblocking filtering procedure, the sample adaptive offset (SAO) procedure, the adaptive loop filter (ALF) procedure, and the bilateral filter procedure may be applied sequentially, or all of them may be applied sequentially. For example, the SAO procedure may be performed after the deblocking filtering procedure is applied to the restored picture. Alternatively, for example, the ALF procedure may be performed after the deblocking filtering procedure is applied to the restored picture. This may be performed in the same manner in the encoding device.

[0114] FIG. 5 shows an example of a video / image encoding method to which an embodiment of the present disclosure can be applied.

[0115] In FIG. 5, the prediction step (S500) may be performed in the prediction unit (220) of the aforementioned encoding device (200), residual processing (S510) based on the prediction result may be performed in the residual processing unit (230), and the step (S520) of encoding image information including prediction information and residual information may be performed in the entropy encoding unit (240). S500 may include an inter / intra prediction procedure according to the present disclosure, S510 may include a residual processing procedure according to the present disclosure, and S520 may include an encoding procedure according to the present disclosure.

[0116] The encoding procedure may optionally include not only a procedure for encoding information for picture restoration (e.g., prediction information, residual information, partitioning information, etc.) and outputting it in the form of a bitstream, but also a procedure for generating a restored picture for the current picture and a procedure for applying in-loop filtering to the restored picture.

[0117] The encoding device (200) can derive (modified) residual samples from quantized transform coefficients through the inverse quantization unit (234) and the inverse transform unit (235), and can generate a restored picture based on the (modified) residual samples and the predicted samples which are the outputs of S500. The restored picture thus generated may be identical to the restored picture generated by the decoding device (300) described above. A modified restored picture may be generated through an in-loop filtering procedure on the restored picture, which may be stored in a buffer or memory, and, as in the case of the decoding device, may be used as a reference picture in the inter-prediction procedure during the subsequent encoding of the picture.

[0118] As described above, depending on the case, part or all of the in-loop filtering procedure may be omitted. When the in-loop filtering procedure is performed, (in-loop) filtering-related information (parameters) may be encoded in the entropy encoding unit (240) and output in the form of a bitstream, and the decoding device (300) may perform the in-loop filtering procedure in the same way as the encoding device based on the filtering-related information.

[0119] Through this in-loop filtering procedure, noise generated during video / image coding, such as blocking artifacts and ringing artifacts, can be reduced, and subjective / objective image quality can be improved. In addition, by performing the in-loop filtering procedure in both the encoding device (200) and the decoding device (300), the same prediction results can be derived in both the encoding device (200) and the decoding device (300), the reliability of picture coding can be increased, and the amount of data that must be transmitted for picture coding can be reduced.

[0120] As described above, the picture restoration procedure can be performed in the encoding device (200) as well as the decoding device (300). Restoration blocks can be generated based on intra prediction / inter prediction for each block unit, and a restored picture containing the restoration blocks can be generated. If the current picture / slice / tile group is an I picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based solely on intra prediction. Meanwhile, if the current picture / slice / tile group is a P or B picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based on intra prediction or inter prediction. In this case, inter prediction may be applied to some blocks within the current picture / slice / tile group, and intra prediction may be applied to the remaining blocks.

[0121] The color components of the picture may include a luminance component and a chroma component, and unless explicitly limited in the present disclosure, embodiments according to the present disclosure may be applied to the luminance component and the chroma component.

[0122] Figure 6 illustrates an exemplary hierarchical structure for a coded video / image.

[0123] Referring to Fig. 6, the coded image is divided into a Video Coding Layer (VCL) that handles the decoding processing of the image and the image itself, a subsystem that transmits and stores the encoded information, and a Network Abstraction Layer (NAL) that exists between the VCL and the subsystem and is responsible for network adaptation functions.

[0124] In VCL, VCL data containing compressed image data (slice data) can be generated, or parameter sets containing information such as Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS), or SEI (Supplemental Enhancement Information) messages that are additionally required in the decoding process of the image can be generated.

[0125] In NAL, a NAL unit can be created by adding header information (NAL unit header) to the Raw Byte Sequence Payload (RBSP) generated in VCL. In this case, the RBSP refers to slice data, parameter sets, SEI messages, etc. generated in VCL. The NAL unit header may include NAL unit type information specified according to the RBSP data included in the NAL unit.

[0126] As shown in FIG. 6, NAL units can be classified into VCL NAL units and Non-VCL NAL units depending on the RBSP generated in VCL. A VCL NAL unit may refer to a NAL unit containing information about an image (slice data), and a Non-VCL NAL unit may refer to a NAL unit containing information necessary to decode an image (parameter set or SEI message).

[0127] The aforementioned VCL NAL unit and Non-VCL NAL unit can be transmitted over a network by attaching header information according to the data specifications of the underlying system. For example, the NAL unit can be transformed into a data format of a specified specification, such as H.266 / VVC file format, RTP (Real-time Transport Protocol), TS (Transport Stream), etc., and transmitted over various networks.

[0128] As described above, the NAL unit type can be determined according to the RBSP data structure included in the NAL unit, and information about this NAL unit type can be stored in the NAL unit header and signaled.

[0129] For example, NAL units can be broadly classified into VCL NAL unit types and Non-VCL NAL unit types depending on whether they contain information about the image (slice data). VCL NAL unit types can be classified according to the properties and types of the picture included in the VCL NAL unit, while Non-VCL NAL unit types can be classified according to the types of parameter sets.

[0130] The following is an example of a NAL unit type specified according to the type of parameter set included in the Non-VCL NAL unit type.

[0131] - APS (Adaptation Parameter Set) NAL unit: Type for the NAL unit containing the APS

[0132] - DPS(Decoding Parameter Set) NAL unit: Type for the NAL unit containing the DPS

[0133] - VPS (Video Parameter Set) NAL unit: Type for the NAL unit containing the VPS

[0134] - SPS (Sequence Parameter Set) NAL unit: Type for the NAL unit containing the SPS

[0135] - PPS(Picture Parameter Set) NAL unit: Type for the NAL unit containing the PPS

[0136] The above-described NAL unit types have syntax information for the NAL unit type, and said syntax information can be stored in the NAL unit header and signaled. For example, said syntax information may be nal_unit_type, and NAL unit types may be specified by the nal_unit_type value.

[0137] A slice header (slice header syntax, slice header information) may include information / parameters that can be commonly applied to the slice. The APS (APS syntax) or PPS (PPS syntax) may include information / parameters that can be commonly applied to one or more slices or pictures. The SPS (SPS syntax) may include information / parameters that can be commonly applied to one or more sequences. The VPS (VPS syntax) may include information / parameters that can be commonly applied to multiple layers. The DPS (DPS syntax) may include information / parameters that can be commonly applied to the entire video. The DPS may include information / parameters related to the concatenation of a CVS (coded video sequence). In the present disclosure, High Level Syntax (HLS) may include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, or slice header syntax.

[0138] In the present disclosure, image / video information encoded by an encoding device and signaled in the form of a bitstream includes not only information related to picture partitioning, intra / inter prediction information, residual information, in-loop filtering information, etc., but may also include information included in the slice header, information included in the APS, information included in the PPS, information included in the SPS, information included in the VPS, and / or information included in the DPS.

[0139] A coded picture may consist of one or more slices. Parameters describing the coded picture are signaled within the picture header (PH), and parameters describing the slices are signaled within the slice header (SH). The PH is transmitted as its own NAL unit type. The SH is located at the beginning of the NAL unit containing the slice payload (i.e., slice data).

[0140] Hereinafter, SEI messages related to embodiments of the present disclosure will be described.

[0141] The SEI message related to one embodiment may include an SEI message representing a Neural-Network Post-Filter Characteristic and a Neural-Network Post-Filter Activation (NNPFA) SEI message. An example of the NNPFC (Neural-Network Post-Filter Characteristic) SEI message syntax is shown in Table 1 below.

[0142] [Table 1]

[0143]

[0144]

[0145]

[0146] The NNPFC SEI message can specify a neural network that can be used as a post-processing filter. The use of specified post-processing filters (NNPFs) for specific pictures can be indicated using the Neural-Network Post-filter Activation (NNPFA) SEI message. Here, 'post-processing filter' and 'post-filter' may have the same meaning.

[0147] To use these SEI messages, it may be necessary to define the following variables:

[0148] - The input picture width and height in luma sample units can be represented as CroppedWidth and CroppedHeight, respectively.

[0149] CroppedYPic[idx], which is a luminance sample array of input pictures with index idx in the range of 0 to numInputPics-1, and CroppedCbPic[idx] and CroppedCrPic[idx], which are chroma sample arrays, can be used as inputs to NNPF if they exist.

[0150] - BitDepth Y can represent the bit depth for the luminance sample array of input pictures.

[0151] - BitDepth C can represent the bit depth of the chroma sample arrays (if any) of the input pictures.

[0152] - ChromaFormatIdc can represent a chroma format indicator.

[0153] - When the value of nnpfc_auxiliary_inp_idc is 1, the filtering strength control value array StrengthControlVal[idx] must contain real numbers in the range of 0 to 1 for input pictures having index idx in the range of 0 to numInputPics-1.

[0154] An input picture having index 0 may correspond to a picture in which the NNPF defined by the NNPFC SEI message is activated by the NNPFA SEI message. An input picture having index i in the range of 1 to numInputPics-1 may precede an input picture having index i-1 in the output order.

[0155] The variables SubWidthC and SubHeightC can be derived from ChromaFormatIdc.

[0156] More than one NNPFC SEI message may exist for the same picture. If two or more NNPFC SEI messages with different nnpfc_id values ​​exist or are active for the same picture, these NNPFC SEI messages may have the same or different values ​​for nnpfc_purpose and may have the same or different values ​​for nnpfc_mode_idc.

[0157] nnpfc_purpose represents the purpose of the NNPF specified in Table 2 below, where

[0158] If ( nnpfc_purpose & bitMask ) ≠ 0, the NNPF indicates that it has a purpose associated with the bitMask value in Table 2. If nnpfc_purpose is greater than 0 and ( nnpfc_purpose & bitMask ) = 0, the purpose associated with the bitMask value does not apply to the NNPF. If nnpfc_purpose is 0, the NNPF is determined by the application and may be used as specified in nnpfc_application_purpose_tag_uri.

[0159] All NNPFC SEI messages with a specific nnpfc_id value within CLVS must have the same nnpfc_purpose value.

[0160] In the bitstream according to the present disclosure, the value of nnpfc_purpose must be in the range of 0 to 255. If the value of nnpfc_purpose is 256 to 65,535, it may be reserved for future use.

[0161] [Table 2]

[0162]

[0163] The variables ChromaUpsamplingFlag, ResolutionResamplingFlag, PictureRateUpsamplingFlag, BitDepthUpsamplingFlag, ColourizationFlag, and TemporalExtrapolationFlag, which respectively specify whether nnpfc_purpose represents the purpose of an NNPF including chroma upsampling, resolution resampling, picture rate upsampling, bit depth upsampling, colorization, and temporal extrapolation, are derived as follows:

[0164] ChromaUpsamplingFlag = ((nnpfc_purpose & 0x02) > 0)? 1:0

[0165] ResolutionResamplingFlag = ( ( nnpfc_purpose & 0x04 ) > 0 ) ? 1:0

[0166] PictureRateUpsamplingFlag = ((nnpfc_purpose & 0x08) > 0)? 1:0

[0167] BitDepthUpsamplingFlag = ( ( nnpfc_purpose & 0x10 ) > 0 ) ? 1:0

[0168] ColourizationFlag = ( ( nnpfc_purpose & 0x20 ) > 0 ) ? 1:0

[0169] TemporalExtrapolationFlag = ((nnpfc_purpose & 0x40) > 0)? 1:0

[0170] SpatialExtrapolationFlag = ((nnpfc_purpose & 0x80) > 0)? 1:0

[0171] If the reserved value of nnpfc_purpose is used in the future, the syntax of this SEI message may be extended to syntax elements whose existence is conditioned by whether nnpfc_purpose is that value or is identical to any of the set of values ​​containing that value.

[0172] If ChromaFormatIdc is 3, ChromaUpsamplingFlag must be equal to 0.

[0173] If ChromaUpsamplingFlag is 1, ColourizationFlag must be equal to 0.

[0174] If PictureRateUpsamplingFlag or TemporalExtrapolationFlag is 1 and the input picture with index 0 is associated with a frame packing arrangement SEI message with fp_arrangement_type 5, then all input pictures must be associated with a frame packing arrangement SEI message with fp_arrangement_type 5 and the same fp_current_frame_is_frame0_flag value.

[0175] If TemporalExtrapolationFlag is 1, the extrapolated pictures generated by NNPF are placed after all input pictures of NNPF in the output order. If TemporalExtrapolationFlag is 1 and there is a decoded output picture after the current picture where NNPF is enabled in the output order, the extrapolated pictures generated by NNPF are placed before that decoded output picture in the output order.

[0176] nnpfc_id contains an identification number that can be used to identify an NNPF. The value of nnpfc_id is 0 or greater, or 2 32It must be within the range of -2 or less. 256 or more and 511 or less and 2 31 More than 2 32 nnpfc_id values ​​in the range of -2 or less are reserved for future use and may not exist in bitstreams suitable for versions according to the present disclosure.

[0177] If it is the first NNPFC SEI message in decoding order among the NNPFC SEI messages with a specific nnpfc_id value within the current CLVS, the following applies:

[0178] - The corresponding SEI message specifies the base NNPF.

[0179] - The SEI message applies to the currently decoded picture and all decoded pictures following the current layer in output order up to the end of the current CLVS.

[0180] If nnpfc_base_flag is 1, it indicates that the SEI message specifies the base NNPF. If nnpfc_base_flag is 0, it indicates that the SEI message specifies an update to the base NNPF.

[0181] The following constraints apply to the nnpfc_base_flag value:

[0182] - If it is the first NNPFC SEI message in decoding order with a specific nnpfc_id value within the current CLVS, the nnpfc_base_flag value must be 1.

[0183] - All NNPFC SEI messages within the same CLVS that have a specific nnpfc_id value and nnpfc_base_flag is 1 must have the same SEI payload content.

[0184] If nnpfc_base_flag is 0, the following applies:

[0185] - The SEI message defines an update for the base NNPF that is earlier in decoding order and has the same nnpfc_id value. Updates are not cumulative, and each update is applied to the base NNPF specified by the first NNPFC SEI message in decoding order that has the same nnpfc_id value within the current CLVS. The NNPF defined by the SEI message is obtained by applying the update defined by the SEI message to the base NNPF that has the same nnpfc_id value.

[0186] - The SEI message applies to the current decoded picture of the current layer in output order and all subsequent decoded pictures of the current layer until the end of the current CLVS or until immediately before the decoded picture that is located after the current decoded picture in output order within the current CLVS and is associated with the subsequent NNPFC SEI message in decoding order, where nnpfc_base_flag is 0 and has the corresponding nnpfc_id value, whichever is earlier.

[0187] If nnpfc_mode_idc is equal to 0, it indicates that neural network information is contained in the NNPFC SEI message, and that neural network information is in the format of an ISO / IEC 15938-17 bitstream.

[0188] When nnpfc_mode_idc is equal to 1, neural network information is identified by a URI indicated by nnpfc_uri, and its format is identified by a tag URI, nnpfc_tag_uri.

[0189] The nnpfc_mode_idc value must be in the range of 0 to 255. Values ​​of 2 to 255 are reserved for future use and may not exist in bitstreams suitable for versions according to the present disclosure.

[0190] nnpfc_alignment_zero_bit_a must be equal to 0.

[0191] nnpfc_tag_uri contains a tag URI with syntax and semantics defined in IETF RFC 4151 and identifies the format and related information of the neural network used for updating the base NNPF specified by nnpfc_uri or the base NNPF having the same nnpfc_id value.

[0192] nnpfc_tag_uri enables the unique identification of the format of neural network data specified by nnpfc_uri without a central registry.

[0193] If nnpfc_tag_uri is "tag:iso.org,2023:15938-17", the neural network data identified by nnpfc_uri indicates that it conforms to ISO / IEC 15938-17.

[0194] nnpfc_uri contains a URI with syntax and semantics defined in IETF Internet Standard 66 and identifies a neural network used for updating a base NNPF or a base NNPF with the same nnpfc_id value.

[0195] If nnpfc_property_present_flag is 1, it indicates that syntax elements for filter attributes, including the objective, input format, output format, and complexity, exist. If nnpfc_property_present_flag is 0, it indicates that syntax elements for filter attributes do not exist.

[0196] If nnpfc_base_flag is 1, nnpfc_property_present_flag must be 1.

[0197] When nnpfc_property_present_flag is 0, the values ​​of all syntax elements that can exist only when nnpfc_property_present_flag is 1 are each inferred as the corresponding syntax element values ​​of an NNPFC SEI message containing a base NNPF that this SEI message provides updates for.

[0198] If an NNPFC SEI message nnpfcCurr with a specific nnpfc_id value within the current CLVS is not the first NNPFC SEI message in decoding order, nor is it a repetition of the first NNPFC SEI message with that nnpfc_id value (in which case the value of nnpfc_base_flag is 0), and the value of nnpfc_property_present_flag is 1, the following constraints apply:

[0199] - The values ​​of the syntax elements located after nnpfc_property_present_flag and before nnpfc_complexity_info_present_flag in the decoding order of an NNPFC SEI message must be identical to the values ​​of the corresponding syntax elements of the first NNPFC SEI message in the decoding order that has the corresponding nnpfc_id value in the current CLVS.

[0200] - nnpfc_complexity_info_present_flag must be equal to 0, or nnpfc_complexity_info_present_flag must be equal to 1 in both the first NNPFC SEI message in decoding order with the corresponding nnpfc_id value within the current CLVS (hereinafter referred to as nnpfcBase) and nnpfcCurr, and all of the following constraints apply:

[0201] -nnpfcCurr's nnpfc_parameter_type_idc must be the same as nnpfcBase's nnpfc_parameter_type_idc.

[0202] - nnpfc_log2_parameter_bit_length_minus3 of nnpfcCurr (if present) must be less than or equal to nnpfc_log2_parameter_bit_length_minus3 of nnpfcBase.

[0203] - If nnpfc_num_parameters_idc of nnpfcBase is 0, nnpfc_num_parameters_idc of nnpfcCurr must also be 0.

[0204] - Otherwise (if nnpfc_num_parameters_idc of nnpfcBase is greater than 0), nnpfc_num_parameters_idc of nnpfcCurr must be greater than 0 and less than or equal to nnpfc_num_parameters_idc of nnpfcBase.

[0205] - If nnpfc_num_kmac_operations_idc of nnpfcBase is 0, nnpfc_num_kmac_operations_idc of nnpfcCurr must also be 0.

[0206] - Otherwise (if nnpfc_num_kmac_operations_idc of nnpfcBase is greater than 0), nnpfc_num_kmac_operations_idc of nnpfcCurr must be greater than 0 and less than or equal to nnpfc_num_kmac_operations_idc of nnpfcBase.

[0207] - If nnpfc_total_kilobyte_size of nnpfcBase is 0, nnpfc_total_kilobyte_size of nnpfcCurr must also be 0.

[0208] - Otherwise (if nnpfc_total_kilobyte_size of nnpfcBase is greater than 0), nnpfc_total_kilobyte_size of nnpfcCurr must be greater than 0 and less than or equal to nnpfc_total_kilobyte_size of nnpfcBase.

[0209] The value of nnpfc_num_input_pics_minus1 plus 1 specifies the number of pictures used as input to NNPF. The value of nnpfc_num_input_pics_minus1 must be in the range of 0 to 63. If PictureRateUpsamplingFlag is 1, the value of nnpfc_num_input_pics_minus1 must be greater than 0.

[0210] The variable numInputPics, which specifies the number of pictures used as input to NNPF, is derived as follows:

[0211] numInputPics = nnpfc_num_input_pics_minus1 + 1

[0212] If nnpfc_input_pic_filtering_flag[ i ] is 1, it indicates that the NNPF generates a corresponding output picture for the i-th input picture. If nnpfc_input_pic_filtering_flag[ i ] is 0, it indicates that the NNPF does not generate a corresponding output picture for the i-th input picture. Each NNPF-generated picture is stored in the NNPF's output tensor. If nnpfc_num_input_pics_minus1 is 0, it is inferred that nnpfc_input_pic_filtering_flag

[0000] is equal to 1. If PictureRateUpsamplingFlag is 0 and nnpfc_num_input_pics_minus1 is greater than 0, nnpfc_input_pic_filtering_flag[ i ] must be equal to 1 for at least one of i in the range 0 or greater and nnpfc_num_input_pics_minus1 or less.

[0213] If nnpfc_absent_input_pic_zero_flag is 1, NNPF indicates that an input picture not present in the bitstream is expected to be represented as a sample array with a sample value of 0. If nnpfc_absent_input_pic_zero_flag is 0, NNPF indicates that an input picture inputPicA not present in the bitstream is expected to be represented as the input picture inputPicB that is closest to inputPicA in output order and is present in the bitstream.

[0214] nnpfc_out_sub_c_flag specifies the values ​​of the variables outSubWidthC and outSubHeightC when ChromaUpsamplingFlag is 1. If nnpfc_out_sub_c_flag is 1, it specifies that outSubWidthC is 1 and outSubHeightC is 1. If nnpfc_out_sub_c_flag is 0, it specifies that outSubWidthC is 2 and outSubHeightC is 1. If ChromaFormatIdc is 2 and nnpfc_out_sub_c_flag exists, the value of nnpfc_out_sub_c_flag must be 1.

[0215] nnpfc_out_colour_format_idc specifies the color format of the NNPF-generated picture when ColourizationFlag is 1, and consequently, the values ​​of the variables outSubWidthC and outSubHeightC. If nnpfc_out_colour_format_idc is 1, it specifies that the color format of the NNPF-generated picture is 4:2:0 format and that both outSubWidthC and outSubHeightC are 2. If nnpfc_out_colour_format_idc is 2, it specifies that the color format is 4:2:2 format and that outSubWidthC is 2 and outSubHeightC is 1. If nnpfc_out_colour_format_idc is 3, it specifies that the color format is 4:4:4 format and that both outSubWidthC and outSubHeightC are 1. The value of nnpfc_out_colour_format_idc must not be equal to 0.

[0216] When both ChromaUpsamplingFlag and ColourizationFlag are 0, outSubWidthC and outSubHeightC are inferred to be equal to SubWidthC and SubHeightC, respectively.

[0217] The value of nnpfc_pic_width_num_minus1 plus 1 and the value of nnpfc_pic_width_denom_minus1 plus 1 specify the numerator and denominator, respectively, of the width resampling ratio of the NNPF-generated picture for CroppedWidth. The values ​​of nnpfc_pic_width_num_minus1 and nnpfc_pic_width_denom_minus1 must both be in the range of 0 to 65,535.

[0218] The value of ( nnpfc_pic_width_num_minus1 + 1 ) ÷ ( nnpfc_pic_width_denom_minus1 + 1 ) must be in the range of 1 / 16 or greater and 16 or less. If nnpfc_pic_width_num_minus1 and nnpfc_pic_width_denom_minus1 do not exist, it is inferred that the values ​​of nnpfc_pic_width_num_minus1 and nnpfc_pic_width_denom_minus1 are both equal to 0.

[0219] The variables nnpfcOutputPicWidth and nnpfcOutputPicHeight represent the width and height of the luminance sample array of the NNPF-generated picture, respectively, when SpatialExtrapolationFlag is 0.

[0220] The variables nnpfcOutputPicWidth1 and nnpfcOutputPicHeight1 represent the width and height, respectively, of the luminance sample array of the NNPF-generated picture when SpatialExtrapolationFlag is 1.

[0221] The variable nnpfcOutputPicWidth is derived as follows:

[0222] nnpfcOutputPicWidth = Ceil( CroppedWidth * ( nnpfc_pic_width_num_minus1 + 1 ) ÷ ( nnpfc_pic_width_denom_minus1 + 1 ) )

[0223] When SpatialExtrapolation is 1, nnpfcOutputPicWidth1 is derived as follows:

[0224] nnpfcOutputPicWidth1 = nnpfcOutputPicWidth + outSubWidthC * (nnpfc_spatial_extrapolation_left_offset+nnpfc_spatial_extrapolation_right_offset )

[0225] As a requirement for bitstream conformance, when SpatialExtrapolation is 0, nnpfcOutputPicWidth must be greater than 0 and nnpfcOutputPicWidth % outSubWidthC must be equal to 0.

[0226] As a requirement for bitstream conformance, when SpatialExtrapolation is 1, nnpfcOutputPicWidth1 must be greater than 0 and nnpfcOutputPicWidth1 % outSubWidthC must be equal to 0.

[0227] The value of nnpfc_pic_height_num_minus1 plus 1 and the value of nnpfc_pic_height_denom_minus1 plus 1 specify the numerator and denominator, respectively, of the height resampling ratio of the NNPF-generated picture for CroppedHeight. The values ​​of nnpfc_pic_height_num_minus1 and nnpfc_pic_height_denom_minus1 must both be in the range of 0 to 65,535.

[0228] The value of ( nnpfc_pic_height_num_minus1 + 1 ) ÷ ( nnpfc_pic_height_denom_minus1 + 1 ) must be in the range of 1 / 16 or greater and 16 or less. If nnpfc_pic_height_num_minus1 and nnpfc_pic_height_denom_minus1 do not exist, their values ​​are all inferred to be equal to 0.

[0229] The variable nnpfcOutputPicHeight is derived as follows:

[0230] nnpfcOutputPicHeight = Ceil( CroppedHeight * ( nnpfc_pic_height_num_minus1 + 1 ) ÷ ( nnpfc_pic_height_denom_minus1 + 1 ) )

[0231] When SpatialExtrapolation is 1, nnpfcOutputPicHeight1 is derived as follows:

[0232] nnpfcOutputPicHeight1 = nnpfcOutputPicHeight + outSubHeightC *

[0233] (nnpfc_spatial_extrapolation_top_offset+nnpfc_spatial_extrapolation_bottom_offset )

[0234] As a requirement for bitstream conformance, when SpatialExtrapolation is 0, nnpfcOutputPicHeight must be greater than 0 and nnpfcOutputPicHeight % outSubHeightC must be equal to 0.

[0235] As a requirement for bitstream conformance, when SpatialExtrapolation is 1, nnpfcOutputPicHeight1 must be greater than 0 and nnpfcOutputPicHeight1 % outSubHeightC must be equal to 0.

[0236] If ResolutionResamplingFlag is 1, at least one of the following conditions must be true:

[0237] - The value of nnpfcOutputPicWidth is not equal to CroppedWidth.

[0238] - The value of nnpfcOutputPicHeight is not equal to CroppedHeight.

[0239] - SpatialExtrapolationFlag is 1.

[0240] nnpfc_interpolated_pics[ i ] specifies the number of interpolated pictures generated by NNPF between the i-th input picture and the (i+1)-th input picture of NNPF. The value of nnpfc_interpolated_pics[ i ] must be in the range of 0 to 63. If nnpfc_interpolated_pics[ i ] syntax elements exist, the value of nnpfc_interpolated_pics[ i ] must be greater than 0 for at least one of i in the range of 0 to (nnpfc_num_input_pics_minus1 - 1).

[0241] For any NNPF, PictureRateUpsamplingFlag is 1, and for the NNPFA SEI message that enabled this NNPF, nnpfa_persistence_flag is 1, and for only one value of i in the range from 0 to (numInputPics - 1), the value of nnpfc_interpolated_pics[i] is greater than 0.

[0242] nnpfc_extrapolated_pics_minus1 plus 1 specifies the number of extrapolated pictures generated by NNPF after all input pictures of NNPF in the output order. The value of nnpfc_extrapolated_pics_minus1 must be in the range of 0 to 62.

[0243] The variable NumInpPicsInOutputTensor, which specifies the number of pictures existing in the output tensor of the NNPF that have corresponding input pictures; the variable InpIdx[ idx ], which specifies the input picture index for a list of input pictures arranged in reverse order of output (i.e., the input picture index of the idx-th picture existing in the output tensor of the NNPF that has a corresponding input picture); and the variable numPicsInOutputTensor, which specifies the total number of pictures existing in the output tensor of the NNPF, are derived as shown in Equation 1 below:

[0244] [Equation 1]

[0245]

[0246] nnpfc_spatial_extrapolation_left_offset, nnpfc_spatial_extrapolation_right_offset, nnpfc_spatial_extrapolation_top_offset and nnpfc_spatial_extrapolation_bottom_offset specify the spatial extrapolation area. When each of these offset values ​​is 0 or greater, the luminance samples with horizontal picture coordinates from outSubWidthC * nnpfc_spatial_extrapolation_left_offset to nnpfcOutputPicWidth1 - (outSubWidthC * nnpfc_spatial_extrapolation_right_offset) and vertical picture coordinates from outSubHeightC * nnpfc_spatial_extrapolation_top_offset to nnpfcOutputPicHeight1 - (outSubHeightC * nnpfc_spatial_extrapolation_bottom_offset) correspond to the spatial region of the input picture. The values ​​of nnpfc_spatial_extrapolation_left_offset, nnpfc_spatial_extrapolation_right_offset, nnpfc_spatial_extrapolation_top_offset, and nnpfc_spatial_extrapolation_bottom_offset must be in the range of -65,536 to 65,536. At least one of them must be greater than 0.

[0247] If nnpfc_component_last_flag is 1, it indicates that the last dimension of the input tensor inputTensor and the output tensor outputTensor for the NNPF is used for the current channel. If nnpfc_component_last_flag is 0, it indicates that the third dimension of the input tensor and output tensor is used for the current channel.

[0248] The first dimension of the input and output tensors is used for the batch index, which is a common practice in some neural network frameworks. The equations in the semantics of this SEI message use a batch size with a batch index of 0, but the batch size used as input for the neural network inference process is determined by the post-processing implementation.

[0249] For example, when nnpfc_inp_order_idc is 3 and nnpfc_auxiliary_inp_idc is 1, the input tensor has 7 channels, which include four luminance matrices, two chroma matrices, and one auxiliary input matrix. In this case, the DeriveInputTensors() process derives each of these 7 channels of the input tensor one by one, and when a specific channel among these channels is processed, that channel is referred to as the "current channel" during the process.

[0250] nnpfc_inp_format_idc indicates a method for converting sample values ​​of an input picture into input values ​​for an NNPF. The value of nnpfc_inp_format_idc must be in the range of 0 to 255. An nnpfc_inp_format_idc value of 2 to 255 is reserved for future use and may not exist in a bitstream suitable for the version according to the present disclosure.

[0251] If nnpfc_inp_format_idc is 0, the input value for NNPF is a real number, and the functions InpY() and InpC() are defined as follows:

[0252] InpY(x) = x ÷ ( ( 1 << BitDepthY ) - 1 )

[0253] InpC(x) = x ÷ ( ( 1 << BitDepthC ) - 1 )

[0254] If nnpfc_inp_format_idc is 1, the input value for NNPF is an unsigned integer, and the functions InpY() and InpC() are defined as in Equation 2 below:

[0255] [Equation 2]

[0256]

[0257] The variable inpTensorBitDepthY is derived from the syntax element nnpfc_inp_tensor_luma_bitdepth_minus8 defined below. The variable inpTensorBitDepthC is derived from the syntax element nnpfc_inp_tensor_chroma_bitdepth_minus8 defined below.

[0258] If nnpfc_auxiliary_inp_idc is greater than 0, it indicates that auxiliary input data exists in the input tensor of the NNPF. If nnpfc_auxiliary_inp_idc is 0, it indicates that auxiliary input data does not exist in the input tensor. If nnpfc_auxiliary_inp_idc is 1, 2, or 3, it specifies that auxiliary input data is derived as defined in the equation below.

[0259] If nnpfc_auxiliary_inp_idc is 2 or 3, nnpfc_spatial_extrapolation_prompt_present_flag must be 1.

[0260] The value of nnpfc_auxiliary_inp_idc must be in the range of 0 to 255. An nnpfc_auxiliary_inp_idc value of 2 to 255 may be reserved for future use and may not exist in a bitstream suitable for the version according to the present disclosure.

[0261] When nnpfc_auxiliary_inp_idc is 1, the auxiliary input data consists of strengthControlScaledVal[ i ].

[0262] When nnpfc_auxiliary_inp_idc is 2, the auxiliary input data consists of nnpfc_prompt character values.

[0263] When nnpfc_auxiliary_inp_idc is 3, the auxiliary input data consists of strengthControlScaledVal[ i ] and nnpfc_prompt character values.

[0264] If nnpfc_inband_prompt_flag is 1, it specifies that the text prompt string to be included in the input tensor is included in this NNPFC SEI message or in the NNPFA SEI message that enables the NNPF defined by this NNPFC SEI message. If nnpfc_inband_prompt_flag is 0, it specifies that the text prompt string to be included in the input tensor is provided to the decoding system by external means.

[0265] nnpfc_alignment_zero_bit_c must be equal to 0.

[0266] nnpfc_prompt specifies a text string prompt used as input for NNPF, for example, to generate the contents of a spatially extrapolated image region. If nnpfc_prompt exists, nnpfc_prompt must not be a null string.

[0267] The variable nnpfcPrompt, which specifies the text prompt string to be provided to the input tensor for a specific picture picA where NNPF is enabled, is derived as follows:

[0268] - If nnpfc_inband_prompt_flag is 1 and nnpfa_prompt_update_flag is 1, nnpfcPrompt is set to nnpfa_prompt.

[0269] - Otherwise, if nnpfc_inband_prompt_flag is 1 and nnpfa_prompt_update_flag is 0, nnpfcPrompt is set to nnpfc_prompt.

[0270] - Otherwise, if nnpfc_inband_prompt_flag is 0 and a text prompt string is provided by an external means, nnpfcPrompt is set to that text prompt string.

[0271] - Otherwise (if nnpfc_inband_prompt_flag is 0 and no text prompt string is provided by external means), nnpfcPrompt is set to a null string.

[0272] nnpfc_inp_order_idc represents a method of forming an input tensor for NNPF by sorting an array of samples from the input picture.

[0273] The value of nnpfc_inp_order_idc must be in the range of 0 to 255. An nnpfc_inp_order_idc value of 4 to 255 is reserved for future use and may not exist in a bitstream suitable for the version according to the present disclosure. A decoder suitable for the version according to the present disclosure may ignore NNPFC SEI messages in which nnpfc_inp_order_idc is 4 to 255.

[0274] If ChromaFormatIdc is not 1, nnpfc_inp_order_idc must not be 3.

[0275] If ChromaFormatIdc is 0, nnpfc_inp_order_idc must be 0.

[0276] If ChromaUpsamplingFlag is 1, nnpfc_inp_order_idc must not be 0.

[0277] Table 3 below contains an explanation (informational explanation) of the nnpfc_inp_order_idc value. Figure 7, mentioned in Table 3, is a diagram illustrating an interleaved method for inducing a luminous channel.

[0278] [Table 3]

[0279]

[0280] nnpfc_inp_tensor_bitlength_minus8 + 8 can represent the bit depth of luminance sample values ​​in an input integer tensor. inpTensorBitDepth Y The value of can be derived as follows:

[0281] inpTensorBitDepth Y = nnpfc_inp_tensor_luma_bitdepth_minus8 + 8

[0282] The requirement for bitstream conformity is that the value of nnpfc_inp_tensor_luma_bitdepth_minus8 must be in the range of 0 to 24.

[0283] nnpfc_inp_tensor_chroma_bitdepth_minus8 + 8 can represent the bit depth of chroma sample values ​​in an input integer tensor. inpTensorBitDepth C The value of can be derived as follows:

[0284] inpTensorBitDepth C = nnpfc_inp_tensor_chroma_bitdepth_minus8 + 8

[0285] The requirement for bitstream conformity is that the value of nnpfc_inp_tensor_chroma_bitdepth_minus8 must be in the range of 0 to 24.

[0286] When nnpfc_out_format_idc is equal to 0, the sample value output by NNPF is a real number, and indicates that the range of values ​​from 0 to 1 is linearly mapped to the range of unsigned integer values ​​from 0 to (1 << bitDepth) - 1 for any desired bit depth bitDepth for subsequent post-processing or display.

[0287] When nnpfc_out_format_idc is equal to 1, it indicates that the luminance sample value output by NNPF is an unsigned integer with a range of 0 to (1 << outTensorBitDepthY) - 1, and the chroma sample value output by NNPF is an unsigned integer with a range of 0 to (1 << outTensorBitDepthC) - 1.

[0288] The value of nnpfc_out_format_idc must be in the range of 0 to 255.

[0289] The case where the value of nnpfc_out_format_idc is 2 or greater and 255 or less is reserved for future definition and may not exist in a bitstream suitable for the version according to the present disclosure. A decoder suitable for the version according to the present disclosure may ignore NNPFC SEI messages in which nnpfc_out_format_idc is in the range of 2 or greater and 255 or less.

[0290] nnpfc_out_order_idc indicates the output order of samples generated from NNPF.

[0291] The value of nnpfc_out_order_idc must be in the range of 0 to 255.

[0292] If the value of nnpfc_out_order_idc is 4 or greater and 255 or less, it is reserved for future use and may not exist in a bitstream suitable for the version according to the present disclosure. A decoder suitable for the version according to the present disclosure may ignore NNPFC SEI messages in which nnpfc_out_order_idc is in the range of 4 or greater and 255 or less.

[0293] If ChromaUpsamplingFlag is 1, nnpfc_out_order_idc must not be equal to 0 or 3.

[0294] If ColourizationFlag is 1, nnpfc_out_order_idc must not be equal to 0.

[0295] Table 4 below contains descriptions (informative descriptions) of the nnpfc_out_order_idc values.

[0296] [Table 4]

[0297]

[0298] nnpfc_out_tensor_luma_bitdepth_minus8 + 8 specifies the bit depth of the lumina sample values ​​in the output integer tensor. The value of nnpfc_out_tensor_luma_bitdepth_minus8 must be in the range of 0 to 24.

[0299] The variable outTensorBitDepthY is derived as follows:

[0300] outTensorBitDepthY = nnpfc_out_tensor_luma_bitdepth_minus8 + 8

[0301] nnpfc_out_tensor_chroma_bitdepth_minus8 + 8 specifies the bit depth of the chroma sample values ​​within the output integer tensor. The value of nnpfc_out_tensor_chroma_bitdepth_minus8 must be in the range of 0 to 24. The variable outTensorBitDepthC is derived as follows:

[0302] outTensorBitDepthC = nnpfc_out_tensor_chroma_bitdepth_minus8 + 8

[0303] If BitDepthUpsamplingFlag is 1, the value of nnpfc_out_format_idc must be 1, and at least one of the following conditions must be true:

[0304] - nnpfc_out_tensor_luma_bitdepth_minus8 exists and outTensorBitDepthY is greater than BitDepthY.

[0305] - nnpfc_out_tensor_chroma_bitdepth_minus8 exists and outTensorBitDepthC is greater than BitDepthC.

[0306] If nnpfc_inp_tensor_luma_bitdepth_minus8, nnpfc_inp_tensor_chroma_bitdepth_minus8, nnpfc_out_tensor_luma_bitdepth_minus8 and nnpfc_out_tensor_chroma_bitdepth_minus8 all exist and outTensorBitDepthY is greater than inpTensorBitDepthY, then outTensorBitDepthC must not be less than inpTensorBitDepthC. If nnpfc_inp_tensor_luma_bitdepth_minus8, nnpfc_inp_tensor_chroma_bitdepth_minus8, nnpfc_out_tensor_luma_bitdepth_minus8 and nnpfc_out_tensor_chroma_bitdepth_minus8 all exist and outTensorBitDepthC is greater than inpTensorBitDepthC, then outTensorBitDepthY must not be less than inpTensorBitDepthY.

[0307] If nnpfc_separate_colour_description_present_flag is 1, it indicates that the color primary, transfer characteristics, matrix coefficients, and a combination of scaling and offset values ​​applied in association with the matrix coefficients are specified separately within the SEI message syntax structure for the picture generated by NNPF.

[0308] When nnpfc_separate_colour_description_present_flag is 0, it indicates that the combination of color primaries, transfer characteristics, matrix coefficients, and scaling and offset values ​​associated with matrix coefficients applied to the picture generated by NNPF is the same as that implied by the VUI parameters vui_colour_primaries, vui_transfer_characteristics, vui_matrix_coeffs, and vui_full_range_flag directed or inferred for CLVS.

[0309] nnpfc_colour_primaries has the same semantics as the vui_colour_primaries syntax element, but differs in the following respects:

[0310] - nnpfc_colour_primaries specifies the color primaries of the resulting picture after applying the NNPF specified by this SEI message, rather than the color primaries used in CLVS.

[0311] - If nnpfc_colour_primaries does not exist in the NNPFC SEI message, the value of nnpfc_colour_primaries is inferred to be equal to vui_colour_primaries.

[0312] nnpfc_transfer_characteristics has the same semantics as the vui_transfer_characteristics syntax element, but differs in the following respects:

[0313] - nnpfc_transfer_characteristics specifies the transfer characteristics of the resulting picture to which the NNPF specified by this SEI message has been applied, rather than the transfer characteristics used in CLVS.

[0314] - If nnpfc_transfer_characteristics does not exist in the NNPFC SEI message, the value of nnpfc_transfer_characteristics is inferred to be equal to vui_transfer_characteristics.

[0315] nnpfc_matrix_coeffs describes the equations used to derive luminance and chroma signals from green, blue, red, or Y, Z, X bases. The semantics are applied to the resulting picture of the NNPF specified by this SEI message, and are identical to the semantics of MatrixCoefficients specified in Rec. ITU-T H.273 | ISO / IEC 23091-2, except that BitDepthY and BitDepthC are equal to outTensorBitDepthY and outTensorBitDepthC, respectively.

[0316] If nnpfc_matrix_coeffs does not exist in the NNPFC SEI message, the value of nnpfc_matrix_coeffs is inferred to be equal to vui_matrix_coeffs.

[0317] nnpfc_matrix_coeffs must not be equal to 0 unless both of the following two conditions are true:

[0318] -nnpfc_out_tensor_chroma_bitdepth_minus8 is equal to nnpfc_out_tensor_luma_bitdepth_minus8.

[0319] - nnpfc_out_order_idc is 2, outSubHeightC is 1, and outSubWidthC is 1.

[0320] nnpfc_matrix_coeffs must not be equal to 8 unless one of the following conditions is true:

[0321] -nnpfc_out_tensor_chroma_bitdepth_minus8 is equal to nnpfc_out_tensor_luma_bitdepth_minus8.

[0322] -nnpfc_out_tensor_chroma_bitdepth_minus8 is equal to nnpfc_out_tensor_luma_bitdepth_minus8 + 1, nnpfc_out_order_idc is 2, outSubHeightC is 1, and outSubWidthC is 1.

[0323] nnpfc_full_range_flag represents the scaling and offset values ​​applied in association with the matrix coefficients specified by nnpfc_matrix_coeffs. Its semantics are identical to the VideoFullRangeFlag parameter defined in Rec. ITU-T H.273 | ISO / IEC 23091-2. If nnpfc_full_range_flag is not present, its value is inferred to be 0.

[0324] If nnpfc_chroma_loc_info_present_flag is 1, it indicates that the nnpfc_chroma_sample_loc_type_frame syntax element exists in the NNPFC SEI message. If nnpfc_chroma_loc_info_present_flag is 0, it indicates that the corresponding syntax element does not exist. If nnpfc_chroma_loc_info_present_flag does not exist, its value is inferred to be 0. If ColourizationFlag is 0 or nnpfc_out_colour_format_idc is not 1, the value of nnpfc_chroma_loc_info_present_flag must be 0.

[0325] If nnpfc_chroma_sample_loc_type_frame is not 6 and nnpfc_out_colour_format_idc is 1, nnpfc_chroma_sample_loc_type_frame specifies the chroma sample location of the output picture. If nnpfc_chroma_sample_loc_type_frame is 6 and nnpfc_out_colour_format_idc is 1, it indicates that the location of the chroma sample is unknown, unspecified, or specified by other means not specified herein. The value of nnpfc_chroma_sample_loc_type_frame must be in the range of 0 to 6.

[0326] nnpfc_overlap represents the number of horizontal and vertical sample overlaps between adjacent input tensors of NNPF. The value of nnpfc_overlap must be in the range of 0 to 16,383. If SpatialExtrapolationFlag is 1, the value of nnpfc_overlap is inferred to be equal to 0.

[0327] When nnpfc_constant_patch_size_flag is 1, it indicates that NNPF accepts only the patch size specified by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1 as input. When nnpfc_constant_patch_size_flag is 0, NNPF indicates that it accepts an arbitrary patch size as input, where the width of the extended patch (i.e., the patch including the overlapping area) is inpPatchWidth + 2 * nnpfc_overlap, and this value is a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 * nnpfc_overlap, and the height of the extended patch is inpPatchHeight + 2 * nnpfc_overlap, and this value is a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 * nnpfc_overlap.

[0328] If SpatialExtrapolationFlag is 1, the value of nnpfc_constant_patch_size_flag is inferred to be equal to 1.

[0329] When nnpfc_constant_patch_size_flag is 1, nnpfc_patch_width_minus1 + 1 represents the number of horizontal samples of the patch size required for the NNPF input. The value of nnpfc_patch_width_minus1 must be in the range of 0 to Min(32,766, CroppedWidth - 1).

[0330] When nnpfc_constant_patch_size_flag is 1, nnpfc_patch_height_minus1 + 1 represents the number of vertical samples of the patch size required for the NNPF input. The value of nnpfc_patch_height_minus1 must be in the range of 0 to Min(32,766, CroppedHeight - 1).

[0331] When nnpfc_constant_patch_size_flag is 0, nnpfc_extended_patch_width_cd_delta_minus1 plus 1 and 2 * nnpfc_overlap represents the common divisor of all allowed values ​​for the extended patch width required for the NNPF input. The value of nnpfc_extended_patch_width_cd_delta_minus1 must be in the range of 0 to Min(32,766, CroppedWidth - 1).

[0332] When nnpfc_constant_patch_size_flag is 0, nnpfc_extended_patch_height_cd_delta_minus1 plus 1 and 2 * nnpfc_overlap represents the common divisor of all allowed values ​​for the extended patch height required for the NNPF input. The value of nnpfc_extended_patch_height_cd_delta_minus1 must be in the range of 0 to Min(32,766, CroppedHeight - 1).

[0333] The variables inpPatchWidth and inpPatchHeight represent the width and height of the patch size, respectively.

[0334] If nnpfc_constant_patch_size_flag is 0, the following applies:

[0335] - The values ​​of inpPatchWidth and inpPatchHeight are provided by external means not specified in this disclosure or are set by a post-processor.

[0336] - The value of inpPatchWidth + 2 * nnpfc_overlap must be a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 * nnpfc_overlap, and inpPatchWidth must be less than or equal to CroppedWidth.

[0337] - The value of inpPatchHeight + 2 * nnpfc_overlap must be a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 * nnpfc_overlap, and inpPatchHeight must be less than or equal to CroppedHeight.

[0338] Otherwise (when nnpfc_constant_patch_size_flag is 1), the value of inpPatchWidth is set to nnpfc_patch_width_minus1 + 1 and the value of inpPatchHeight is set to nnpfc_patch_height_minus1 + 1.

[0339] The variables outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, and outPatchCHeight are derived as follows:

[0340] outPatchWidth = (nnpfcOutputPicWidth * inpPatchWidth) / CroppedWidth

[0341] outPatchHeight = (nnpfcOutputPicHeight * inpPatchHeight) / CroppedHeight

[0342] horCScaling = SubWidthC / outSubWidthC

[0343] verCScaling = SubHeightC / outSubHeightC

[0344] outPatchCWidth = outPatchWidth * horCScaling

[0345] outPatchCHeight = outPatchHeight * verCScaling

[0346] As a requirement for bitstream conformance, when SpatialExtrapolation is 0, outPatchWidth * CroppedWidth must be equal to nnpfcOutputPicWidth * inpPatchWidth, and outPatchHeight * CroppedHeight must be equal to nnpfcOutputPicHeight * inpPatchHeight.

[0347] As a requirement for bitstream conformance, when SpatialExtrapolation is 1, outPatchWidth * CroppedWidth must be equal to nnpfcOutputPicWidth1 * inpPatchWidth, and outPatchHeight * CroppedHeight must be equal to nnpfcOutputPicHeight1 * inpPatchHeight.

[0348] nnpfc_padding_type indicates the padding method applied when referencing sample positions outside the boundaries of the input picture, as described in Table 5 below. The value of nnpfc_padding_type must be in the range of 0 to 15. Values ​​of 5 to 15 are reserved for future use and may not exist in bitstreams suitable for the version according to the present disclosure. A decoder suitable for the version according to the present disclosure may ignore NNPFC SEI messages in which nnpfc_padding_type is in the range of 5 to 15.

[0349] [Table 5]

[0350]

[0351] nnpfc_luma_padding_val represents the lumina value to be used for padding when nnpfc_padding_type is 4. The value of nnpfc_luma_padding_val must be in the range of 0 or greater ( 1 << BitDepthY ) - 1 or less.

[0352] nnpfc_cb_padding_val represents the Cb value to be used for padding when nnpfc_padding_type is 4. The value of nnpfc_cb_padding_val must be in the range of 0 or greater ( 1 << BitDepthC ) - 1 or less.

[0353] nnpfc_cr_padding_val represents the Cr value to be used for padding when nnpfc_padding_type is 4. The value of nnpfc_cr_padding_val must be in the range of 0 or greater ( 1 << BitDepthC ) - 1 or less.

[0354] The function InpSampleVal( y, x, picHeight, picWidth, croppedPic, cIdx ) with inputs of vertical sample position y, horizontal sample position x, picture height picHeight, picture width picWidth, sample array croppedPic, and component index cIdx (0 for Luma, 1 for Cb, 2 for Cr) returns a value of sampleVal derived as shown in Equation 3 below:

[0355] In the input of the function InpSampleVal(), vertical positions are listed before horizontal positions to be compatible with the input tensor conventions of some inference engines.

[0356] [Equation 3]

[0357]

[0358] If nnpfc_auxiliary_inp_idc is equal to 1, the variable strengthControlScaledVal is derived as shown in Equation 4 below:

[0359] [Equation 4]

[0360]

[0361] A patch is a rectangular array of samples from the components of a picture (e.g., luminance or chroma components).

[0362] As a requirement for bitstream conformance, outPatchWidth must be > 0. As a requirement for bitstream conformance, outPatchHeight must be > 0.

[0363] The NNPF-generated picture with index i contains sample arrays FilteredYPic[i], FilteredCbPic[i], and FilteredCrPic[i]. The NNPF-generated picture does not contain an overlap region.

[0364] The NNPF process consists of outputting NNPF generated pictures in increasing order of index, all NNPF generated pictures interpolated by NNPF are output, and the NNPF generated pictures corresponding to the input pictures are output as defined in the semantics of the NNPFA SEI message.

[0365] If nnpfc_complexity_info_present_flag is 1, it specifies that there is one or more syntax elements representing the complexity of the NNPF associated with nnpfc_id. If nnpfc_complexity_info_present_flag is 0, it specifies that there are no syntax elements representing the complexity of the NNPF associated with nnpfc_id.

[0366] If nnpfc_parameter_type_idc is 0, it indicates that the neural network uses only integer parameters. If nnpfc_parameter_type_idc is 1, it indicates that the neural network can use floating-point or integer parameters. If nnpfc_parameter_type_idc is 2, it indicates that the neural network uses only binary parameters. If nnpfc_parameter_type_idc is 3, it is reserved for future use and may not exist in a bitstream suitable for the version according to the present disclosure. A decoder suitable for the version according to the present disclosure may ignore NNPFC SEI messages such as nnpfc_parameter_type_idc being 3.

[0367] When nnpfc_log2_parameter_bit_length_minus3 is 0, 1, 2, or 3, it indicates that the neural network does not use parameters with bit lengths greater than 8, 16, 32, or 64, respectively. If nnpfc_parameter_type_idc exists and nnpfc_log2_parameter_bit_length_minus3 does not exist, the neural network does not use parameters with a bit length greater than 1.

[0368] nnpfc_num_parameters_idc represents the maximum number of neural network parameters of the NNPF in powers of 2048. If nnpfc_num_parameters_idc is 0, it indicates that the maximum number of neural network parameters is unknown. The value of nnpfc_num_parameters_idc must be in the range of 0 to 52. Values ​​greater than 52 are reserved for future use and may not exist in bitstreams suitable for the version according to the present disclosure. A decoder suitable for the version according to the present disclosure may ignore NNPFC SEI messages where nnpfc_num_parameters_idc is greater than 52.

[0369] If the value of nnpfc_num_parameters_idc is greater than 0, the variable maxNumParameters is derived as follows:

[0370] maxNumParameters = ( 2 048 << nnpfc_num_parameters_idc ) - 1

[0371] As a requirement for bitstream conformance, the number of neural network parameters of the NNPF must be less than or equal to maxNumParameters.

[0372] If nnpfc_num_kmac_operations_idc is greater than 0, it indicates that the maximum number of multiply-accumulate operations per sample in the NNPF is nnpfc_num_kmac_operations_idc * 1,000 or less. If nnpfc_num_kmac_operations_idc is 0, it indicates that the maximum number of multiply-accumulate operations in the network is unknown. The value of nnpfc_num_kmac_operations_idc is between 0 and 2 32 - It must be within the range of 2 or less.

[0373] If nnpfc_total_kilobyte_size is greater than 0, it indicates the total size (kilobytes) required to store the neural network's uncompressed parameters. The total number of bits is greater than or equal to the sum of the bits used to store each parameter. nnpfc_total_kilobyte_size is the total number of bits divided by 8,000 and rounded up. If nnpfc_total_kilobyte_size is 0, it indicates that the total size required to store the neural network parameters is unknown. The value of nnpfc_total_kilobyte_size is between 0 and 2 32 - It must be within the range of 2 or less.

[0374] If nnpfc_num_metadata_extension_bits is equal to 0, it specifies that nnpfc_application_purpose_tag_uri_present_flag, nnpfc_metadata_alignment_zero_bit, nnpfc_application_purpose_tag_uri, nnpfc_scan_type_idc, nnpfc_for_human_viewing_idc, nnpfc_for_machine_analysis_idc and nnpfc_reserved_metadata_extension do not exist.

[0375] If nnpfc_num_metadata_extension_bits is greater than 0, the variable numSpecifiedMetadataExtensionBits is set to the number of bits representing all syntax elements between nnpfc_num_metadata_extension_bits and nnpfc_reserved_metadata_extension. nnpfc_num_metadata_extension_bits being greater than 0 specifies the sum of numSpecifiedMetadataExtensionBits and the length (in bits) of nnpfc_reserved_metadata_extension.

[0376] The value of nnpfc_num_metadata_extension_bits must be in the range of numSpecifiedMetadataExtensionBits or greater and 2048 or less. Values ​​of numSpecifiedMetadataExtensionBits + 1 or greater and 2048 or less are reserved for future use and may not exist in bitstreams suitable for versions according to the present disclosure.

[0377] If nnpfc_application_purpose_tag_uri_present_flag is 1, it indicates that the nnpfc_application_purpose_tag_uri syntax element is present in this NNPFC SEI message. If nnpfc_application_purpose_tag_uri_present_flag is 0, it indicates that the syntax element is not present. If it is not present, nnpfc_application_purpose_tag_uri_present_flag is inferred to be equal to 0.

[0378] nnpfc_metadata_alignment_zero_bit must be equal to 0.

[0379] nnpfc_application_purpose_tag_uri specifies a tag URI with the syntax and semantics defined in IETF RFC 4151 to identify the application determined purpose of NNPF when nnpfc_purpose is equal to 0.

[0380] The nnpfc_application_purpose_tag_uri enables the unique identification of NNPF application determination purposes without a central registration authority.

[0381] If nnpfc_scan_type_idc is 0, it indicates that the preferred display method for pictures output by the NNPF is unknown, unspecified, or specified by external means. If nnpfc_scan_type_idc is 1, it indicates that the pictures output by the NNPF are suitable for a display method using overscan. If nnpfc_scan_type_idc is 2, it indicates that the pictures output by the NNPF contain visually important information across the entire area up to the picture edges and should not be displayed using overscan; instead, they should be displayed using an exact match between the display area and the edges or underscan. Here, "overscan" refers to a display process where a portion near the picture boundaries becomes invisible in the display area. "Underscan" refers to a display process where the entire picture is visible in the display area but does not cover the entire display area. In display processing that does not use either overscan or underscan, the display area exactly matches the picture area. The value of nnpfc_scan_type_idc must not be equal to 2. If it does not exist, the value of nnpfc_scan_type_idc is inferred to be equal to 0.

[0382] If nnpfc_for_human_viewing_idc is 3, it specifies that the intended optimal use of the NNPF process result video includes human viewing. If nnpfc_for_human_viewing_idc is 2, it specifies that the result video is suitable for human viewing but is not specifically optimized for human viewing. If nnpfc_for_human_viewing_idc is 1, it specifies that the result video is unsuitable for human viewing. If nnpfc_for_human_viewing_idc is 0, it specifies that it is unknown whether the result video is suitable for human viewing. If it does not exist, nnpfc_for_human_viewing_idc is inferred to be equal to 0.

[0383] If nnpfc_for_machine_analysis_idc is 3, it specifies that the intended optimal use of the NNPF process result video includes machine analysis. If nnpfc_for_machine_analysis_idc is 2, it specifies that the result video is suitable for machine analysis but is not specifically optimized for machine analysis. If nnpfc_for_machine_analysis_idc is 1, it specifies that the result video is unsuitable for machine analysis. If nnpfc_for_machine_analysis_idc is 0, it specifies that it is unknown whether the result video is suitable for machine analysis. If it does not exist, nnpfc_for_machine_analysis_idc is inferred to be equal to 0.

[0384] As a requirement for bitstream conformance, the values ​​of nnpfc_for_human_viewing_idc and nnpfc_for_machine_analysis_idc must not both be equal to 1.

[0385] When the decoding system displays video for human viewing, it is suggested to omit all NNPFs where nnpfc_for_human_viewing_idc is 1. When the decoding system performs machine analysis, it is suggested to omit all NNPFs where nnpfc_for_machine_analysis_idc is 1.

[0386] nnpfc_reserved_metadata_extension must not exist in a bitstream suitable for the version according to the present disclosure. However, a decoder suitable for the version according to the present disclosure must ignore the existence and value of nnpfc_reserved_metadata_extension. If it exists, the length (in bits) of nnpfc_reserved_metadata_extension is equal to nnpfc_num_metadata_extension_bits - numSpecifiedMetadataExtensionBits.

[0387] nnpfc_alignment_zero_bit_b must be equal to 0.

[0388] nnpfc_payload_byte[i] contains the i-th byte of a bitstream compliant with ISO / IEC 15938-17. For all existing i values, the byte sequence of nnpfc_payload_byte[i] must be a complete bitstream compliant with ISO / IEC 15938-17.

[0389] Table 6 below shows examples of NNFPA SEI message syntax.

[0390] [Table 6]

[0391]

[0392] The NNPFA SEI message can enable or disable the possible use of the target neural network post-processing filter identified by nnpfa_target_id and nnpfa_base_flag for post-processing filtering of the picture set. For a specific picture with the NNPF enabled, the target NNPF can be derived as follows:

[0393] - If nnpfa_target_base_flag is 1, the target NNPF is a base NNPF with the same nnpfc_id as nnpfa_target_id.

[0394] - Otherwise (when nnpfa_target_base_flag is 0), the target NNPF is the NNPF specified by the last NNPFC SEI message having the same nnpfc_id as nnpfa_target_id, which precedes the first VCL NAL unit of the current picture in the decoding order and is not a repetition of the NNPFC SEI message containing the base NNPF.

[0395] When NNPFs are used for different purposes or filter different color components, multiple NNPFA SEI messages may exist for the same picture.

[0396] nnpfa_target_id may represent a target NNPF specified by one or more NNPFC SEI messages associated with the current picture and having an nnpfc_id identical to nnpfa_target_id. nnpfa_target_id represents the nnpfc_id of a target NNPF, which is specified by one or more NNPFC SEI messages associated with the current picture and having an nnpfc_id identical to nnpfa_target_id. The value of nnpfa_target_id is 0 or greater (2 32 - 2) It must be within the following range.

[0397] An NNPFA SEI message with an nnpfa_target_id of a specific value must not exist in the current PU unless one or more of the following conditions are true:

[0398] - Within the current CLVS, if an NNPFC SEI message exists in a PU that precedes the current PU in decoding order where the nnpfc_id is identical to the corresponding nnpfa_target_id value

[0399] - If an NNPFC SEI message exists in the current PU where nnpfc_id is identical to the corresponding nnpfa_target_id value

[0400] If a PU contains both an NNPFC SEI message with a specific value of nnpfc_id and an NNPFA SEI message where nnpfa_target_id is the same as the nnpfc_id value, the NNPFC SEI message must precede the NNPFA SEI message in the decoding order.

[0401] If nnpfa_cancel_flag is 1, it indicates that the persistence of the target NNPF set by all previous NNPFA SEI messages having the same nnpfa_target_id as the current SEI message is canceled. That is, the target NNPF is no longer used unless it is reactivated by another NNPFA SEI message with the same nnpfa_target_id having nnpfa_cancel_flag of 0.

[0402] If nnpfa_cancel_flag is 0, nnpfa_persistence_flag, nnpfa_target_base_flag, nnpfa_no_prev_clvs_flag, nnpfa_no_foll_clvs_flag (where nnpfa_persistence_flag is 1), and nnpfa_num_output_entries follow.

[0403] nnpfa_persistence_flag specifies the persistence of the target NNPF for the current layer. If nnpfa_persistence_flag is 0, it specifies that the target NNPF can be used for post-processing filtering only for the current picture. If nnpfa_persistence_flag is 1, it specifies that the target NNPF can be used for post-processing filtering for the current picture and all subsequent pictures of the current layer until one or more of the following conditions in the output order become true:

[0404] - When a new CLVS of the current layer starts

[0405] - When the bitstream ends

[0406] - When the picture of the current layer associated with the NNPFA SEI message having the same nnpfa_target_id as the current SEI message is output after the current picture in the output order

[0407] For this subsequent picture, the target NNPF is not applied because it is a picture of the current layer associated with an NNPFA SEI message having the same nnpfa_target_id as the current SEI message.

[0408] The set of pictures associated with the NNPFC SEI message corresponding to the target NNPF is called nnpfcTargetPictures, and the set of pictures in which the target NNPF is currently activated by the NNPFA SEI message is called nnpfaTargetPictures. As a requirement for bitstream conformity, all pictures included in nnpfaTargetPictures must also be included in nnpfcTargetPictures.

[0409] If nnpfa_target_base_flag is 1, it specifies that the target NNPF is a base NNPF with nnpfc_id identical to nnpfa_target_id. If nnpfa_target_base_flag is 0, it indicates that the target NNPF is an NNPF that precedes the first VCL NAL unit of the current picture in decoding order and is specified by the last NNPFC SEI message with nnpfc_id identical to nnpfa_target_id, rather than a repetition of the NNPFC SEI message containing the base NNPF.

[0410] The NNPFA message can activate a base NNPF with a specific nnpfc_id value while the update of the base NNPF is enabled, thereby switching the target NNPF from the updated NNPF to the base NNPF.

[0411] In the case where nnpfa_target_base_flag is 0 in an NNPFA SEI message, there must be at least one NNPFC SEI message preceding that NNPFA SEI message in the decoding order in which nnpfc_id is the same as nnpfa_target_id and nnpfc_base_flag is 0.

[0412] If nnpfa_no_prev_clvs_flag is 1, it specifies that the input pictures of the NNPF do not originate from previous CLVS. If nnpfa_no_prev_clvs_flag is 0, it specifies that the input pictures of the NNPF may or may not originate from previous CLVS.

[0413] If the current CLVS is spliced ​​from another bitstream adjacent to the previous CLVS, and this NNPFA SEI message selects one or more input pictures from one or more previous CLVS, potentially having a negative effect on the output of the target NNPF, the value of nnpfa_no_prev_clvs_flag may be changed from 0 to 1.

[0414] If nnpfa_no_foll_clvs_flag is 1, and this NNPFA SEI message persists to the last PU in the output order of CLVS, this NNPFA SEI message is treated as if it persists within the bitstream to the last PU in the output order of the current layer. If this NNPFA SEI message does not persist to the last PU in the output order of CLVS, or if nnpfa_no_foll_clvs_flag is 0, the value of nnpfa_no_foll_clvs_flag has no specific effect.

[0415] If subsequent CLVS are spliced ​​from another bitstream adjacent to the current CLVS, the value of nnpfa_no_foll_clvs_flag for the picture rate upsampling NNPF can be changed from 0 to 1. As a result, the NNPF process interpolates pictures to the end of the current CLVS using only the input pictures originating from the current CLVS.

[0416] nnpfa_num_output_entries specifies the number of nnpfa_output_flag[i] syntax elements present in the NNPFA SEI message. The value of nnpfa_num_output_entries must be in the range from 0 to NumInpPicsInOutputTensor. If PictureRateUpsamplingFlag is 0 and nnpfa_num_output_entries is equal to NumInpPicsInOutputTensor, nnpfa_output_flag[i] must be 1 for at least one value of i in the range from 0 to nnpfa_num_output_entries - 1.

[0417] nnpfa_output_flag[i] indicates whether the NNPF result picture generated in correspondence with the input picture at index InpIdx[i] is output by the NNPF process activated by the corresponding NNPFA SEI message.

[0418] If nnpfa_output_flag[i] is 1, it specifies that the NNPF result picture generated corresponding to the input picture with index InpIdx[i] is output by the NNPF process activated by this NNPFA SEI message, where the NNPF process is specified in the semantics of the NNPFC SEI message. If nnpfa_output_flag[i] is 0, it specifies that the corresponding NNPF result picture is not output. If nnpfa_num_output_entries is less than NumInpPicsInOutputTensor, nnpfa_output_flag[i] is inferred to be 1 for each value of i in the range from nnpfa_num_output_entries to NumInpPicsInOutputTensor - 1.

[0419] If nnpfa_prompt_update_flag is 1, it specifies that nnpfa_prompt syntax elements exist and nnpfa_alignment_zero_bit syntax elements may exist. If nnpfa_prompt_update_flag is 0, it specifies that nnpfa_prompt syntax elements and nnpfa_alignment_zero_bit syntax elements do not exist. If they do not exist, the value of nnpfa_prompt_update_flag is inferred to be 0.

[0420] If nnpfc_prompt_present_flag is 0, even if nnpfa_prompt_update_flag exists, its value must be 0.

[0421] nnpfa_alignment_zero_bit must be 0.

[0422] nnpfa_prompt specifies the text string prompt used as input for the target NNPF. If nnpfa_prompt_update_flag is 1, nnpfa_prompt must not be a null string. If nnpfa_prompt exists, the nnpfc_prompt text in the DeriveInputTensors() process is replaced with the nnpfa_prompt text.

[0423] nnpfa_num_input_pic_shift specifies the number of input pictures to be shifted from the list of candidate input pictures to obtain the final input pictures for the target NNPF. If none exist, the value of nnpfa_num_input_pic_shift is inferred to be 0. The value of nnpfa_num_input_pic_shift must be in the range of 0 to 63.

[0424] Meanwhile, if nnpfc_input_pic_filtering_flag[ i ] is equal to 1, it indicates that the NNPF generates a corresponding output picture for the i-th input picture. If nnpfc_input_pic_filtering_flag[ i ] is equal to 0, it indicates that the NNPF does not generate a corresponding output picture for the i-th input picture. Each picture generated by the NNPF is stored in the NNPF's output tensor. If nnpfc_num_input_pics_minus1 is equal to 0, the value of nnpfc_input_pic_filtering_flag

[0000] is inferred to be equal to 1. If PictureRateUpsamplingFlag is equal to 0 and nnpfc_num_input_pics_minus1 is greater than 0, then for at least one value of i in the range 0 or greater and nnpfc_num_input_pics_minus1 or less (inclusive), nnpfc_input_pic_filtering_flag[ i ] must be equal to 1.

[0425] A temporal extrapolation NNPF can use only one picture as input, and the NNPF may not filter the input picture(s) unless it has other purposes as well. Therefore, the inference rule “if nnpfc_num_input_pics_minus1 is equal to 0, nnpfc_input_pic_filtering_flag

[0000] is inferred to be equal to 1” is not valid for an NNPF that has only a temporal extrapolation purpose.

[0426] To resolve this issue, the constraint can be updated as follows:

[0427] If nnpfc_input_pic_filtering_flag[ i ] is equal to 1, it indicates that the NNPF generates a corresponding output picture for the i-th input picture. In this disclosure, “i-th” may mean having an index value of i. If nnpfc_input_pic_filtering_flag[ i ] is equal to 0, it indicates that the NNPF does not generate a corresponding output picture for the i-th input picture. Each picture generated by the NNPF is stored in the output tensor of the NNPF. If nnpfc_num_input_pics_minus1 is equal to 0, the value of nnpfc_input_pic_filtering_flag

[0000] is inferred to be equal to 1. If PictureRateUpsamplingFlag is equal to 0 and TemporalExtrapolationFlag is equal to 0, and nnpfc_num_input_pics_minus1 is greater than 0, then for at least one value of i in the range 0 or greater and nnpfc_num_input_pics_minus1 or less (inclusive), nnpfc_input_pic_filtering_flag[ i ] must be equal to 1.

[0428] Meanwhile, the update of the above constraint is valid only when the output picture is not updated while the SEI message is active.

[0429] If the list of output pictures is updated while the SEI message is active, a similar constraint is also required to ensure that temporal extrapolation filtering can be used without requiring additional output pictures to be output beyond what the NNPF requires.

[0430] Accordingly, in a method according to one embodiment, the current constraint in the NNPFA SEI message can be relaxed so that, when the output picture list is updated and the update includes flags for all output pictures associated with the input picture, it includes temporal extrapolation for cases where there is no output picture associated with the input picture.

[0431] Specifically, nnpfa_num_output_entries specifies the number of nnpfa_output_flag[ i ] syntax elements present in the NNPFA SEI message. The value of nnpfa_num_output_entries must be in the range from 0 to NumInpPicsInOutputTensor (inclusive). If PictureRateUpsamplingFlag is equal to 0, TemporalExtrapolationFlag is equal to 0, and nnpfa_num_output_entries is equal to NumInpPicsInOutputTensor, then nnpfa_output_flag[ i ] must be equal to 1 for at least one value of i in the range from 0 to nnpfa_num_output_entries - 1 (inclusive).

[0432] In other words, if PictureRateUpsamplingFlag is 1 or TemporalExtrapolationFlag is 1, the constraint that nnpfa_output_flag[i] must be equal to 1 for at least one value of i does not apply. A PictureRateUpsamplingFlag equal to 0 means that the purpose of the neural network post-processing filter corresponds to picture rate upsampling. A TemporalExtrapolationFlag equal to 1 means that the purpose of the neural network post-processing filter corresponds to temporal extrapolation. When PictureRateUpsamplingFlag is 1, since the purpose of the neural network post-processing filter is to generate an interpolated image that fills the gaps between frames for picture rate upsampling, it may be unnecessary to impose a constraint on the output of the filtered input picture. In addition, when TemporalExtrapolationFlag is 1, since the purpose of the neural network post-processing filter is to generate at least one future picture through temporal extrapolation, it may be unnecessary to constrain the output of the image filtered from the input picture in the same way.

[0433] Accordingly, in one embodiment, the existing constraint can be relaxed by including the condition that the values ​​of PictureRateUpsamplingFlag and TemporalExtrapolationFlag are 0 in the condition that the value of nnpfa_output_flag[ i ] must be 1. As the constraint is relaxed in this way, the constraint that a picture generated based on a neural network post-processing filter must be output is not applied when the value of TemporalExtrapolationFlag is 1, i.e., when the objective of NNPF is temporal extrapolation, and when PictureRateUpsamplingFlag is 1, i.e., when the objective of NNPF is picture rate upsampling.

[0434] FIG. 8 is a diagram illustrating a method for encoding image information according to one embodiment.

[0435] The above description may be applied to the NNPFC SEI message and NNPFA SEI message generated or derived by the method according to one embodiment.

[0436] This can be performed by an encoding device (200) according to one embodiment. Accordingly, all or part of the description of the encoding device (200) described above and the description of the encoding method described with reference to FIG. 5 may also be applied to the method according to one embodiment. That is, the descriptions described above may also be applied to the method according to one embodiment to the extent that they do not conflict with the descriptions described below.

[0437] The encoding device (200) may include a memory and a processor electrically connected to the memory, and the operation of the aforementioned encoding device (200), the encoding method, or the method described below may be executed by the processor of the encoding device (200).

[0438] Terms or names used in this disclosure (e.g., names of syntax elements or names of variables, etc.) are merely examples, and the scope of the embodiments is not limited to these terms. Even if a term is not used in this disclosure, if substantial features such as the function performed, the definition thereof, or the method by which it is derived are identical or similar to those of this disclosure, it may be considered to be included within the scope of the embodiments described in this disclosure.

[0439] In addition, the method according to one embodiment may include other operations in addition to the operations described below, and some of the operations described below may be omitted depending on the example.

[0440] Referring to FIG. 8, a method according to one embodiment may include the step of determining at least one neural network that can be used as a neural network post-processing filter (S1010), the step of determining whether the neural network post-processing filter is activated (S1020), and the step of encoding image information including an NNPFC SEI message containing information regarding at least one neural network that can be used as a neural network post-processing filter and an NNPFA SEI message indicating whether the neural network post-processing filter is activated (S1030).

[0441] As described above, the NNPFC SEI message can specify a neural network that can be used as a post-processing filter. The use or activation of specified post-processing filters (NNPFs) for specific pictures can be indicated using the NNPFA SEI message. The NNPFA SEI message enables or disables the use of the target neural network post-processing filter for a set of pictures identified by nnpfa_target_id and nnpfa_target_base_flag.

[0442] A method according to one embodiment may generate an NNPFC SEI message containing information regarding at least one neural network that can be used as a neural network post-processing filter, and an NNPFA SEI message indicating whether the neural network post-processing filter is activated, based on the results of steps S1010 and S1020, and may encode image information including the generated NNPFC SEI message and NNPFA SEI message.

[0443] The above NNPFA SEI message may include NNPFA output information indicating whether a picture generated based on a neural network post-processing filter corresponding to an input picture is output. For example, the NNPFA output information may be a syntax element represented by nnpfa_output_flag[i]. nnpfa_output_flag[i] indicates whether an NNPF result picture generated corresponding to an input picture having index InpIdx[i] is output by an NNPF process activated by the corresponding NNPFA SEI message.

[0444] If nnpfa_output_flag[i] is 1, it specifies that the NNPF result picture generated corresponding to the input picture with index InpIdx[i] is output by the NNPF process activated by the corresponding NNPFA SEI message, where the NNPF process is specified in the semantics of the NNPFC SEI message. If nnpfa_output_flag[i] is 0, it specifies that the corresponding NNPF result picture is not output.

[0445] NNPFA output information indicates that a picture generated based on the neural network post-processing filter is output based on a condition that the purpose of the neural network post-processing filter does not correspond to temporal extrapolation.

[0446] For example, when the above condition is satisfied, the NNPFA output information may indicate that a picture generated based on the neural network post-processing filter is output. For example, when the NNPFA output information is nnpfa_output_flag[i], the value of nnpfa_output_flag[i] may be 1 when the above condition is satisfied.

[0447] Meanwhile, when the above condition is satisfied, the output of a picture generated based on a neural network post-processing filter, that is, the value of nnpfa_output_flag[i] being 1, may be a constraint. In other words, if the above condition is satisfied, a picture generated based on a neural network post-processing filter must be output. That is, if the above condition is satisfied, the value of nnpfa_output_flag[i] must be 1.

[0448] According to one embodiment, the condition of the constraint that a picture generated based on a neural network post-processing filter must be output includes that the purpose of the neural network post-processing filter does not correspond to temporal extrapolation. That is, if the purpose of the neural network post-processing filter corresponds to temporal extrapolation, this constraint is not applied, so it is possible for the picture generated based on the neural network post-processing filter not to be output. By relaxing unnecessary constraints in this way, coding efficiency can be increased.

[0449] Whether the objective of a neural network post-processing filter corresponds to temporal extrapolation can be determined based on the value of the temporal extrapolation flag. If the value of the temporal extrapolation flag is 1, it indicates that the objective of the neural network post-processing filter corresponds to temporal extrapolation, and if the value of the temporal extrapolation flag is 0, it indicates that the objective of the neural network post-processing filter does not correspond to temporal extrapolation. For example, the temporal extrapolation flag can be represented by the variable TemporalExtrapolationFlag.

[0450] The NNPFC SEI message may include NNPFC purpose information indicating the purpose of the neural network post-processing filter, and the value of the temporal extrapolation flag may be derived based on the NNPFC purpose information. For example, the NNPFC purpose information may be the aforementioned syntax element nnpfc_purpose. Accordingly, the NNPFC purpose information may indicate the purpose of the NNPF according to the definition in Table 2.

[0451] The variables ChromaUpsamplingFlag, ResolutionResamplingFlag, PictureRateUpsamplingFlag, BitDepthUpsamplingFlag, ColourizationFlag, and TemporalExtrapolationFlag, which respectively specify whether nnpfc_purpose represents the purpose of an NNPF including chroma upsampling, resolution resampling, picture rate upsampling, bit depth upsampling, colorization, and temporal extrapolation, are derived as follows:

[0452] ChromaUpsamplingFlag = ((nnpfc_purpose & 0x02) > 0)? 1:0

[0453] ResolutionResamplingFlag = ( ( nnpfc_purpose & 0x04 ) > 0 ) ? 1:0

[0454] PictureRateUpsamplingFlag = ((nnpfc_purpose & 0x08) > 0)? 1:0

[0455] BitDepthUpsamplingFlag = ( ( nnpfc_purpose & 0x10 ) > 0 ) ? 1:0

[0456] ColourizationFlag = ( ( nnpfc_purpose & 0x20 ) > 0 ) ? 1:0

[0457] TemporalExtrapolationFlag = ((nnpfc_purpose & 0x40) > 0)? 1:0

[0458] SpatialExtrapolationFlag = ((nnpfc_purpose & 0x80) > 0)? 1:0

[0459] The condition that a picture generated based on a neural network post-processing filter must be output, that is, the condition that the value of the NNPFA output information must be 1, is applied, can be satisfied based on the fact that the purpose of the neural network post-processing filter does not correspond to picture rate upsampling and temporal extrapolation, and that the number of NNPFA output information present in the NNPFA SEI message is equal to the number of pictures having a corresponding input picture among the pictures included in the NNPF output tensor.

[0460] That is, if the purpose of the neural network post-processing filter corresponds to picture rate upsampling or temporal extrapolation, the constraint that nnpfa_output_flag[i] must be equal to 1 is not applied. In this way, coding efficiency can be improved by relaxing the conditions so that unnecessary constraints are not applied.

[0461] The variable PictureRateUpsamplingFlag can indicate whether the purpose of the neural network post-processing filter corresponds to picture rate upsampling. If the value of PictureRateUpsamplingFlag is 0, it indicates that the purpose of the neural network post-processing filter does not correspond to picture rate upsampling, and if the value of PictureRateUpsamplingFlag is 1, it indicates that the purpose of the neural network post-processing filter corresponds to picture rate upsampling. Similar to TemporalExtrapolationFlag, the value of PictureRateUpsamplingFlag can also be derived based on NNPFC purpose information (e.g., nnpfc_purpose).

[0462] The number of NNPFA output information (e.g., nnpfa_output_flag[ i ]) present in the above NNPFA SEI message can be represented by the syntax element nnpfa_num_output_entries. The number of pictures among the pictures included in the NNPF output tensor that have a corresponding input picture can be represented by the variable NumInpPicsInOutputTensor.

[0463] Therefore, if PictureRateUpsamplingFlag is equal to 0, TemporalExtrapolationFlag is equal to 0, and nnpfa_num_output_entries is equal to NumInpPicsInOutputTensor, then the picture generated based on the neural network post-processing filter should be output. nnpfa_num_output_entries being equal to NumInpPicsInOutputTensor may mean that the NNPFA SEI message has specified output for all input-corresponding pictures (pictures that have corresponding input pictures) present in the output tensor.

[0464] Specifically, if PictureRateUpsamplingFlag is equal to 0, TemporalExtrapolationFlag is equal to 0, and nnpfa_num_output_entries is equal to NumInpPicsInOutputTensor, then for at least one value of i in the range from 0 to nnpfa_num_output_entries - 1 (inclusive), nnpfa_output_flag[i] must be equal to 1.

[0465] A bitstream can be generated or acquired by encoding image information according to the method described above. The generated or acquired bitstream can be stored non-temporarily on a computer-readable storage medium. Additionally, the generated or acquired bitstream can be transmitted externally (e.g., to a decoding device) through a transmitter of a transmission device. At this time, the generation or acquisition of the bitstream can be executed by at least one processor provided in the transmission device.

[0466] FIG. 9 is a diagram illustrating a method for decoding image information according to one embodiment.

[0467] The above description may be applied to the NNPFC SEI message and NNPFA SEI message obtained or derived by the method according to one embodiment.

[0468] A method according to one embodiment may be performed by a decoding device (300) according to one embodiment. Accordingly, all or part of the description of the decoding device (300) described above and the description of the decoding method described with reference to FIG. 4 may also be applied to a method according to one embodiment. That is, the descriptions described above may also be applied to a method according to one embodiment to the extent that they do not conflict with the descriptions described below.

[0469] In addition, the description regarding the NNPFC SEI message and NNPFA SEI message described above may be applied to the method according to one embodiment.

[0470] The decoding device (300) may include a memory and a processor electrically connected to the memory, and the operation of the decoding device (300) described above, the encoding method, or the method described below may be executed by the processor of the decoding device (300).

[0471] Terms or names used in this disclosure (e.g., names of syntax elements or names of variables, etc.) are merely examples, and the scope of the embodiments is not limited to these terms. Even if a term is not used in this disclosure, if substantial features such as the function performed, the definition thereof, or the method by which it is derived are identical or similar to those of this disclosure, it may be considered to be included within the scope of the embodiments described in this disclosure.

[0472] In addition, the method according to one embodiment may include other operations in addition to the operations described below, and some of the operations described below may be omitted depending on the example.

[0473] Referring to FIG. 9, a method according to one embodiment may include the step of obtaining an NNPFC SEI message and an NNPFA SEI message from a bitstream (S1110), the step of determining at least one neural network that can be used as a neural network post-processing filter based on the NNPFC SEI message (S1120), and the step of determining whether to activate the neural network post-processing filter based on the NNPFA SEI message (S1130).

[0474] As described above, the NNPFC SEI message can specify a neural network that can be used as a post-processing filter. The use or activation of specific post-processing filters (NNPFs) for specific pictures can be indicated using the NNPFA SEI message. The NNPFA SEI message enables or disables the use of the target neural network post-processing filter for a set of pictures identified by nnpfa_target_id and nnpfa_target_base_flag.

[0475] The above NNPFA SEI message may include NNPFA output information indicating whether a picture generated based on a neural network post-processing filter corresponding to an input picture is output. For example, the NNPFA output information may be a syntax element represented by nnpfa_output_flag[i]. nnpfa_output_flag[i] indicates whether an NNPF result picture generated corresponding to an input picture having index InpIdx[i] is output by an NNPF process activated by the corresponding NNPFA SEI message.

[0476] If nnpfa_output_flag[i] is 1, it specifies that the NNPF result picture generated corresponding to the input picture with index InpIdx[i] is output by the NNPF process activated by the corresponding NNPFA SEI message, where the NNPF process is specified in the semantics of the NNPFC SEI message. If nnpfa_output_flag[i] is 0, it specifies that the corresponding NNPF result picture is not output.

[0477] NNPFA output information indicates that a picture generated based on the neural network post-processing filter is output based on a condition that the purpose of the neural network post-processing filter does not correspond to temporal extrapolation.

[0478] For example, when the above condition is satisfied, the NNPFA output information may indicate that a picture generated based on the neural network post-processing filter is output. For example, when the NNPFA output information is nnpfa_output_flag[i], the value of nnpfa_output_flag[i] may be 1 when the above condition is satisfied.

[0479] Meanwhile, when the above condition is satisfied, the output of a picture generated based on a neural network post-processing filter, that is, the value of nnpfa_output_flag[i] being 1, may be a constraint. In other words, if the above condition is satisfied, a picture generated based on a neural network post-processing filter must be output. That is, if the above condition is satisfied, the value of nnpfa_output_flag[i] must be 1.

[0480] According to one embodiment, the condition of the constraint that a picture generated based on a neural network post-processing filter must be output includes that the purpose of the neural network post-processing filter does not correspond to temporal extrapolation. That is, if the purpose of the neural network post-processing filter corresponds to temporal extrapolation, this constraint is not applied, so it is possible for the picture generated based on the neural network post-processing filter not to be output. By relaxing unnecessary constraints in this way, coding efficiency can be increased.

[0481] Whether the objective of a neural network post-processing filter corresponds to temporal extrapolation can be determined based on the value of the temporal extrapolation flag. If the value of the temporal extrapolation flag is 1, it indicates that the objective of the neural network post-processing filter corresponds to temporal extrapolation, and if the value of the temporal extrapolation flag is 0, it indicates that the objective of the neural network post-processing filter does not correspond to temporal extrapolation. For example, the temporal extrapolation flag can be represented by the variable TemporalExtrapolationFlag.

[0482] The NNPFC SEI message may include NNPFC purpose information indicating the purpose of the neural network post-processing filter, and the value of the temporal extrapolation flag may be derived based on the NNPFC purpose information. For example, the NNPFC purpose information may be the aforementioned syntax element nnpfc_purpose. Accordingly, the NNPFC purpose information may indicate the purpose of the NNPF according to the definition in Table 2.

[0483] The variables ChromaUpsamplingFlag, ResolutionResamplingFlag, PictureRateUpsamplingFlag, BitDepthUpsamplingFlag, ColourizationFlag, and TemporalExtrapolationFlag, which respectively specify whether nnpfc_purpose represents the purpose of an NNPF including chroma upsampling, resolution resampling, picture rate upsampling, bit depth upsampling, colorization, and temporal extrapolation, are derived as follows:

[0484] ChromaUpsamplingFlag = ((nnpfc_purpose & 0x02) > 0)? 1:0

[0485] ResolutionResamplingFlag = ( ( nnpfc_purpose & 0x04 ) > 0 ) ? 1:0

[0486] PictureRateUpsamplingFlag = ((nnpfc_purpose & 0x08) > 0)? 1:0

[0487] BitDepthUpsamplingFlag = ( ( nnpfc_purpose & 0x10 ) > 0 ) ? 1:0

[0488] ColourizationFlag = ( ( nnpfc_purpose & 0x20 ) > 0 ) ? 1:0

[0489] TemporalExtrapolationFlag = ((nnpfc_purpose & 0x40) > 0)? 1:0

[0490] SpatialExtrapolationFlag = ((nnpfc_purpose & 0x80) > 0)? 1:0

[0491] The condition that a picture generated based on a neural network post-processing filter must be output, that is, the condition that the value of the NNPFA output information must be 1, is applied, can be satisfied based on the fact that the purpose of the neural network post-processing filter does not correspond to picture rate upsampling and temporal extrapolation, and that the number of NNPFA output information present in the NNPFA SEI message is equal to the number of pictures having a corresponding input picture among the pictures included in the NNPF output tensor.

[0492] That is, if the purpose of the neural network post-processing filter corresponds to picture rate upsampling or temporal extrapolation, the constraint that nnpfa_output_flag[i] must be equal to 1 is not applied. In this way, coding efficiency can be improved by relaxing the conditions so that unnecessary constraints are not applied.

[0493] The variable PictureRateUpsamplingFlag can indicate whether the purpose of the neural network post-processing filter corresponds to picture rate upsampling. If the value of PictureRateUpsamplingFlag is 0, it indicates that the purpose of the neural network post-processing filter does not correspond to picture rate upsampling, and if the value of PictureRateUpsamplingFlag is 1, it indicates that the purpose of the neural network post-processing filter corresponds to picture rate upsampling. Similar to TemporalExtrapolationFlag, the value of PictureRateUpsamplingFlag can also be derived based on NNPFC purpose information (e.g., nnpfc_purpose).

[0494] The number of NNPFA output information (e.g., nnpfa_output_flag[ i ]) present in the above NNPFA SEI message can be represented by the syntax element nnpfa_num_output_entries. The number of pictures among the pictures included in the NNPF output tensor that have a corresponding input picture can be represented by the variable NumInpPicsInOutputTensor.

[0495] Therefore, if PictureRateUpsamplingFlag is equal to 0, TemporalExtrapolationFlag is equal to 0, and nnpfa_num_output_entries is equal to NumInpPicsInOutputTensor, then the picture generated based on the neural network post-processing filter should be output. nnpfa_num_output_entries being equal to NumInpPicsInOutputTensor may mean that the NNPFA SEI message has specified output for all input-corresponding pictures (pictures that have corresponding input pictures) present in the output tensor.

[0496] Specifically, if PictureRateUpsamplingFlag is equal to 0, TemporalExtrapolationFlag is equal to 0, and nnpfa_num_output_entries is equal to NumInpPicsInOutputTensor, then for at least one value of i in the range from 0 to nnpfa_num_output_entries - 1 (inclusive), nnpfa_output_flag[i] must be equal to 1.

[0497] When a neural network post-processing filter and whether it is activated are determined according to the method described above, the neural network post-processing filter can be applied to the input picture according to the result.

[0498] FIG. 10 is a diagram illustrating an exemplary content streaming system to which an embodiment according to the present disclosure can be applied.

[0499] As illustrated in FIG. 10, a content streaming system to which an embodiment of the present disclosure is applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0500] The above encoding server compresses content input from multimedia input devices, such as smartphones, cameras, and camcorders, into digital data to generate a bitstream and transmits it to the streaming server. As another example, if multimedia input devices, such as smartphones, cameras, and camcorders, generate the bitstream directly, the encoding server may be omitted.

[0501] The bitstream above may be generated by a video encoding method and / or video encoding device to which an embodiment of the present disclosure is applied, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0502] The streaming server transmits multimedia data to a user device based on a user request through a web server, and the web server can act as a medium to inform the user of available services. When a user requests a desired service from the web server, the web server transmits it to the streaming server, and the streaming server can transmit multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server can perform the role of controlling commands and responses between each device within the content streaming system.

[0503] The streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, to provide a seamless streaming service, the streaming server can store the bitstream for a certain period of time.

[0504] Examples of the above user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs (head-mounted displays)), digital TVs, desktop computers, digital signage, etc.

[0505] Each server within the above-mentioned content streaming system can be operated as a distributed server, and in this case, data received from each server can be processed in a distributed manner.

[0506] The scope of the present disclosure includes software or machine-executable instructions (e.g., operating system, application, firmware, program, etc.) that enable an operation according to a method of various embodiments to be executed on a device or computer, and a non-transitory computer-readable medium on which such software or instructions, etc. are stored and executable on a device or computer.

[0507] An embodiment according to the present disclosure can be used to encode / decode images.

Claims

1. A step of acquiring NNPFC (neural-network post-filter characteristics) SEI messages and NNPFA (neural-network post-filter activation) SEI messages from a bitstream; A step of determining at least one neural network that can be used as a neural network post-processing filter based on the above NNPFC SEI message; and Based on the above NNPFA SEI message, the method includes a step of determining whether to activate the neural network post-processing filter, and The above NNPFA SEI message is, It includes NNPFA output information indicating whether a picture generated based on the neural network post-processing filter is output in correspondence with the input picture, and The above NNPFA output information is, A method indicating that a picture generated based on the neural network post-processing filter is output based on a condition including that the purpose of the neural network post-processing filter does not correspond to temporal extrapolation.

2. In Paragraph 1, The above conditions are, A method in which satisfaction is determined based on the value of a temporal extrapolation flag indicating whether the purpose of the above neural network post-processing filter corresponds to the above temporal extrapolation.

3. In Paragraph 2, The above NNPFC SEI message is, It includes NNPFC objective information indicating the objective of the above neural network post-processing filter, and The value of the above temporal extrapolation flag is, A method derived based on the above NNPFC objective information.

4. In Paragraph 1, The above conditions are, A method in which the purpose of the above neural network post-processing filter does not correspond to picture rate upsampling and the above temporal extrapolation, and is satisfied based on the fact that the number of NNPFA output information present in the NNPFA SEI message is equal to the number of pictures having corresponding input pictures among the pictures included in the NNPF output tensor.

5. A step of determining at least one neural network that can be used as a neural network post-processing filter; A step of determining whether to activate the above neural network post-processing filter; and The method includes the step of encoding image information including an NNPFC (neural-network post-filter characteristics) SEI message containing information regarding at least one neural network that can be used as a neural network post-processing filter, and an NNPFA (neural-network post-filter activation) SEI message indicating whether the neural network post-processing filter is activated. The above NNPFA SEI message is, It includes NNPFA output information indicating whether a picture generated based on the neural network post-processing filter corresponding to the input picture is output, and The above NNPFA output information is, A method indicating that a picture generated based on the neural network post-processing filter is output based on a condition including that the purpose of the neural network post-processing filter does not correspond to temporal extrapolation.

6. In Paragraph 5, The above conditions are, A method in which satisfaction is determined based on the value of a temporal extrapolation flag indicating whether the purpose of the above neural network post-processing filter corresponds to the above temporal extrapolation.

7. In Paragraph 6, The above NNPFC SEI message is, It includes NNPFC objective information indicating the objective of the above neural network post-processing filter, and The value of the above temporal extrapolation flag is, A method derived based on the above NNPFC objective information.

8. In Paragraph 5, The above conditions are, A method in which the purpose of the above neural network post-processing filter does not correspond to picture rate upsampling and the above temporal extrapolation, and is satisfied based on the fact that the number of NNPFA output information present in the NNPFA SEI message is equal to the number of pictures having corresponding input pictures among the pictures included in the NNPF output tensor.

9. A computer-readable storage medium that non-transiently stores a bitstream generated by the method of claim 5 above.

10. Step of generating a bitstream; and The method includes the step of transmitting data including the bitstream above; and The step of generating the above bitstream is, A step of determining at least one neural network that can be used as a neural network post-processing filter; A step of determining whether to activate the above neural network post-processing filter; and The method includes the step of encoding image information including an NNPFC (neural-network post-filter characteristics) SEI message containing information regarding at least one neural network that can be used as a neural network post-processing filter, and an NNPFA (neural-network post-filter activation) SEI message indicating whether the neural network post-processing filter is activated. The above NNPFA SEI message is, It includes NNPFA output information indicating whether a picture generated based on the neural network post-processing filter corresponding to the input picture is output, and The above NNPFA output information is, A method indicating that a picture generated based on the neural network post-processing filter is output based on a condition including that the purpose of the neural network post-processing filter does not correspond to temporal extrapolation.