Image encoding / decoding method, recording medium having stored bit stream therein, and method for transmitting bit stream
By acquiring and parsing NNPFC and NNPFA SEI messages, the activation status of the neural network post-processing filter is determined, which solves the problem of high transmission and storage costs in high-resolution image encoding and realizes efficient image encoding and decoding.
Patent Information
- Application Number
- CN202480005193.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-05
- Filing Date
- 2024-07-05
- Publication Date
- 2025-07-25
AI Technical Summary
In the encoding and decoding process of high resolution and high quality images, the prior art has problems with high transmission and storage costs due to the increase in the amount of information, and it is difficult to effectively determine whether the basic filter or the update filter is activated.
By acquiring the neural network post-filter characteristics (NNPFC) and neural network post-filter activation (NNPFA) SEI messages, the activation state of the neural network post-processing filter is determined, and this information is encoded and decoded in the bitstream to improve encoding/decoding efficiency.
More efficient image encoding and decoding is realized, the activation state of the filter is clarified, transmission and storage costs are reduced, and storage methods are provided for non-transitory computer-readable recording media.
Smart Images

Figure CN120380751A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image decoding method, an image encoding method, a method of transmitting a bitstream, and a method of determining whether to activate a neural network post-filter. Background Art
[0002] Recently, the demand for high-resolution and high-quality images such as high-definition (HD) images and ultra-high-definition (UHD) images has been increasing in various fields. As the resolution and quality of image data are increased, the amount of information or bits to be transmitted relatively increases compared to existing image data. The increase in the amount of information or bits to be transmitted results in an increase in transmission cost and storage cost.
[0003] Therefore, there is a need for efficient image compression techniques to effectively transmit, store, and reproduce information on high-resolution and high-quality images. Summary of the Invention
[0004] Technical Problem
[0005] The present invention relates to providing an image encoding / decoding method and apparatus having improved encoding / decoding efficiency.
[0006] The present disclosure also relates to providing a method of clearly determining a filter to be activated between a base filter and an updated filter.
[0007] The present disclosure also relates to providing a method of preventing a situation where it is unclear whether to activate a base filter or an updated filter.
[0008] The present disclosure also relates to providing a non-transitory computer-readable recording medium for storing a bitstream generated using the image encoding method according to the present disclosure.
[0009] The present disclosure also relates to providing a non-transitory computer-readable recording medium for storing a bitstream received and decoded by the image decoding apparatus according to the present disclosure and used for image reconstruction.
[0010] The present disclosure also relates to providing a method of transmitting a bitstream generated using the image encoding method according to the present disclosure.
[0011] The technical objects of the present disclosure are not limited to those described above, and other technical objects not described above will be clearly understood by those skilled in the art to which the present disclosure pertains from the following description.
[0012] Technical Solution
[0013] An image decoding method performed by an image decoding device according to an aspect of the present disclosure may include: obtaining a neural network post-filter characteristic (NNPFC) supplementary enhancement information (SEI) message and a neural network post-filter activation (NNPFA) SEI message; determining at least one neural network to be used as a neural network post-filter (NNPF) based on the NNPFC SEI message; and determining whether to activate a target NNPF to be applied to a current picture based on the NNPFA SEI message, where the NNPFA SEI message includes target identification information and target base flag information, where the target NNPF is determined based on the target identification information and the target base flag information, where based on the target base flag information not indicating that the target NNPF is a base NNPF, there is at least one NNPFC SEI message in decoding order before the NNPFA SEI message, where the identification information is equal to the target identification information and the value of the base flag information is equal to the value of the target base flag information.
[0014] An image encoding method performed by an image encoding device according to an aspect of the present disclosure may include: encoding at least one neural network that can be used as a post-processing filter into a neural network post-filter characteristic (NNPFC) SEI message; and encoding whether a target NNPF that can be applied to a current picture is activated into a neural network post-filter activation (NNPFA) SEI message, where the NNPFA SEI message includes target identification information and target base flag information, where based on the target base flag information not indicating that the target NNPF is a base NNPF, there is at least one NNPFC SEI message in decoding order before the NNPFA SEI message, where the identification information is equal to the target identification information and the value of the base flag information is equal to the value of the target base flag information.
[0015] A computer-readable digital storage medium according to an aspect of the present disclosure may store a bitstream generated by an image encoding method or an image encoding device.
[0016] A sending method according to an aspect of the present disclosure may send a bitstream generated by an image encoding method or an image encoding device.
[0017] The features of the present disclosure outlined above are merely illustrative aspects of the detailed description of the present disclosure and do not limit the scope of the present disclosure.
[0018] Beneficial effects
[0019] According to the present disclosure, an image encoding / decoding method and device with improved encoding / decoding efficiency can be provided.
[0020] According to the present disclosure, the applied neural network post-filter can also be clarified among neural network post-filters that supplement enhancement information (SEI) messages with several previous neural network post-filter characteristics (NNPFC).
[0021] According to the present disclosure, there can also be provided a non-transitory computer-readable recording medium for storing a bitstream generated using the image encoding method according to the present disclosure.
[0022] According to the present disclosure, there can also be provided a non-transitory computer-readable recording medium for storing a bitstream received and decoded by the image decoding device according to the present disclosure and used for image reconstruction.
[0023] According to the present disclosure, there can also be provided a method for transmitting a bitstream generated using an image encoding method.
[0024] The effects of the present disclosure are not limited to the effects described above, and from the following description, those skilled in the technical field to which the present disclosure pertains will clearly understand other effects not described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a view schematically showing a video compilation system to which an embodiment of the present disclosure is applicable.
[0026] Figure 2 is a diagram schematically showing an image encoding device to which an embodiment of the present disclosure can be applied.
[0027] Figure 3 is a schematic diagram showing an image decoding device to which an embodiment of the present disclosure can be applied.
[0028] Figure 4 Exemplarily shows the hierarchical structure of a compiled video / image to which an embodiment of the present disclosure can be applied.
[0029] Figure 5 is a diagram for explaining an interleaving method for deriving a luminance channel.
[0030] Figure 6 is a diagram showing a problem situation that may occur when a target NNPF to be activated is identified based on a neural network post-filter activation (NNPFA) SEI message.
[0031] Figure 7 is an example of a situation where an embodiment of the present invention is applied.
[0032] Figure 8 is another example of a situation where an embodiment of the present invention is applied, and Figure 9 is an example of an image encoding method to which an embodiment of the present disclosure is applicable.
[0033] Figure 10 It is an example of an image decoding method to which the embodiments of the present disclosure are applicable.
[0034] Figure 11 It shows the process of determining the target NNPF according to an embodiment of the present disclosure.
[0035] Figure 12 It shows the process of determining the target NNPF according to an embodiment of the present disclosure.
[0036] Figure 13 It is a diagram illustrating a content streaming system to which the embodiments of the present disclosure can be applied. Detailed implementation manners
[0037] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that they can be easily implemented by those skilled in the art. However, the present disclosure can be implemented in various different forms and is not limited to the embodiments described herein.
[0038] In the description of the present disclosure, if the detailed description of relevant known functions or configurations makes the scope of the present disclosure unnecessarily ambiguous, its detailed description will be omitted. In the drawings, parts irrelevant to the description of the present disclosure are omitted, and similar reference numerals are attached to similar parts.
[0039] In the present disclosure, when a component is "connected", "coupled" or "linked" to another component, it may include not only a direct connection relationship but also an indirect connection relationship with an intermediate component present. In addition, when a component "includes" or "has" other components, unless otherwise stated, this means that other components can be further included, rather than excluding other components.
[0040] In the present disclosure, unless otherwise stated, terms such as first, second, etc. may be used only for the purpose of distinguishing one component from other components, and do not limit the order or importance of the components. Therefore, within the scope of the present disclosure, the first component in one embodiment may be referred to as the second component in another embodiment, and similarly, the second component in one embodiment may be referred to as the first component in another embodiment.
[0041] In the present disclosure, components distinguished from each other are intended to clearly describe each feature, and do not mean that these components must be separated. That is, multiple components can be integrated and implemented in one hardware or software unit, or one component can be distributed and implemented in multiple hardware or software units. Therefore, such embodiments in which components are integrated or distributed are included in the scope of the present disclosure even if not otherwise stated.
[0042] In the present disclosure, the components described in various embodiments are not necessarily essential components, and some components may be optional components. Therefore, embodiments consisting of a subset of the components described in the embodiments are also included within the scope of the present disclosure. In addition, embodiments including other components in addition to the components described in each embodiment are included within the scope of the present invention.
[0043] The present disclosure relates to the encoding and decoding of images, and unless redefined in the present disclosure, the terms used in the present disclosure may have the general meanings commonly used in the technical field to which the present disclosure pertains.
[0044] The present disclosure presents various embodiments of video / image encoding, and unless otherwise specified, these embodiments can be executed in combination with each other.
[0045] Unless redefined in the present disclosure, the terms used in the present disclosure may have their usual meanings in the technical field to which the present disclosure pertains.
[0046] In the present disclosure, a "picture" generally refers to a unit representing an image for a specific time period, and a slice / tile is a compilation unit that forms part of a picture, and a picture can be composed of one or more slices / tiles. Additionally, a slice / tile can include one or more CTUs (compilation tree units). A picture can be composed of one or more tile groups. A tile group can include one or more tiles. A patch can represent a rectangular area of the CTU rows of the tiles in a picture. In this document, tile groups and slices can be used interchangeably. For example, in this document, a tile group / tile group header can be referred to as a slice / slice header.
[0047] In the present disclosure, a "pixel" or "picture element" can mean the smallest unit that constitutes a picture (or image). Additionally, "sample" can be used as a term corresponding to a pixel. A sample generally can represent a pixel or the value of a pixel, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0048] In the present disclosure, a "unit" can represent a basic unit of image processing. The unit can include at least one of a specific area of a picture and information related to that area. A unit can include one luminance block and two chrominance (e.g., Cb, Cr) blocks. In some cases, the unit can be used interchangeably with terms such as "sample array", "block", or "region". Generally, an M-by-N block can include a sample (or sample array) or set (or array) of transform coefficients of M columns and N rows.
[0049] In the present invention, "current block" may mean one of "current compilation block", "current compilation unit", "compilation target block", "decoding target block", or "processing target block". When performing prediction, "current block" may mean "current prediction block" or "prediction target block". When performing transformation (inverse transformation) / quantization (dequantization), "current block" may represent "current transformation block" or "transformation target block". When performing filtering, "current block" may mean "filtering target block".
[0050] In addition, in the present invention, unless explicitly stated as a chrominance block, "current block" may mean a block including both a luminance component block and a chrominance component block or the "luminance block of the current block". The chrominance component block of the current block can be expressed by an explicit description including a chrominance component block such as "chrominance block" or "current chrominance block".
[0051] In the present disclosure, the terms ",", "and" should be interpreted as indicating "and / or". For example, the expression "For example, table and" and "" can mean "with and / or B with". In addition, "In addition, the expression" and "" and "", B, C. can mean "at least one of A, B, and / or C".
[0052] In the present disclosure, the term "or" should be interpreted as indicating "and / or". For example, the expression "For example, example or B example" can include 1) only "", 2) only "", and / or 3) both "" and "" and B and. In other words, in the present disclosure, the term "or" should be interpreted as indicating "additionally or alternatively".
[0053] Overview of the video compilation system
[0054] Figure 1 is a view schematically showing a video compilation system to which an embodiment of the present disclosure is applicable.
[0055] The video compilation system according to an embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may deliver the encoded video and / or image information or data to the decoding device 20 in the form of a file or a stream via a digital storage medium or a network.
[0056] The encoding device 10 according to an embodiment may include a video source generator 11, an encoder 12, and a transmitter 13. The decoding device 20 according to an embodiment may include a receiver 21, a decoder 22, and a renderer 23. The encoder 12 may be referred to as a video / image encoding device, and the decoder 22 may be referred to as a video / image decoding device. The transmitter 13 may be included in the encoder 12. The receiver 21 may be included in the decoder 22. The renderer 23 may include a display, and the display may be configured as a separate device or an external component.
[0057] The video source generator 11 can obtain video / images through processes such as capturing, synthesizing, or generating video / images. The video source generator 11 can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, and a smart phone, and can (electronically) generate video / images. For example, virtual video / images can be generated by a computer or the like. In this case, the video / image capture process can be replaced by a process of generating relevant data.
[0058] The encoder 12 can encode the input video / image. The encoder 12 can execute a series of programs, such as prediction, transformation, and quantization for compression and compilation efficiency. The encoder 12 is capable of outputting the encoded data (encoded video / image information) in the form of a bitstream.
[0059] The transmitter 13 can send the encoded video / image information or data output in the form of a bitstream to the receiver 21 of the decoding device 20 in the form of a file or a stream through a digital storage medium or a network. The digital storage medium can include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter 13 can include elements for generating a media file in a predetermined file format and can include elements for sending through a broadcast / communication network. The transmitter 13 can be provided as a transmission device separate from the encoder 120. In this case, the transmission device includes: at least one processor that obtains the encoded video / image information or data output in the form of a bitstream; and a transmitter that delivers it in the form of a file or a stream. The receiver 21 can extract / receive the bitstream from the storage medium or the network and send the bitstream to the decoder 22.
[0060] The decoder 22 can decode the video / image by performing a series of processes (such as dequantization, inverse transformation, and prediction corresponding to the operations of the encoder 12).
[0061] The renderer 23 can render the decoded video / image. The rendered video / image can be displayed through a display.
[0062] Overview of the video encoding device
[0063] Figure 2 is a diagram schematically showing an image encoding device to which an embodiment according to the present disclosure can be applied.
[0064] Reference Figure 2, the encoding device 200 includes an image splitter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the image splitter 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or a processor). Additionally, the memory 270 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware component may further include the memory 270 as an internal / external component.
[0065] The image splitter 210 may split an input image (or picture or frame) input to the encoding device 200 into one or more processors. For example, the processor may be referred to as a coding unit (CU). In this case, the coding unit may be recursively split from a coding tree unit (CTU) or a largest coding unit (LCU) according to a quadtree binary tree ternary tree (QTBTTT) structure. For example, a coding unit may be divided into multiple coding units with a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quadtree structure may be applied first, and the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied first. The coding process according to this document may be performed based on the final coding unit that is no longer divided. In this case, the largest coding unit may be used as the final coding unit based on the coding efficiency according to the image characteristics, or, if necessary, the coding unit may be recursively split into coding units with a deeper depth, and the coding unit with the optimal size may be used as the final coding unit. Here, the coding process may include processes such as prediction, transformation, and reconstruction that will be described later. As another example, the processor may further include a predictor (PU) or a transformation unit (TU). In this case, the predictor and the transformation unit may be split or divided from the above-mentioned final coding unit. The predictor may be a unit for sample prediction, and the transformation unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from the transformation coefficients.
[0066] In the encoding device 200, a prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 221 or the intra-frame predictor 222 is subtracted from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, as shown, the unit for subtracting the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) in the encoder 200 may be referred to as the subtractor 231. The predictor may perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a prediction block including the prediction samples of the current block. The predictor may determine whether to apply intra-frame prediction or inter-frame prediction on a per current block or CU basis. As described later in the description of each prediction mode, the predictor may generate various information related to the prediction, such as prediction mode information, and send the generated information to the entropy encoder 240. The information about the prediction may be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0067] The intra-frame predictor 222 may predict the current block by referring to the samples in the current picture. The reference samples may be located in the neighborhood of the current block or may be separately located according to the prediction mode. In intra-frame prediction, the prediction mode may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, the DC mode and the planar mode. According to the level of detail of the prediction direction, the directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used according to the settings. The intra-frame predictor 222 may determine the prediction mode applied to the current block by using the prediction mode applied to the adjacent blocks.
[0068] The inter - frame predictor 221 can derive a predicted block of a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. Here, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, the motion information can be predicted in units of blocks, sub - blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include inter - frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter - frame prediction, adjacent blocks can include spatially adjacent blocks present in the current picture and temporally adjacent blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporally adjacent blocks can be the same or different. Temporally adjacent blocks can be referred to as co - located reference blocks, co - located CUs (colCUs), etc., and the reference picture including the temporally adjacent blocks can be referred to as a co - located picture (colPic). For example, the inter - frame predictor 221 can configure a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter - frame prediction can be performed based on various prediction modes. For example, in the skip mode and the merge mode, the inter - frame predictor 221 can use the motion information of adjacent blocks as the motion information of the current block. In the skip mode, different from the merge mode, the residual signal may not be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of an adjacent block can be used as a motion vector prediction value, and the motion vector of the current block can be indicated by signaling a motion vector difference.
[0069] The predictor 220 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra - frame prediction or inter - frame prediction to predict a block, but also apply both intra - frame prediction and inter - frame prediction simultaneously. This can be referred to as combined inter - frame and intra - frame prediction (CIIP). Additionally, the predictor can be based on the intra - block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or the palette mode can be used for content image / video compilation such as games, e.g., screen content compilation (SCC). IBC basically performs prediction in the current picture, but can be performed similarly to inter - frame prediction because a reference block is derived in the current picture. That is, IBC can use at least one of the inter - frame prediction techniques described in this document. The palette mode can be regarded as an example of intra - frame compilation or intra - frame prediction. When the palette mode is applied, the sample values within the picture can be signaled based on information about the palette table and the palette index.
[0070] The prediction signal generated by a predictor (including the inter-frame predictor 221 and / or the intra-frame predictor 222) can be used to generate a reconstructed signal or to generate a residual signal. The subtraction unit 115 can subtract the prediction signal (prediction block, prediction sample array) output from the predictor 200 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array). The generated residual signal can be sent to the transformation unit 232.
[0071] The transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a karhunen-loève transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, when the relationship information between pixels is represented by a graph, GBT means a transform obtained from the graph. CNT refers to a transform generated based on a prediction signal generated using all previously reconstructed pixels. Additionally, the transform process can be applied to square pixel blocks of the same size, or can be applied to blocks of variable size and non-square shape.
[0072] The quantizer 233 can quantize the transform coefficients and send them to the entropy encoder 240, and the entropy encoder 240 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 233 can rearrange the block-type quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. Information about the transform coefficients can be generated.
[0073] The entropy encoder 240 can perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 240 can encode, together or separately, the information required for video / image reconstruction in addition to the quantized transform coefficients (e.g., the values of syntax elements, etc.). The encoded information (e.g., the encoded video / image information) can be sent or stored in units of the NAL (network abstraction layer) in the form of a bitstream. The video / image information can further include information about various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information can also include general constraint information. In this document, the information and / or syntax elements transmitted / signaled from the encoding device to the decoding device can be included in the video / picture information. The video / image information can be encoded through the above encoding process and included in the bitstream.
[0074] The bitstream can be transmitted over a network or can be stored in a digital storage medium. The network can include a broadcast network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that transmits the signal output from the entropy encoder 240 and / or a storage unit (not shown) that stores the signal can be included as an internal / external element of the encoding device 200, and alternatively, the transmitter can be included in the entropy encoder 240.
[0075] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, the quantized transform coefficients can be dequantized and inverse-transformed by the dequantizer 234 and the inverse-transformer 235 to reconstruct a residual signal (residual block or residual samples).
[0076] The adder 250 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). If there is no residual for the block to be processed, such as in the case of applying the skip mode, the predicted block can be used as the reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture, and can be used for inter-frame prediction of the next picture by filtering as described below.
[0077] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 270 (specifically, the DPB of the memory 270). Various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various information related to filtering, and send the generated information to the entropy encoder 240, as described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoder 240 and output in the form of a bitstream.
[0078] The modified reconstructed picture sent to the memory 270 can be used as a reference picture in the inter-frame predictor 221. When inter-frame prediction is applied by the encoding device, prediction mismatch between the encoding device 200 and the decoding device can be avoided, and the encoding efficiency can be improved.
[0079] The DPB of the memory 270 may store the modified reconstructed picture to be used as a reference picture in the inter - frame predictor 221. The memory 270 may store the motion information of the blocks from which the motion information in the current picture is derived (or encoded) and / or the motion information of the blocks in the already reconstructed pictures. The stored motion information may be sent to the inter - frame predictor 221 and used as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 270 may store the reconstructed samples of the reconstructed blocks in the current picture and may transfer the reconstructed samples to the intra - frame predictor 222.
[0080] Overview of the image decoding device
[0081] Figure 3 is a schematic diagram showing an image decoding device to which embodiments according to the present disclosure may be applied.
[0082] Reference Figure 3 , the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter - frame predictor 331 and an intra - frame predictor 332. The residual processor 320 may include a de - quantizer 321 and an inverse transformer 321. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be configured by hardware components (e.g., a decoder chipset or a processor). Additionally, the memory 360 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0083] When receiving a bitstream including video / image information, the decoding device 300 may reconstruct an image corresponding to the process of processing the video / image information in the Figure 2 encoding device. For example, the decoding device 300 may derive units / blocks based on the block partition - related information obtained from the bitstream. The decoding device 300 may perform decoding using the processor applied in the encoding device. Thus, the decoding processor may be, for example, a compilation unit, and may partition the compilation unit according to the quadtree structure, binary tree structure, and / or ternary tree structure from the coding tree unit or the largest compilation unit. One or more transform units may be derived from the compilation unit. The reconstructed image signal decoded and output by the decoding device 300 may be reproduced by a reproduction device.
[0084] The decoding device 300 is capable of receiving in the form of a bitstream from Figure 2The signal output by the encoding device, and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information can further include information about various parameter sets such as adaptive parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). In addition, the video / image information can also include general constraint information. The decoding device can also decode the picture based on the information about the parameter set and / or the general constraint information. The signaled / received information and / or syntax elements described later in this document can be obtained from the bitstream by decoding using the decoded pictures. For example, the entropy decoder 310 decodes the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and outputs the syntax elements required for image reconstruction and the quantization values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive the bins corresponding to each syntax element in the bitstream, determine the context model using the information of the syntax element to be decoded, the information of the block to be decoded, or the information of the symbols / bins decoded in the previous stage, and perform arithmetic decoding on the bins by predicting the occurrence probability of the bins according to the determined context model, and generate symbols corresponding to the values of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using the information of the decoded symbols / bins of the context model of the next symbol / bin. The information related to the prediction in the information decoded by the entropy decoder 310 can be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values (i.e., the quantized transform coefficients and the related parameter information) obtained by performing entropy decoding in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive a residual signal (residual block, residual sample, residual sample array). In addition, the information about the filtering in the information decoded by the entropy decoder 310 can be provided to the filter 350. Meanwhile, the receiver (not shown) for receiving the signal output by the encoding device can be further configured as an internal / external component of the decoding device 300, or the receiver can be a component of the entropy decoder 310.
[0085] Meanwhile, the decoding device according to this document can be referred to as a video / image / picture decoding device, and the decoding device can be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 310, and the sample decoder can include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0086] The dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scanning order executed in the encoding device. The dequantizer 321 can perform dequantization on the quantized transform coefficients by using quantization parameters (e.g., quantization step information) and obtain the transform coefficients.
[0087] The inverse transformer 322 inverse-transforms the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0088] The predictor 330 can perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor can determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output from the entropy decoder 310, and can determine a specific intra / inter prediction mode.
[0089] The predictor 330 can generate a prediction signal based on various prediction methods (techniques) described below, which are the same as those mentioned in the description of the predictor 220 of the image encoding device 200.
[0090] The intra predictor 332 can predict the current block by referring to samples in the current picture. The reference samples can be located in the neighborhood of the current block, or can be separately located according to the prediction mode. In intra prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra predictor 331 can determine the prediction mode applied to the current block by using the prediction mode applied to an adjacent block.
[0091] The inter predictor 331 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between an adjacent block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information about the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the adjacent block can include a spatially adjacent block present in the current picture and a temporally adjacent block present in the reference picture. For example, the inter predictor 332 can configure a motion information candidate list based on adjacent blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. The inter prediction can be performed based on various prediction modes, and the information about prediction can include information indicating the inter prediction mode of the current block.
[0092] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the predicted signal (predicted block, predicted sample array) output from a predictor (including the inter-frame predictor 332 and / or the intra-frame predictor 331). If there is no residual for the block to be processed, such as when the skip mode is applied, the predicted block can be used as the reconstructed block. The adder 340 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture, can be output through filtering as described below, or can be used for inter-frame prediction of the next picture.
[0093] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 360 (specifically, the DPB of the memory 360). The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0094] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can store the motion information of the blocks from which the motion information in the current picture is derived (or decoded) and / or the motion information of the blocks in the already reconstructed pictures. The stored motion information can be sent to the inter-frame predictor 260 for use as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 360 can store the reconstructed samples of the reconstructed blocks in the current picture and transmit the reconstructed samples to the intra-frame predictor 331.
[0095] In the present disclosure, the embodiments described in the filter 260, the inter-frame predictor 221, and the intra-frame predictor 222 of the encoding device 100 can be the same as or respectively correspond to the filter 350, the inter-frame predictor 332, and the intra-frame predictor 331 of the decoding device 300. This can also be applied to the unit 332 and the intra-frame predictor 331.
[0096] Figure 4 Exemplarily, a hierarchical structure of the compiled video / image to which the embodiments according to the present disclosure can be applied is shown.
[0097] Reference Figure 4 , the compiled image / video is divided into a VCL (compiled coding layer) that processes the image / video and its own decoding process, a subsystem that transmits and stores the compiled information, and a NAL (network abstraction layer) that is responsible for functions and exists between the VCL and the subsystem.
[0098] In VCL, VCL data including compressed image data (slice data) can be generated, or a parameter set including a picture parameter set (PSP), a sequence parameter set (SPS), and a video parameter set (VPS) or supplementary enhancement information (SEI) message that is additionally required for the image decoding process can be generated.
[0099] In NAL, a NAL unit can be generated by adding header information (NAL unit header) to the raw byte sequence payload (RBSP) generated in VCL. In this case, the RBSP refers to slice data, a parameter set, an SEI message, etc. generated in VCL. The NAL unit header can include NAL unit type information specified according to the RBSP data included in the corresponding NAL unit.
[0100] As shown in the figure, NAL units can be classified into VCL NAL units and non-VCL NAL units according to the RBSP generated in VCL. A VCL NAL unit can represent a NAL unit including information about an image (slice data), and a non-VCL NAL unit can represent a NAL unit including information required for decoding an image (parameter set or SEI message).
[0101] The above VCL NAL units and non-VCL NAL units can be transmitted over a network by attaching header information according to the data standard of the subsystem. For example, a NAL unit can be transformed into a data format of a predetermined standard, such as the H.266 / VVC file format, the real-time transport protocol (RTP), the transport stream (TS), etc., and transmitted over various networks.
[0102] As described above, a NAL unit can be specified by a NAL unit type according to the RBSP data structure included in the corresponding NAL unit, and information about the NAL unit type can be stored and signaled in the NAL unit header.
[0103] For example, a NAL unit can be classified into a VCL NAL unit type and a non-VCL NAL unit type according to whether the NAL unit includes information about an image (slice data). The VCL NAL unit type can be classified according to the nature and type of the picture included in the VCL NAL unit, and the non-VCL NAL unit type can be classified according to the type of the parameter set.
[0104] The following are examples of NAL unit types specified according to the type of the parameter set included in the non-VCL NAL unit type.
[0105] APS (Adaptive Parameter Set) NAL unit: The type of the NAL unit including APS
[0106] DPS (Decoding Parameter Set) NAL unit: The type of NAL unit including DPS
[0107] VPS (Video Parameter Set) NAL unit: The type of NAL unit including VPS
[0108] SPS (Sequence Parameter Set) NAL unit: The type of NAL unit including SPS
[0109] PPS (Picture Parameter Set) NAL unit: The type of NAL unit including PPS
[0110] The aforementioned NAL unit types may have syntax information for the NAL unit type, and the syntax information may be stored in and signaled in the NAL unit header. For example, the syntax information may be nal_unit_type, and the NAL unit type may be specified by the nal_unit_type value.
[0111] The slice header (slice header syntax) may include information / parameters that can be commonly applied to the slice. APS (APS syntax) or PPS (PPS syntax) may include information / parameters that can be commonly applied to one or more slices or pictures. SPS (SPS syntax) may include information / parameters that can be commonly applied to one or more sequences. VPS (VPS syntax) may include information / parameters that can be commonly applied to multiple layers. DPS (DPS syntax) may include information / parameters that can generally be applied to the entire video. DPS may include information / parameters related to the concatenation of the compiled video sequence (CVS). The high-level syntax (HLS) in this document may include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, and slice header syntax.
[0112] In this document, the image / video information encoded by the encoding device and signaled to the decoding device in the form of a bitstream includes not only information related to partitions in the picture, intra / inter prediction information, residual information, in-loop filter information, etc., but also information in the slice header, information in the picture header, information in APS, information in PPS, information in SPS, information in VPS, and / or information in DPS.
[0113] General post - processing filtering process using NNPF
[0114] The input to the process is the BitstreamToFilter. The output of the process is a list ListNnpfOutputPics of NNPF output pictures. First, decode the BitstreamToFilter, and set the list CroppedDecodedPictures to the list of cropped decoded pictures in output order produced from decoding the BitstreamToFilter. Second, for each cropped decoded picture in CroppedDecodedPictures for which one or more NNPFs are activated, repeatedly call the filtering process for one picture in output order. The order of the pictures in ListNnpfOutputPics is the output order.
[0115] Within ListNnpfOutputPics, there should be no more than one picture related to any particular output time instance. When for any particular picture in CroppedDecodedPictures there are multiple NNPFs that are activated and although any NNPF can be chosen, only one NNPF is allowed to be chosen for application, the above constraint will apply regardless of which NNPF is chosen to be applied to the particular picture.
[0116] The specified filtering process applies to each cropped decoded picture (referred to as the current picture) in CroppedDecodedPictures for which one or more NNPFs are activated. When applying an NNPF to the current picture, the NNPF generates a filtered and / or interpolated picture by applying the NNPF process specified in the semantics of the NNPF C SEI message to the current picture in a patch-by-patch manner.
[0117] When applying an NNPF to the current picture, the order of the pictures generated by the NNPF by applying the NNPF process stored in the output tensor of the NNPF is the output order. When the applied NNPF is the last NNPF applied to the current picture, the pictures generated by the NNPF and output by the NNPF process are included in ListNnpfOutputPics in the same order as when the pictures were stored in the output tensor of the NNPF.
[0118] Neural network post - filter characteristics SEI message (NNPFC)
[0119] The combinations in Tables 1 to 3 represent the NNPF C syntax structure.
[0120] [Table 1]
[0121]
[0122] [Table 2]
[0123]
[0124] [Table 3]
[0125]
[0126] The NNPFC syntax structures of Tables 1 to 3 can be signaled in the form of Supplemental Enhancement Information (SEI) messages. The SEI messages signaling the NNPFC syntax structures of Tables 1 to 3 can be referred to as NNPFC SEI messages.
[0127] The Neural Network Post-Filter Characteristics (NNPFC) SEI message specifies the neural network that can be used as a post-processing filter. The Neural Network Post-Filter Activation (NNPFA) SEI message is used to indicate the use of the specified neural network post-processing filter (NNPF) for a particular picture.
[0128] The use of this SEI message requires the definition of the following variables:
[0129] The width and height of the input picture in terms of luma samples, which are represented by CroppedWidth and CroppedHeight respectively in this document.
[0130] The luma sample array CroppedYPic[idx] and the chroma sample arrays CroppedCbPic[idx] and CroppedCrPic[idx] (when present) of the input picture with an index idx in the range of 0 to numInputPics - 1 (including 0 and numInputPics - 1), which are used as inputs for the NNPF.
[0131] The bit depth BitDepthY of the luma sample array for the input picture.
[0132] The bit depth BitDepthC of the chroma sample array (if any) for the input picture.
[0133] The chroma format indicator, which is represented by ChromaFormatIdc in this document.
[0134] When nnpfc_auxiliary_inp_idc is equal to 1, the filter strength control value array StrengthControlVal[idx] shall contain real numbers in the range of 0 to 1 (including 0 and 1) for the input picture with an index idx in the range of 0 to numInputPics - 1 (including 0 and numInputPics - 1).
[0135] The input picture with index 0 corresponds to the NNPF defined by the NNPFC SEI message and is the picture activated by the NNPFA SEI message. The input picture with index i in the range of 1 to numInputPics - 1 (inclusive of 1 and numInputPics - 1) is before the input picture with index i - 1 in the output order.
[0136] The variables SubWidthC and SubHeightC are derived from ChromaFormatIdc. For the same picture, there can be more than one NNPFC SEI message. When there are or are activated more than one NNPFC SEI messages with different nnpfc_id values for the same picture, they can have the same or different nnpfc_purpose and nnpfc_mode_idc values.
[0137] nnpfc_purpose indicates the purpose of the NNPF as specified in Table 3, where (nnpfc_purpose & bitMask) not equal to 0 indicates that the NNPF has the purpose associated with the bitMask value in Table 3. When nnpfc_purpose is greater than 0 and (nnpfc_purpose & bitMask) equals 0, the purpose associated with the bitMask value does not apply to the NNPF. When nnpfc_pupose equals 0, the NNPF determined by the application can be used.
[0138] In the bitstream compliant with this version of this document, the value of nnpfc_purpose shall be in the range of 0 to 63 (inclusive of 0 and 63). The values from 64 to 65535 (inclusive of 64 and 65535) for nnpfc_purpose are reserved for future use by ITU - T|ISO / IEC and shall not be present in the bitstream compliant with this version of this document. The decoder compliant with this version of this document shall ignore the NNPFC SEI messages with nnpfc_purpose in the range of 64 to 65535 (inclusive of 64 and 65535).
[0139] [Table 4]
[0140]
[0141] The following derived variables chromaUpsamplingFlag, resolutionResamplingFlag, pictoureRateUpsamplingFlag, bitDepthUpsamplingFlag, and coulouzationFlag specify whether nnpfc_purpose indicates the purpose of the NNPF to include chroma upsampling, resolution resampling, picture rate upsampling, bit depth upsampling, and CoulourizationFlag, respectively:
[0142] [Table 5]
[0143]
[0144] When future reserved values of nnpfc_purpose are used, the syntax of this SEI message can be extended with syntax elements that are conditional on nnpfc_purpose being equal to the said value.
[0145] When ChromaFormatIdc is equal to 3, chromaUpsamplingFlag shall be equal to 0.
[0146] When ChromaFormatIdc or chromaUpsamplingFlag is not equal to 0, colourizationFlag shall be equal to 0.
[0147] When pictureRateUpsamplingFlag is equal to 1 and the input picture with index 0 is associated with a frame-packing arrangement SEI message with fp_arrangement_type equal to 5, all input pictures are associated with a frame-packing arrangement SEI message with the same value of fp_arrangement_type and fp_current_frame_is_frame0_flag equal to 5.
[0148] nnpfc_id contains an identification number that can be used to identify the NNPF. The value of nnpfc_id shall be in the range of 0 to 2 32 -2 (including 0 and 2 32 -2). Values of nnpfc_id from 256 to 511 (including 256 and 511) and from 2 31 to 2 32 -2 (including 2 31 and 2 32 -2) are reserved for future use by ITU-T|ISO / IEC. Encountering an nnfc_id in the range of 256 to 511 (including 256 and 511) or in the range of 2 31 to 232 -2 (including 2 31 and 2 32 -2), decoders compliant with this version of this document for NNPFC SEI messages within this range shall ignore the said SEI messages.
[0149] When the NNPFC SEI message is the first NNPFC SEI message with a specific nnpfc_id value within the current CLVS in decoding order, the following applies:
[0150] - This SEI message specifies the base NNPF.
[0151] - This SEI message is related to the current decoded picture and all subsequent decoded pictures of the current layer in output order until the end of the current CLVS.
[0152] nnpfc_base_flag being equal to 1 specifies that the SEI message specifies the base NNPF. nnpfc_base_flag being equal to 1 specifies that the SEI message specifies an update related to the base NNPF.
[0153] The following constraints apply to the value of nnpfc_base_flag:
[0154] - When the NNPFC SEI message is the first NNPFC SEI message with a specific nnpfc_id value within the current CLVS in decoding order, the value of nnpfc_base_flag shall be equal to 1.
[0155] - When the NNPFC SEI message nnpfcB is not the first NNPFC SEI message with a specific nnpfc_id value within the current CLVS in decoding order and the value of nnpfc_base_flag is equal to 1, the NNPFC SEI message shall be a repetition of the first NNPFC SEI message nnnpfcA with the same nnpfc_id value in decoding order, i.e., the payload content of nnpfcB shall be the same as the payload content of nnpfcA.
[0156] When nnpfc_base_flag is equal to 0, the following applies:
[0157] - This SEI message defines an update relative to a previous base NNPF in decoding order having the same nnpfc_id value. The updates are not cumulative, but each update is applied to the base NNPF which is the NNPF specified by the first NNPFC SEI message in decoding order having a specific npfc_id value within the current CLVS. The NNPF defined by this SEI message is obtained by applying the update defined by this SEI message relative to the base NNPF having the same nnpfc_id value.
[0158] - This SEI message pertains to the current decoded picture and all subsequent decoded pictures (in output order) of the current layer until the end of the current CLVS or until but not including the decoded picture that follows the current decoded picture in output order within the current CLVS and is associated with a subsequent NNPFC SEI message (in decoding order) having an nnpfc_base_flag equal to 0 and the said specific nnpfc_id value within the current CLVS (whichever is earlier).
[0159] An nnpfc_mode_idc equal to 0 indicates that this SEI message contains an ISO / IEC 15938-17 bitstream that specifies a base NNPF (when nnpfc_base_flag is equal to 1) or an update relative to a base NNPF having the same nnpfc_id value (when nnpfc_base_flag is equal to 0).
[0160] When nnpfc_base_flag is equal to 1, an nnpfc_mode_idc equal to 1 specifies that the base NNPF associated with the nnpfc_id value is the neural network identified by the URI indicated by nnpfc_uri having a format identified by the tag URI nnpfc_tag_uri. When nnpfc_base_flag is equal to 0, an nnpfc_mode_idc equal to 1 specifies that the update relative to the base NNPF having the same nnpfc_id value is defined by the URI indicated by nnpfc_uri, which has a format identified by the tag URI nnpfc_tag_uri.
[0161] In the bitstream conforming to this version of this document, the value of nnpfc_mode_idc shall be in the range of 0 to 1 (including the endpoints). Values of nnpfc_mode_idc from 2 to 255 (including the endpoints) are reserved for future use and shall not be present in the bitstream conforming to this version of this document. The decoder conforming to this version of this document shall ignore NNPFC SEI messages having nnpfc_mode_idc in the range of 2 to 255 (including 2 and 255). Values of nnpfc_mode_idc greater than 255 shall not be present in the bitstream conforming to this version of this document and are not reserved for future use.
[0162] In the bitstream conforming to this version of this document, nnpfc_reserved_zero_bit_a shall be equal to 0. The decoder shall ignore NNPFC SEI messages in which nnpfc_reserved_zero_bit_a is not equal to 0.
[0163] nnpfc_tag_uri includes a tag URI having the syntax and semantics as specified in IETF RFC4151, which identifies the format and associated information of the neural network used as the base NNPF or an update with respect to the base NNPF having the same nnpfc_id value specified by nnpfc_uri.
[0164] nnpfc_tag_uri enables the unique identification of the format of the neural network data specified by nnrpf_uri without the need for a central registration authority.
[0165] nnpfc_tag_uri being equal to "g:iso.org,2023:15938-17" indicates that the neural network data identified by nnpfc_uri conforms to ISO / IEC 15938-17.
[0166] nnpfc_uri includes a URI having the syntax and semantics as specified in IETF Internet Standard 66, which identifies the neural network used as the base NNPF or an update with respect to the base NNPF having the same nnpfc_id value.
[0167] nnpfc_property_present_flag being equal to 1 specifies that the syntax elements related to filter purpose, input formatting, output formatting, and complexity are present. nnpfc_property_present_flag being equal to 0 specifies that the syntax elements related to filter purpose, input formatting, output formatting, and complexity are not present.
[0168] When nnpfc_base_flag is equal to 1, nnpfc_property_present_flag shall be equal to 1.
[0169] When nnpfc_property_present_flag is equal to 0, the values of all syntax elements that can be present only when nnpfc_property_present_flag is equal to 1 are inferred to be equal to their corresponding syntax elements in the NNPFC SEI message that includes the base NNPF for which this SEI message provides an update, respectively.
[0170] When the NNPFC SEI message nnpfcCurr is not the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS, is not a repeat of the first NNPFC SEI message with that particular npfc_id (i.e., the value of nnpfc_base_flag is equal to 0), and the value of nnpfc_property_present_flag is equal to 1, the following constraints apply:
[0171] - The value of nnpfc_purpose in the NNPFC SEI message shall be the same as the value of nnpfc_purpose in the first NNPFC SEI message in decoding order that has the particular nnpfc_id value within the current CLVS.
[0172] - The values of the syntax elements in the NNPFC SEI message that are after nnpfc_property_present_flag and before nnpfc_complexity_info_present_flag in decoding order shall be the same as the values of the corresponding syntax elements in the first NNPFC SEI message in decoding order that has the particular npfc_id value within the current CLVS.
[0173] - In the first NNPFC SEI message, nnpfc_complexity_info_present_flag shall be equal to 0 or both nnpfc_complexity_info_present_flag shall be equal to 1, in decoding order, the first NNPFC SEI message has the particular nnpfc_id value within the current CLVS (hereinafter referred to as nnpfcBase) and all of the following apply:
[0174] - The nnpfc_parameter_type_idc in npfcCurr shall be equal to the nnpfc_parameter_type_idc in nnpfcBase.
[0175] - When present, nnpfc_log2_parameter_bit_length_minus3 in npfcCurr should be less than or equal to nnpfc_log2_parameter_bit_length_minus3 in nnpfcBase.
[0176] - If nnpfc_num_parameters_idc in nnpfcBase is equal to 0, then nnpfc_num_parameters_idc in npfcCurr should be equal to 0.
[0177] - Otherwise (nnpfc_num_parameters_idc in nnpfcBase is greater than 0), nnpfc_num_parameters_idc in npfcCurr should be greater than 0 and less than or equal to nnpfc_num_parameters_idc in nnpfcBase.
[0178] - If nnpfc_num_kmac_operations_idc in nnpfcBase is equal to 0, then nnpfc_num_kmac_operations_idc should be equal to 0.
[0179] - Otherwise (nnpfc_num_kmac_operations_idc in nnpfcBase is greater than 0), nnpfc_num_kmac_operations_idc in npfcCurr should be greater than 0 and less than or equal to nnpfc_num_kmac_operations_idc in nnpfcBase.
[0180] - If nnpfc_total_kilobyte_size in nnpfcBase is equal to 0, then nnpfc_total_kilobyte_size in npfcCurr should be equal to 0.
[0181] - Otherwise (nnpfc_total_kilobyte_size in nnpfcBase is greater than 0), nnpfc_total_kilobyte_size in npfcCurr should be greater than 0 and less than or equal to nnpfc_total_kilobyte_size in nnpfcBase.
[0182] nnpfc_num_input_pics_minus1 plus 1 specifies the number of pictures used as input to the NNPF. The value of nnpfc_num_input_pics_minus1 shall be in the range of 0 to 63, inclusive. When pictureRateUpsamplingFlag is equal to 1, the value of nnpfc_num_input_pics_minus1 shall be greater than 0.
[0183] The variable numInputPics, which specifies the number of pictures used as input to the NNPF, is derived as follows:
[0184] [Equation 1]
[0185]
[0186] nnpfc_input_pic_output_flag[i] being equal to 1 indicates that for the i-th input picture, the NNPF generates the corresponding output picture. nnpfc_input_pic_output_flag[i] being equal to 0 means that for the i-th input picture, the NNPF does not generate the corresponding output picture. When nnpfc_num_input_pics_minus1 is equal to 0, it is inferred that nnpfc_input_pic_output_flag[0] is equal to 1. When pictureRateUpsamplingFlag is equal to 0 and nnpfc_num_input_pics_minus1 is greater than 0, for at least one value of i in the range of 0 to nnpfc_num_input_pics_minus1, inclusive, nnpfc_input_pic_output_flag[i] shall be equal to 1.
[0187] The nnpfc_absent_input_pic_zero_flag being equal to 1 indicates that the input picture not present in the NNPF is represented by a sample array with sample values equal to 0. The nnpfc_absent_input_pic_flag being equal to 0 indicates that the input picture not present in the NNPF is represented by the input picture nearest in output order within the bitstream. When chromaUpsamplingFlag is equal to 1, the nnpfc_out_sub_c_flag specifies the values of the variables outSubWidthC and outSubHeightC. The nnpfc_out_sub_c_flag being equal to 1 specifies that outSubWidthC is equal to 1 and outSubHeightC is equal to 1. The nnpfc_out_sub_c_flag being equal to 0 specifies that outSubWidthC is equal to 2 and outSubHeightC is equal to 1. When ChromaFormatIdc is equal to 2 and the nnpfc_out_sub_c_flag exists, the value of the nnpfc_out_sub_c_flag shall be equal to 1.
[0188] When colourizationFlag is equal to 1, the nnpfc_out_colour_format_idc specifies the colour format of the NNPF output and thereby specifies the values of the variables outSubWidthC and outSubHeightC. The nnpfc_out_colour_format_idc being equal to 1 specifies that the colour format of the NNPF output is the 4:2:0 format and both outSubWidthC and outSubHeightC are equal to 2. The nnpfc_out_colour_format_idc being equal to 2 specifies that the colour format of the NNPF output is the 4:2:2 format and outSubWidthC is equal to 2 and outSubHeightC is equal to 1. The nnpfc_out_colour_format_idc being equal to 3 indicates that the colour format output of the NNPF is the 4:4:4 format and both outSubWidthC and outSubHeightC are equal to 1. The value of the nnpfc_out_colour_format_idc shall not be equal to 0.
[0189] When both chromaUpsamplingFlag and colourizationFlag are equal to 0, it is inferred that outSubWidthC and outSubHeightC are equal to SubWidthC and SubHeightC respectively. nnpfc_pic_width_num_minus1 plus 1 and nnpfc_pic_width_denom_minus1 plus 1 specify the numerator and denominator of the resampling rate of the NNPF output picture width relative to CroppedWidth respectively. The value of (nnpfc_pic_width_num_minusl + 1) ÷ (nnpfc_pic_width_denom_minusl + 1) should be in the range of l÷16 to l16 (including l÷16 and l16). When nnpfc_pic_width_num_minus1 and nnpfc_pic_width_denom_minus1 do not exist, the values of nnpfc_pic_width_num_minus1 and nnpfc_pic_width_denom_minus1 are both inferred to be equal to 0.
[0190] The variable nnpfcOutputPicWidth, representing the width of the luminance sample array of the picture produced by applying the NNPF identified by nnpfc_id to the input picture, is derived as follows:
[0191] [Equation 2]
[0192]
[0193] The requirement for bitstream consistency is that the value of nnpfcOutputPicHeight % outSubHeightC should be equal to 0.
[0194] nnpfc_pic_height_num_minus1 plus 1 and nnpfc_pic_height_denom_minus1 plus 1 respectively specify the numerator and denominator of the resampling rate of the NNPF output picture height relative to CroppedHeight. The value of (nnpfc_pic_height_num_minusl + 1) ÷ (nnpfc_pic_height_denom_minusl + 1) should be in the range of 1÷16 to 116 (including 1÷16 and 116). When nnpfc_pic_height_num_minus1 and nnpfc_pic_height_denom_minus1 do not exist, the values of nnpfc_pic_height_num_minus1 and nnpfc_pic_height_denom_minus1 are both inferred to be equal to 0.
[0195] The variable nnpfcOutputPicHeight, representing the height of the luminance sample array of the picture generated by applying the NNPF identified by nnpfc_id to the input picture, is derived as follows:
[0196] [Equation 3]
[0197]
[0198] The requirement for bitstream consistency is that the value of nnpfcOutputPicHeight % outSubHeightC should be equal to 0.
[0199] When nnpfc_pic_width_num_minus1, nnpfc_pic_width_denom_minus1, nnpfc_pic_height_num_minus1, and nnpfc_pic_height_denom_minus1 exist, at least one of the following should be true:
[0200] - The value of nnpfcOutputPicWidth is not equal to CroppedWidth.
[0201] - The value of nnpfcOutputPicHeight is not equal to CroppedHeight.
[0202] nnpfc_interpolated_pics[i] specifies the number of interpolated pictures generated by the NNPF between the i-th picture and the (i + 1)-th picture used as input to the NNPF. The value of nnpfc_interpolated_pics[i] shall be in the range of 0 to 63, inclusive of 0 and 63.
[0203] For at least one value of i in the range of 0 to nnpfc_num_input_pics_minus-1-1, inclusive of 0 and nnpfc_num_input_pics_minus-1-1, the value of nnpfc_interpolated_pics[i] shall be greater than 0.
[0204] The variable NumInpPicsInOutputTensor specifies the number of pictures with corresponding input pictures that are present in the output tensor of the NNPF, InpIdx[idx] specifies the input picture index of the idx-th picture that is present in the output tensor of the NNPF and has a corresponding input picture, and numOutputPics specifies the total number of pictures present in the output tensor of the NNPF, and are derived as follows:
[0205] [Table 6]
[0206]
[0207] nnpfc_component_last_flag being equal to 1 indicates that the last dimension in the input tensor inputTensor to the NNPF and the output tensor outputTensor generated from the NNPF is used for the current channel. nnpfc_component_last_flag being equal to 0 indicates that the third dimension in the input tensor inputTensor to the NNPF and the output outputTensor generated from the NNPF is used for the current channel.
[0208] The first dimension in the input tensor and the output tensor is used for the batch index, which is a practice in some neural network frameworks. Although the formulas in the semantics of this SEI message use the batch size corresponding to a batch index equal to 0, it depends on the post-processing implementation to determine the batch size used as input to the neural network inference.
[0209] For example, when nnpfc_inp_order_idc is equal to 3 and nnpfc_auxiliary_inp_idc is equal to 1, there are 7 channels in the input tensor, including four luminance matrices, two chrominance matrices, and one auxiliary input matrix. In this case, the process DerivedInputTensors() will derive each of these 7 channels of the input tensor one by one, and when processing a specific channel among these channels, this channel is called the current channel during this process.
[0210] nnpfc_inp_format_idc indicates the method of converting the sample values of the input picture into the input values for the NNPF. When nnpfc_inp_format_idc is equal to 0, the input values for the NNPF are real numbers, and the functions InpY() and InpC() are specified as follows:
[0211] [Equation 4]
[0212]
[0213] When nnpfc_inp_format_idc is equal to 1, the input values for the NNPF are unsigned integers, and the functions InpY() and InpC() are specified as follows:
[0214] [Table 7]
[0215]
[0216] The variable inpTensorBitDepthY is derived from the syntax element nnpfc_inp_tensor_luma_bitdepth_minus8 specified as follows. The variable inpTensorBitDepthC is derived from the syntax element nnpfc_inp_tensor_chroma_bitdepth_minus8 specified as follows.
[0217] Values of nnpfc_inp_format_idc greater than 1 are reserved for future specifications of ITU-T|ISO / IEC and shall not be present in the bitstream conforming to this version of this document. A decoder conforming to this version of this document shall ignore the NNPFC SEI message containing the reserved value of nnpfc_inp_format_idc.
[0218] An nnpfc_auxiliary_inp_idc greater than 0 indicates that auxiliary input data exists in the input tensor of the NNPF, an nnpfc_auxiliary_inp_idc equal to 0 indicates that auxiliary input data does not exist in the input tensor, and an nnpfc_auxiliary_inp_idc equal to 1 specifies that the auxiliary input data is derived as specified in Equation 85.
[0219] In the bitstream conforming to this version of this document, the value of nnpfc_auxiliary_inp_idc shall be in the range of 0 to 1 (including 0 and 1). The values of nnpfc_auxiliary_inp_idc from 2 to 255 (including 2 and 255) are reserved for future use and shall not exist in the bitstream conforming to this version of this document. The decoder conforming to this version of this document shall ignore NNPFC SEI messages with an nnfc_auxiliary_inp_idc in the range of 2 to 255 (including 2 and 255). Values of nnpfc_auxiliary_inp_idc greater than 255 shall not exist in the bitstream conforming to this version of this document and are not reserved for future use.
[0220] The nnpfc_inp_order_idc indicates the method of sorting the sample array of the input picture to form the input tensor of the NNPF.
[0221] In the bitstream conforming to this version of this document, the value of nnpfc_inp_order_idc shall be in the range of 0 to 3 (including 0 and 3). The values of nnpfc_inp_order_idc from 4 to 255 (including 4 and 255) are reserved for future use by ITU-T|ISO / IEC and shall not exist in the bitstream conforming to this version of this document. The decoder conforming to this version of this document shall ignore NNPFC SEI messages with an nnfc_inp_order_idc in the range of 4 to 255 (including 4 and 255). Values of nnpfc_inp_order_idc greater than 255 shall not exist in the bitstream conforming to this version of this document and are not reserved for future use.
[0222] When ChromaFormatIdc is not equal to 1, nnpfc_inp_order_idc shall not be equal to 3.
[0223] When ChromaFormatIdc is equal to 0, nnpfc_inp_order_idc shall be equal to 0.
[0224] When chromaUpsamplingFlag is equal to 1, nnpfc_inp_order_idc shall not be equal to 0.
[0225] Table 8 contains the information description of the nnpfc_inp_order_idc value.
[0226] [Table 8]
[0227]
[0228] nnpfc_inp_tensor_luma_bitdepth_minus8 plus 8 specifies the bit depth of the luma sample values in the input integer tensor. The value of inpTensorBitDepthY is derived as follows:
[0229] [Equation 5]
[0230]
[0231] The requirement for bitstream consistency is that the value of nnpfc_inp_tensor_luma_bitdepth_minus8 shall be in the range of 0 to 24, inclusive (0 and 24).
[0232] nnpfc_inp_tensor_chroma_bitdepth_minus8 plus 8 specifies the bit depth of the chroma sample values in the input integer tensor. The value of inpTensorBitDepthC is derived as follows:
[0233] [Equation 6]
[0234]
[0235] The requirement for bitstream consistency is that the value of nnpfc_inp_tensor_chroma_bitdepth_minus8 shall be in the range of 0 to 24, inclusive (0 and 24).
[0236] When nnpfc_auxiliary_inp_idc is equal to 1, the variable strengthControlScaledVal is derived as follows:
[0237] [Table 9]
[0238]
[0239] A patch is a rectangular array of samples from a component (e.g., luma or chroma component) of a picture.
[0240] The process DerivedInputTensors() for deriving the input tensor inputTensor for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft is specified as follows. The horizontal sample coordinate cLeft specifies the upper left sample position of the patch of samples included in the input tensor:
[0241] [Table 10]
[0242]
[0243] [Table 11]
[0244]
[0245] [Table 12]
[0246]
[0247] nnpfc_out_format_idc equal to 0 indicates that the sample values output by the NNPF are real numbers, where the value range from 0 to 1 (including 0 and 1) is linearly mapped to the unsigned integer value range from 0 to (1<<bitDepth)-1 (including 0 and (1<<bitDepth)-1) for subsequent post-processing or display.
[0248] nnpfc_out_format_idc equal to 1 indicates that the luminance sample values output by the NNPF are unsigned integers in the range from 0 to (1<<outTensorBitDepthY)-1 (including 0 and (1<<outTensorBitDepthY)-1), and the chrominance sample values output by the NNPF are unsigned integers in the range from 0 to (1<<outTensorBitDepthC)-1 (including 0 and (1<<outTensorBitDepthC)-1).
[0249] Values of nnpfc_out_format_idc greater than 1 are reserved for future specifications of ITU-T|ISO / IEC and shall not be present in the bitstream compliant with this version of this document. Decoders compliant with this version of this document shall ignore NNPFC SEI messages containing reserved values of nnpfc_out_format_idc.
[0250] nnpfc_out_order_idc indicates the output order of the samples produced by the NNPF.
[0251] In the bitstream conforming to this version of this document, the value of nnpfc_out_order_idc shall be in the range of 0 to 3 (including 0 and 3). The values of nnpfc_out_order_idc from 4 to 255 (including 4 and 255) are reserved for future use by ITU-T|ISO / IEC and shall not be present in the bitstream conforming to this version of this document. The decoder conforming to this version of this document shall ignore NNPFC SEI messages with nnfc_out_order_idc in the range of 4 to 255 (including 4 and 255). Values of nnpfc_out_order_idc greater than 255 shall not be present in the bitstream conforming to this version of this document and are not reserved for future use.
[0252] When chromaUpsamplingFlag is equal to 1, nnpfc_out_order_idc shall not be equal to 0 or 3.
[0253] When colourizationFlag is equal to 1, nnpfc_out_order_idc shall not be equal to 0.
[0254] Table 13 contains the information description of the values of nnpfc_out_order_idc.
[0255] [Table 13]
[0256]
[0257] nnpfc_out_tensor_luma_bitdepth_minus8 plus 8 specifies the bit depth of the luma sample values in the output integer tensor. The value of nnpfc_out_tensor_luma_bitdepth_minus8 shall be in the range of 0 to 24 (including 0 and 24). The value of outTensorBitDepthY is derived as follows:
[0258] [Equation 7]
[0259]
[0260] nnpfc_out_tensor_chroma_bitdepth_minus8 plus 8 specifies the bit depth of the chroma sample values in the output integer tensor. The value of nnpfc_out_tensor_chroma_bitdepth_minus8 shall be in the range of 0 to 24 (including 0 and 24). The value of outTensorBitDepthC is derived as follows:
[0261] [Equation 8]
[0262]
[0263] When bitDepthUpsamplingFlag is equal to 1, the value of nnpfc_out_format_idc will be equal to 1 and at least one of the following conditions will be true:
[0264] There exists nnpfc_out_tensor_luma_bitdepth_minus8, and outTensorBitDepthY is greater than BitDepthY.
[0265] - nnpfc_out_tensor_chroma_bitdepth_minus8 exists, and outTensorBitDepthC is greater than BitDepthC.
[0266] When there exist nnpfc_inp_tensor_luma_bitdepth_minus8, nnpfc_inp_tensor_chroma_bitdepth_minus8, nnpfc_out_tensor_luma_bitdepth_minus8 and nnpfc_out_tensor_chroma_bitdepth_minus8 and outTensorBitDepthY is greater than inpTensorBitDepthY, outTensorBitDepthC will not be less than inpTensorBitDepthC.
[0267] When there exist nnpfc_inp_tensor_luma_bitdepth_minus8, nnpfc_inp_tensor_chroma_bitdepth_minus8, nnpfc_out_tensor_luma_bitdepth_minus8 and nnpfc_out_tensor_chroma_bitdepth_minus8 and outTensorBitDepthC is greater than inpTensorBitDepthC, outTensorBitDepthY will not be less than inpTensorBitDepthY.
[0268] The StoreOutputTensors() process derives the sample values in the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor outputTensor for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft of the top-left sample position of the patch for the samples specified to be included in the input tensor, as specified below:
[0269] [Table 14]
[0270]
[0271] [Table 15]
[0272]
[0273] The nnpfc_separate_colour_description_present_flag being equal to 1 indicates that different combinations of primary colours, transfer characteristics, matrix coefficients, and scaling and offset values related to the matrix coefficients of the picture generated by NNPF are specified in the SEI message syntax structure. The nnpfc_separate_colour_description_present_flag being equal to 0 indicates that the combination of primary colours, transfer characteristics, matrix coefficients, and scaling and offset values related to the matrix coefficients of the picture generated by NNPF is the same as that indicated in the VUI parameters of CLVS.
[0274] nnpfc_colour_primaries has the same semantics as specified for the vui_colour_primaries syntax element in Subclause 7.3, with the following differences:
[0275] nnpfc_colour_primaries specifies the colour primaries of the picture produced by the NNPF specified in the applied SEI message, rather than the colour primaries for CLVS.
[0276] - When nnpfc_colour_primaries is not present in the NNPFC SEI message, it is inferred that the value of nnpfc_colour_primaries is equal to vui_colour_primaries.
[0277] nnpfc_transfer_characteristics has the same semantics as specified for the vui_transfer_characteristics syntax element in Subclause 7.3, with the following differences:
[0278] - nnpfc_transfer_characteristics specifies the transfer characteristics of pictures generated by the NNPF specified in the applied SEI message, rather than the transfer characteristics for CLVS.
[0279] - When nnpfc_transfer_characteristics does not exist in the NNPF SEI message, it is inferred that the value of nnpfc_transfer_characteristics is equal to vui_transfer_characteristics.
[0280] nnpfc_matrix_coeffs describes the equations used to derive the luma and chroma signals from the green, blue, and red or Y, Z, and X primary colors. Its semantics apply to the pictures generated by the NNPF specified in this applied SEI message and are the same as those specified for MatrixCoefficients in Rec. ITU-T H.273|ISO / IEC 23091-2, where BitDepthY and BitDepthC are equal to outTensorBitDepthY and outTensorBitDepthC, respectively.
[0281] When nnpfc_matrix_coeffs does not exist in the NNPF SEI message, it is inferred that the value of nnpfc_matrix_coeffs is equal to vui_matrix_coeffs.
[0282] nnpfc_matrix_coeffs shall not be equal to 0 unless both of the following conditions are true:
[0283] - nnpfc_out_tensor_chroma_bitdepth_minus8 is equal to nnpfc_out_tensor_luma_bitdepth_minus8.
[0284] - nnpfc_out_order_idc is equal to 2, outSubHeightC is equal to 1, and outSubWidthC is equal to 1.
[0285] nnpfc_matrix_coeffs shall not be equal to 8 unless one of the following conditions is true:
[0286] - nnpfc_out_tensor_chroma_bitdepth_minus8 is equal to nnpfc_out_tensor_luma_bitdepth_minus8.
[0287] - npfc_out_tensor_chroma_bitdepth_minus8 is equal to nnpfc_out_tensor_luma_bitdepth_minus8 + 1, nnpfc_out_order_idc is equal to 2, outSubHeightC is equal to 1, and outSubWidthC is equal to 1.
[0288] The nnpfc_full_range_flag indicates the scaling and offset values applied in association with the matrix coefficients specified by nnpfc_matrix_coeffs. Its semantics are the same as those specified for the VideoFullRangeFlag parameter in Rec. ITU-T H.273|ISO / IEC 23091-2. When absent, it is inferred that the value of nnpfc_full_range_flag is equal to 0.
[0289] The npfc_chroma_loc_info_present_flag being equal to 1 indicates the presence of the nnpfc_chroma_sample_loc_type_frame syntax element in the NNPFC SEI message, and the npfc_chroma_loc_info_present_flag being equal to 0 indicates the absence of the nnpfc_chroma_sample_loc_type_frame syntax element in the NNPFC SEI message. When colourizationFlag is equal to 0 or nnpfc_out_colour_format_idc is not equal to 1, the value of nnpfc_chroma_loc_info_present_flag shall be equal to 0.
[0290] The nnpfc_chroma_sample_loc_type_frame specifies the location of the chroma sampling of the output picture when it is not equal to 6 and nnpfc_out_colour_format_idc is equal to 1, as Figure 1 shown. The nnpfc_chroma_sample_loc_type_frame being equal to 6 and nnpfc_out_colour_format_idc being equal to 1 indicates that the location of the chroma sampling is unknown or not specified, or is specified by some other means not specified in this document. The value of nnpfc_chroma_sample_loc_type_frame shall be in the range of 0 to 6 (including 0 and 6).
[0291] The nnpfc_overlap indicates the horizontal and vertical sample counts of the overlap of adjacent input tensors of the NNPF. The value of nnpfc_overlap should be in the range of 0 to 16383 (including 0 and 16383). The nnpfc_constant_patch_size_flag being equal to 1 indicates that the NNPF exactly accepts the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1 as input. The nnpfc_constant_patch_size_flag being equal to 0 indicates that the NNPF accepts any patch size with width inpPatchWidth and height inpPatchHeight, such that the width of the extended patch (i.e., the patch plus the overlap region) which is equal to inpPatchWidth + 2 * nnpfc_overlap is a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1+2*npfc_overlap, and the height of the extended patch which is equal to inpPatchHeight + 2 * nnpfc_overlap is a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1+2*npfc_overlap.
[0292] When the nnpfc_constant_patch_size_flag is equal to 1, nnpfc_patch_width_minus1 plus 1 indicates the horizontal sample count of the patch size required for the input of the NNPF. The value of nnpfc_patch_width_minus1 should be in the range of 0 to Min(32766, CroppedWidth - 1) (including 0 and Min(32766, CroppedWidth - 1)).
[0293] When the nnpfc_constant_patch_size_flag is equal to 1, nnpfc_patch_height_minus1 plus 1 indicates the vertical sample count of the patch size required for the input of the NNPF. The value of nnpfc_patch_height_minus1 should be in the range of 0 to Min(32766, CroppedHeight - 1) (including 0 and Min(32766, CroppedHeight - 1)).
[0294] When nnpfc_constant_patch_size_flag is equal to 0, nnpfc_extended_patch_width_cd_delta_minus1 plus 1 plus 2 * npfc_overlap indicates the greatest common divisor of all allowed values of the width of the extended patch required for the input to the NNPF. The value of nnpfc_extended_patch_width_cd_delta_minus1 shall be in the range of 0 to Min(32766, CroppedWidth - 1) (including 0 and Min(32766, CroppedWidth - 1)).
[0295] When nnpfc_constant_patch_size_flag is equal to 0, nnpfc_extended_patch_height_cd_delta_minus1 plus 1 plus 2 * npfc_overlap indicates the greatest common divisor of all allowed values of the height of the extended patch required for the input to the NNPF. The value of nnpfc_extended_patch_height_cd_delta_minus1 shall be in the range of 0 to Min(32766, CroppedHeight - 1) (including 0 and Min(32766, CroppedHeight - 1)).
[0296] Let the variables inpPatchWidth and inpPatchHeight be the patch size width and the patch size height, respectively.
[0297] If nnpfc_constant_patch_size_flag is equal to 0, then the following applies:
[0298] - The values of inpPatchWidth and inpPatchHeight are provided by external means not specified in this document, or are set by the post-processor itself.
[0299] The value of -inpPatchWidth + 2*nnpfc_overlap shall be a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2*nnpfc_overlap, and inpPatchWidth shall be less than or equal to CroppedWidth. The value of inpPatchHeight + 2*nnpfc_overlap shall be a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2*npfc_overlap, and inpPatchHeight shall be less than or equal to CroppedHeight.
[0300] Otherwise (when npfc_constant_patch_size_flag equals 1), the value of inpPatchWidth is set to be equal to nnpfc_patch_width_minus1 + 1, and the value of inpPatchHeight is set to be equal to nnpfc_patch_height_minus1 + 1.
[0301] The variables outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, and outPatchCHeight are derived as follows:
[0302] [Table 16]
[0303]
[0304] The requirement for bitstream consistency is that outPatchWidth * CroppedWidth shall be equal to nnpfcOutputPicWidth * inpPatchWidth, and outPatchHeight * CroppedHeight shall be equal to nnpfcOutputPicHeight * inpPatchHeight.
[0305] The nnpfc_padding_type indicates the padding process when referring to sample positions outside the boundaries of the input picture as described in Table 17. In the bitstream conforming to this version of this document, the value of nnpfc_padding_type shall be in the range of 0 to 4 (inclusive of 0 and 4). The values of nnpfc_padding_type from 5 to 15 (inclusive of 5 and 15) are reserved for future use and shall not be present in the bitstream conforming to this version of this document. The decoder conforming to this version of this document shall ignore NNPFC SEI messages with nnfc_padding_type in the range of 5 to 15 (inclusive of 5 and 15). Values of nnpfc_padding_type greater than 15 shall not be present in the bitstream conforming to this version of this document and are not reserved for future use.
[0306] [Table 17]
[0307]
[0308] The nnpfc_luma_padding_val indicates the luma value to be used for padding when nnpfc_padding_type is equal to 4. The value of nnpfc_luma_padding_val shall be in the range of 0 to (1<<BitDepthY)-1 (inclusive of 0 and (1<<BitDepthY)-1).
[0309] The nnpfc_cb_padding_val indicates the Cb value to be used for padding when nnpfc_padding_type is equal to 4. The value of nnpfc_cb_padding_val shall be in the range of 0 to (1<<BitDepthC)-1 (inclusive of 0 and (1<<BitDepthC)-1).
[0310] The nnpfc_cr_padding_val indicates the Cr value to be used for padding when nnpfc_padding_type is equal to 4. The value of nnpfc_cr_padding_val shall be in the range of 0 to (1<<BitDepthC)-1 (inclusive of 0 and (1<<BitDepthC)-1).
[0311] The function InpSampleVal(y, x, picHeight, picWidth, croppedPic, cIdx) returns the value of sampleVal derived as follows, where the inputs are the vertical sample position y, the horizontal sample position x, the picture height picHeight, the picture width picWidth, the sample array croppedPic, and the component index cIdx (equal to 0 for luminance, 1 for Cb, and 2 for Cr):
[0312] For the inputs to the function InpSampleVal(), the vertical position is listed before the horizontal position to be compatible with the input tensor convention of some inference engines.
[0313] [Table 18]
[0314]
[0315] The NNPF PostProcessingFilter() is the target NNPF derived as in the semantics of the NNPFA SEI message. The following example process can be used with the NNPF post-processing filter to generate, in a patch-by-patch manner, filtered and / or interpolated pictures that contain the Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, respectively, as indicated by nnpfc_out_order_idc:
[0316] [Table 19]
[0317]
[0318] The NNPF-generated picture with index i contains the sample arrays FilteredYPic[i], FilteredCbPic[i], and FilteredCrPic[i] (when present) derived by Equation 99. The NNPF-generated pictures do not include overlapping regions.
[0319] The NNPF process consists of the process defined by Equation 99, followed by outputting the NNPF-generated pictures in their increasing index order, where all NNPF-generated pictures interpolated by the NNPF are output, and those NNPF-generated pictures corresponding to any input pictures of the NNPF are output, as specified in the semantics of the NNPFA SEI message.
[0320] The nnpfc_complexity_info_present_flag being equal to 1 specifies the existence of one or more syntax elements indicating the complexity of the NNPF associated with the nnpfc_id. The nnpfc_complexity_info_present_flag being equal to 0 specifies the non - existence of syntax elements indicating the NNPF complexity associated with the nnpfc_id.
[0321] The nnpfc_parameter_type_idc being equal to 0 indicates that the neural network uses only integer parameters, the nnpfc_parameter_type_flag being equal to 1 indicates that the neural network can use floating - point or integer parameters, the nnpfc_parameter_type_idc being equal to 2 indicates that the neural network uses only binary parameters, the nnpfc_parameter_type_idc being equal to 3 is reserved for future use by ITU - T|ISO / IEC and shall not be present in the bitstream conforming to this version of the document. A decoder conforming to this version of the document shall ignore the NNPFC SEI message with nnpfc_parameter_type_idc equal to 3.
[0322] The nnfc_log2_parameter_bit_length_minus3 being equal to 0, 1, 2, and 3 indicates that the neural network does not use parameters with bit lengths greater than 8, 16, 32, and 64 respectively. When the nnpfc_parameter_type_idc is present and the nnfc_log2_parameter_bit_length_minus3 is absent, the neural network does not use parameters with bit lengths greater than 1.
[0323] The nnpfc_num_parameters_idc represents the maximum number of NNPF neural network parameters, in units of powers of 2048. The nnpfc_num_parameters_idc being equal to 0 indicates that the maximum number of neural network parameters is unknown. The value of nnpfc_num_parameters_idc shall be in the range from 0 to 52 (inclusive of 0 and 52). Values of nnpfc_num_parameters_idc greater than 52 are reserved for future use by ITU - T|ISO / IEC and shall not be present in the bitstream conforming to this version of the document. A decoder conforming to this version of the document shall ignore the NNPFC SEI message with nnfc_num_parameters_idc greater than 52.
[0324] If the value of nnpfc_num_parameters_idc is greater than zero, then export the following variable maxNumParameters:
[0325] [Equation 9]
[0326]
[0327] The requirement for bitstream consistency is that the number of neural network parameters of the NNPF should be less than or equal to maxNumParameters.
[0328] nnpfc_num_kmac_operations_idc greater than 0 indicates that the maximum number of multiply-accumulate operations per sample of the NNPF is less than or equal to nnpfc_num_kmac_operations_idc * 1000. nnpfc_num_kmac_operations_idc equal to 0 indicates that the maximum number of multiply-accumulate operations of the network is unknown. The value of nnpfc_num_kmac_operations_idc should be in the range of 0 to 2 32 -2 (including 0 and 2 32 -2).
[0329] nnpfc_total_ kilobyte _size greater than 0 indicates the total size in kilobytes required to store the uncompressed parameters of the neural network. The total size in bits is a number equal to or greater than the sum of the bits used to store each parameter. nnpfc_total_kilobyte_size is the total size in bits divided by 8000, rounded up. nnpfc_total_kilobyte_size equal to 0 indicates that the total size required to store the parameters of the neural network is unknown. The value of nnpfc_total_ kilobyte _size should be in the range of 0 to 2 32 -2 (0 and 2 32 -2).
[0330] nnpfc_metadata_extension_num_bits being equal to 0 specifies that nnpfc_reserved_metadata_extension does not exist. nnpfc_metadata_extension_num_bits being greater than 0 specifies the length of nnpfc_reserved_metadata_extension in bits. nnpfc_metadata_extension_num_bits shall be equal to 0 in this version of this document. Values in the range of 1 to 2048 (including the endpoints) of nnpfc_metadata_extension_num_bits are reserved for future use by ITU-T|ISO / IEC and shall not be present in a bitstream conforming to this version of this document. Decoders conforming to this version of this document shall allow any value of nnpfc_metadata_extension_num_bits in the range of 0 to 2048 (including 0 and 2048). Values of nnpfc_metadata_extension_num_bits greater than 2048 shall not be present in a bitstream conforming to this version of this document and are not reserved for future use.
[0331] nnpfc_reserved_metadata_extension shall not be present in a bitstream conforming to this version of this document. However, decoders conforming to this version of this document shall ignore the presence and value of nnpfc_reserved_metadata_extension. When present, the length of nnpfc_reserved_metadata_extension in bits is equal to nnpfc_metadata_extension_num_bits.
[0332] In a bitstream conforming to this version of this document, nnpfc_reserved_zero_bit_b shall be equal to 0. Decoders shall ignore NNPFC SEI messages in which nnpfc_reserved_zero_bit_b is not equal to 0.
[0333] nnpfc_payload_byte[i] contains the i-th byte of a bitstream conforming to ISO / IEC 15938-17. The byte sequence nnpfc_payload_byte[i] for all current values of i shall be a complete bitstream conforming to ISO / IEC 15938-17.
[0334] Neural network post - filter activation SEI message (NNFPA)
[0335] The syntax structure of NNFPA is shown in Table 20.
[0336] [Table 20]
[0337]
[0338] The NNPFA syntax structure in Table 20 can be signaled in the form of an SEI message. The SEI message that signals the NNPFA syntax structure in Table 19 can be referred to as an NNPFA SEI message.
[0339] The Neural Network Post - Filter Activation (NNPFA) SEI message activates or deactivates the possible use of the target Neural Network Post - Processing Filter (NNPF) identified by nnpfa_target_id and nnpfa_target_base_flag for post - processing filtering of a group of pictures. For a specific picture for which the NNPF is activated, the target NNPF is derived as follows:
[0340] - If nnpfa_target_base_flag is equal to 1, the target NNPF is the base NNPF for which nnpfc_id is equal to nnpfa_target_id.
[0341] - Otherwise (nnpfa_target_base_flag is equal to 0), the target NNPF is the NNPF specified by the last NNPFC SEI message for which nnpfc_id is equal to nnpfa_target_id, which is before the first VCL NAL unit of the current picture in decoding order and is not a duplicate of the NNPFC SEI message that includes the base NNPF.
[0342] For example, when the NNPF is intended for different purposes or for filtering different color components, there can be several NNPFA SEI messages for the same picture.
[0343] nnpfa_target_id indicates the target NNPF, which is specified by one or more NNPFC SEI messages that are related to the current picture and have an nnpfc_id equal to nnpfa_target_id.
[0344] The value of nnpfa_target_id shall be in the range from 0 to 2 32 −2 (including 0 and 2 32 −2). The range from 256 to 511 and the nnpfa_target_id values in the range from 2 31 to 2 32 −2 can be reserved for future use. The decoder must ignore nnpfa_target_id values in the range from 256 to 511 or 2 31from 0 to 2 32 NNPFA SEI messages within the range of -2.
[0345] A NNPFA SEI message with a specific value of nnpfa_target_id will not exist in the current PU unless one or both of the following conditions are true:
[0346] - Within the current CLVS, there exists a NNPFC SEI message whose nnpfc_id is equal to the specific value of nnpfa_target_id that exists in the PU before the current PU in decoding order.
[0347] - There exists a NNPFC SEI message with a nnpfc_id equal to the specific value of nnpfa_target_id in the current PU.
[0348] When the PU contains both a NNPFC SEI message with a specific value of nnpfc_id and a NNPFA SEI message with a nnpfa_target_id equal to the specific value of nnpfc_id, the NNPFC SEI message will be before the NNPFA SEI message in decoding order.
[0349] nnpfa_cancel_flag equal to 1 indicates canceling the persistence of the target NNPF established by any previous NNPFA SEI message with the same nnpfa_target_id as the current SEI message, i.e., the target NNPF is no longer used unless it is activated by another NNPFA SEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 0. nnpfa_cancel_flag equal to 0 indicates that nnpfa_target_base_flag, nnpfa_persistence_flag, and nnpfa_num_output_entries follow.
[0350] nnpfa_target_base_flag equal to 1 specifies that the target NNPF is the base NNPF with a nnpfc_id equal to nnpfa_target_id. nnpfa_target_base_flag equal to 0 specifies that the target NNPF is the NNPF specified by the last NNPFC SEI message with a nnpfc_id equal to nnpfa_target_id, which is before the first VCL NAL unit of the current picture in decoding order and is not a duplicate of the NNPFC SEI message containing the base NNPF.
[0351] The nnpfa_persistence_flag specifies the persistence of the target NNPF for the current layer. An nnpfa_persistence_flag equal to 0 specifies that the target NNPF can be used for post-processing filtering for the current picture. An nnpfa_persistence_flag equal to 1 specifies that the target NNPF can be used for post-processing filtering for the current picture and all subsequent pictures of the current layer in output order until one or more of the following conditions are true:
[0352] - A new CLVS for the current layer starts.
[0353] - The bitstream ends.
[0354] - A picture in the current layer associated with an NNPFA SEI message having the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1 is output, which follows the current picture in output order.
[0355] The target NNPF shall not be used for this subsequent picture in the current layer associated with an NNPFA SEI message having the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1.
[0356] Let nnpfcTargetPictures be the set of pictures related to the last NNPFC SEI message having an nnpfc_id equal to nnpfa_target_id that is before the current NNPFA SEI message in decoding order. Let nnpfaTargetPictures be the set of pictures for which the target NNPF is activated by the current NNPFA SEI message. The requirement for bitstream consistency is that any picture included in nnpfTargetPictures should also be included in nnpfcTargetPictures.
[0357] The nnpfa_num_output_entries specifies the number of nnpfa_output_flag[i] syntax elements present in the NNPFA SEI message. The value of nnpfa_num_output_entries shall be in the range of 0 to NumInpPicsInOutputTensor (including 0 and NumInpPicsInOutputTensor).
[0358] nnpfa_output_flag[i] being equal to 1 specifies that the NNPF process activated by the NNPFA SEI message outputs the NNPF-generated picture corresponding to the input picture with index InpIdx[i], where the NNPF process is specified in the semantics of the NNPF C SEI message. nnpfa_output_flag[i] being equal to 0 specifies that the NNPF process activated by the NNPFA SEI message does not output the NNPF-generated image corresponding to the input image with index InpIdx[i]. When nnpfa_num_output_entries is less than NumInpPicsInOutputTensor, for each value of i in the range from nnpfa_num_output_entries to NumInpPicsInOutputTensor - 1 (inclusive of nnpfa_num_output_entries and NumInpPicsInOutputTensor - 1), it is inferred that nnpfa_output_flag[i] is equal to 1.
[0359] Post - filter hint
[0360] The syntax structure of the post-filter hint is shown in Table 21.
[0361] [Table 21]
[0362]
[0363] The syntax structure of the post-filter hint in Table 21 can be signaled in the form of an SEI message. The SEI message that signals the syntax structure of the post-filter hint in Table 21 can be referred to as the post-filter hint SEI message.
[0364] This SEI message provides the coefficients of the post-filter or related information for designing the post-filter, potentially for use in post-processing a set of pictures after decoding and outputting them to obtain improved display quality.
[0365] filter_hint_cancel_flag being equal to 1 indicates that the SEI message cancels the persistence of any previous post-filter hint SEI message applied to the current layer in output order.
[0366] The filter_hint_persistence_flag specifies the persistence of the post-filtering hint SEI message for the current layer. A filter_hint_persistence_flag equal to 0 specifies that the post-filtering hint applies to the current decoded picture. A filter_hint_persistence_flag equal to 1 specifies that the post-filtering hint SEI message applies to the current decoded picture and persists for all subsequent pictures of the current layer in output order until one or more of the following conditions become true:
[0367] - A new CLVS for the current layer begins.
[0368] - The bitstream ends.
[0369] - The picture in the current layer in the AU associated with the post-filtering hint SEI message is output, where the picture is after the current picture in output order.
[0370] The filter_hint_size_y specifies the vertical size of the filter coefficients or associated array. The value of filter_hint_size_y shall be in the range of 1 to 15, inclusive of 1 and 15.
[0371] The filter_hint_size_x specifies the horizontal size of the filter coefficients or associated array. The value of filter_hint_size_x shall be in the range of 1 to 15, inclusive of 1 and 15.
[0372] The filter_hint_type identifies the type of the transmitted filter hint as specified in Table 19. The value of filter_hint_type shall be in the range of 0 to 2, inclusive of 0 and 2. A value of filter_hint_type equal to 3 is reserved for future use and shall not be present in the bitstream conforming to this version of this document. A decoder conforming to this version of this document shall ignore a post-filtering hint SEI message with filter_hint_type equal to 3.
[0373] [Table 19]
[0374]
[0375] The filter_hint_chroma_coeff_present_flag equal to 1 specifies that the filter coefficients for chroma are present. The filter_hint_chroma_coeff_present_flag equal to 0 specifies that there are no filter coefficients for chroma.
[0376] filter_hint_value[cIdx][cy][cx] specifies an element of the filter coefficients or the cross-correlation matrix between the original signal and the decoded signal with 16-bit precision. The value of filter_hint_value[cIdx][cy][cx] shall be in the range of -2 31 +1 to 2 31 -1 (including -2 31 +1 and 2 31 -1). cIdx specifies the relevant color component, cy represents the counter in the vertical direction, and cx represents the counter in the horizontal direction. Depending on the value of filter_hint_type, the following applies:
[0377] - If filter_hint_type is equal to 0, then send the coefficients of a 2-dimensional finite impulse response (FIR) filter with a size of filter_hint_size_y * filter_hint_size_x.
[0378] - Otherwise, if filter_hint_type is equal to 1, then send the filter coefficients of two 1-dimensional FIR filters. In this case, filter_hint_size_y shall be equal to 2. The index cy equal to 0 specifies the filter coefficients of the horizontal filter, and the cy equal to 1 specifies the filter coefficients of the vertical filter. During the filtering process, the horizontal filter is applied first, and the result is filtered by the vertical filter.
[0379] - Otherwise (filter_hint_type is equal to 2), the sent hint specifies the cross-correlation matrix between the original signal s and the decoded signal s'.
[0380] The normalized cross-correlation matrix of the relevant color component with a size of filter_hint_size_y * filter_hint_size_x identified by cIdx is defined as follows:
[0381] [Equation 10]
[0382]
[0383] Where s represents an array of samples of the color component cIdx of the original picture, s' represents the corresponding array of the decoded picture, h represents the vertical height of the relevant color component, w represents the horizontal width of the relevant color component, bitDepth represents the bit depth of the color component, OffsetY is equal to (filter_hint_size_y >> 1), OffsetX is equal to (filter_hint_size_x >> 1), 0 <= cy < filter_hint_size_y and 0 <= cx_hint_size_x.
[0384] The decoder can derive a Wiener post-filter from the cross-correlation matrix of the original and decoded signals and the autocorrelation matrix of the decoded signal.
[0385] Problems of the prior art
[0386] In the current design of the Neural Network Post-Filter Characteristics (NNPFC) and Neural Network Post-Filter Activation Supplementary Enhancement Information (SEI) message, the NNPFC SEI message can activate a base filter (i.e., the NNPFC SEI message with nnpfc_base_flag equal to 1) or update a filter (i.e., the NNPFC SEI message with nnpfc_base_flag equal to 0) based on the nnpfa_target_base_flag.
[0387] There are the following problems in identifying the target NNPF. Figure 6 It is a diagram showing problem situations that may occur when identifying the target NNPF to be activated based on the Neural Network Post-Filter Activation (NNPFA) SEI message.
[0388] When the value of nnpfa_target_base_flag is 0, there are two conflicting specifications to define which NNPFC will be activated.
[0389] - As shown in Table 23, at the beginning of the semantics for NNPFA, the following is specified: Otherwise (nnpfa_target_base_flag is 0), the target NNPF is the NNPF specified by the last NNPFC SEI message with nnpfc_id equal to nnpfa_target_id, which is before the first Video Coding Layer (VCL) Network Abstraction Layer (NAL) unit of the current picture in decoding order and is not a duplicate of the NNPFC SEI message containing the base NNPF.
[0390] [Table 23]
[0391]
[0392] As shown in Table 24, in the second half of the semantics of nnpfa_target_base_flag itself, the following is specified: when nnpfa_target_base_flag is equal to 0, it is specified that the target NNPF is the NNPF specified by the last NNPFC SEI message with nnpfc_id equal to nnpfa_target_id, which is before the first VCL NAL unit of the current picture in decoding order and is not a duplicate of the NNPFC SEI message containing the base NNPF.
[0393] [Table 24]
[0394]
[0395] When nnpfa_target_base_flag is 0 but there is no update of the NNPFC SEI with nnpfc_id equal to the target NNPF's nnpfa_target_id, it is unclear whether the decoder should activate the base filter. This may occur in the Figure 6 example. In the Figure 6 example, the NNPFA SEI message exists in the access unit (AU) including the predictor (PU) with picture order count 3 (POC3). The corresponding NNPFA SEI message has a specific nnpfa_target_id value, and the value of nnpfa_target_base_flag is 0. The nnpfa_target_base_flag with a value of 0 indicates that the target NNPF is not the base NNPF but an updated NNPF.
[0396] Specifically, the target NNPF activated by the NNPFA SEI message is the NNPF specified by the last NNPFC SEI message with nnpfc_id equal to nnpfa_target_id, which is before the first VCL NAL unit of the current picture in decoding order and is not a duplicate of the NNPFC SEI message containing the base NNPF. However, as shown in the Figure 6 example, when there is no update of the NNPFC SEI message with nnpfc_id equal to nnpfa_target_id, it is unclear whether to activate the base NNPF, regardless of the fact that the value of nnpfa_target_base_flag is equal to 0.
[0397] Embodiment
[0398] The following embodiments provide a solution to the above problems. An overview of the embodiments is given below.
[0399] 1. It is constrained that when the value of nnpfa_target_base_flag in the NNPFA SEI message is equal to 0, there should be at least one NNPFC SEI message in which nnpfc_id is equal to nnpfa_target_id and nnpfc_base_flag is equal to 0, and which is before the NNPFA SEI message in decoding order.
[0400] 2. Alternatively, it can be clarified or specified that when nnpfa_target_base_flag is equal to 0, in the case where there is no NNPFC SEI message with nnpfc_id equal to nnpfa_target_id and nnpfa_base_flag equal to 0 that is before the NNPFA SEI message in decoding order, the target NNPF can be the base NNPF.
[0401] The present invention proposes various embodiments to solve the above problems. The embodiments proposed in the present invention can be executed alone or in combination of two or more.
[0402] Hereinafter, the NNPFC can be the NNPFC syntax structure of Tables 1 to 3 above and signaled in the form of an SEI message. In this case, the NNPFC can be an NNPFC SEI message. The NNPFA can be the NNPFA syntax structure of Table 20 above and signaled in the form of an SEI message. In this case, the NNPFA can be an NNPFA SEI message. The post-filtering hint can be the post-filtering hint syntax structure of Table 21 above, and can be signaled in the form of an SEI message. In this case, the post-filtering hint can be a post-filtering hint SEI message.
[0403] Embodiment 1
[0404] Example 1 relates to the above description of Overview 1. According to Example 1, when the value of nnpfa_target_base_flag in the NNPFA SEI message is equal to 0, at least one NNPFC SEI message having an nnpfc_id equal to nnpfa_target_id and an nnpfc_base_flag equal to 0 shall be before the NNPFA SEI message in decoding order. For example, based on the fact that the value of nnpfa_target_base_flag in the NNPFA SEI message is equal to 0, at least one NNPFC SEI message having an nnpfc_id equal to nnpfa_target_id and an nnpfc_base_flag equal to 0 is encoded to be before the NNPFA SEI message in decoding order. In Example 1, the target NNPF for activating a specific picture of NNPF can be derived as follows.
[0405] nnpfa_target_base_flag being equal to 1 indicates that the target NNPF is the base NNPF with nnpfc_id equal to nnpfa_target_id. An nnpfa_target_base_flag equal to 0 does not indicate that the target NNPF is the base NNPF, but rather indicates that the target NNPF is the NNPF specified by the last NNPFC SEI message having an nnpfc_id equal to nnpfa_target_id, which is before the first VCL NAL unit of the current picture in decoding order and is not a duplicate of the NNPFC SEI message containing the base NNPF. When nnpfa_target_base_flag is equal to 0, there is at least one NNPFC SEI with nnpfc_id equal to nnpfa_target_id and nnpfa_base_flag equal to 0, which is before the NNPFA SEI message in decoding order. Thus, when nnpfa_target_base_flag is equal to 0, there is an updated NNPF specified by the NNPFC SEI message before the NNPFA SEI (in decoding order), and the updated NNPF can be activated without uncertainty about whether to activate the base NNPF.
[0406] Figure 7 is an example of a case where an embodiment of the present invention is applied.
[0407] Figure 7The example of shows the case where the NNPFA SEI message exists in the AU including the PU having POC3. The corresponding NNPFA SEI message has a specific nnpfa_target_id value, and the value of nnpfa_target_base_flag is 0. The nnpfa_target_base_flag with the value 0 indicates that the target NNPF is not the base NNPF but the updated NNPF.
[0408] In this case, according to Embodiment 1, there is at least one NNPFC SEI message whose nnpfc_id is equal to nnpfa_target_id and nnpfa_base_flag is equal to 0, and it is before the NNPFA SEI message in the decoding order, as Figure 7 shown. Therefore, since the updated NNPF is specified by the NNPFC SEI message before the NNPFA SEI message (in the decoding order), the target NNPF activated by the NNPFA SEI message can be determined as the updated NNPF.
[0409] Embodiment 2
[0410] Embodiment 2 relates to the above description of Overview 2. According to Embodiment 2, when the NNPFC SEI message with nnpfa_target_base_flag equal to 0, nnpfc_id equal to nnpfa_target_id, and nnpfc_base_flag equal to 0 is before the NNPFA SEI message in the decoding order, the target NNPF can be the base NNPF. For example, when the NNPFC SEI message with nnpfa_target_base_flag equal to 0, nnpfc_id equal to nnpfa_target_id, and nnpfc_base_flag equal to 0 is before the NNPFA SEI message in the decoding order, it can be clarified or specified that the target NNPF is determined as the base NNPF. In Embodiment 2, the target NNPF of a specific picture that activates the NNPF can be derived as follows.
[0411] The nnpfa_target_base_flag being equal to 1 indicates that the target NNPF is the base NNPF with nnpfc_id equal to nnpfa_target_id. The nnpfa_target_base_flag equal to 0 does not indicate that the target NNPF is the base NNPF, but rather indicates that the target NNPF is the NNPF specified by the last NNPFC SEI message with nnpfc_id equal to nnpfa_target_id, which is before the first VCL NAL unit of the current picture in decoding order and is not a repetition of the NNPFC SEI message including the base NNPF.
[0412] When nnpfa_target_base_flag is equal to 0 and there is no NNPFC SEI message with nnpfc_id equal to nnpfa_target_id and nnpfa_base_flag equal to 0 that is before the NNPFA SEI message in decoding order, the target NNPF can be determined as the base NNPF. Thus, when the value of nnpfa_target_base_flag is 0 but there is no updated NNPF to be activated, the base NNPF can be determined as the target NNPF without uncertainty about the target to be activated.
[0413] Figure 8 It is another example of a case where an embodiment of the present invention is applied.
[0414] Figure 8 The example shows a case where there is an NNPFA SEI message in a PU with POC3, the corresponding NNPFA SEI message has a specific nnpfa_target_id value, and the value of nnpfa_target_base_flag is 0. The nnpfa_target_base_flag with a value of 0 indicates that the target NNPF is not the base NNPF but an updated NNPF.
[0415] In this case, when the NNPFC SEI message with nnpfc_id equal to nnpfa_target_id and nnpfa_base_flag equal to 0 is not before the NNPFA SEI message in the decoding order as Figure 8 shown, according to Embodiment 2, the base NNPF can be determined as the target NNPF to be activated by the NNPFA SEI message.
[0416] Image encoding method and image decoding method
[0417] An image encoding method and an image decoding method according to an embodiment of the present invention will be described below.
[0418] Figure 9 is an example of an image encoding method to which embodiments of the present disclosure are applicable, and Figure 10 is an example of an image decoding method to which embodiments of the present disclosure are applicable.
[0419] Referring to Figure 9 , at least one neural network that can be used as a post-processing filter can be determined, and information about the determined neural network can be encoded as at least one NNPFC SEI message (S610).
[0420] It can be determined whether to activate the target NNPF applicable to the current picture, and information about the determined target NNPF can be encoded as an NNPFA SEI message (S620). The process S620 of determining whether to activate the target NNPF can include a process of determining the target NNPF, a process of determining whether to cancel the persistence of the target NNPF, a process of determining whether the target NNPF persists, a process of determining whether to apply the basic NNPF, and the like.
[0421] Post-filter coefficients, related information, etc. for the design of the post-filter can be encoded as a post-filter hint SEI message (S630). The NNPFC SEI message, the NNPFA SEI message, and / or the post-filter hint SEI message can be included in the NNPF SEI message.
[0422] When the NNPF SEI message is applied to the current picture in the image decoding apparatus 300, the target NNPF can be determined or activated by various embodiments of the present invention.
[0423] For example, it can be determined which NNPF will be activated based on the NNPFA SEI message. Here, the target NNPF to be activated can be identified from the target identification information and the target base flag information included in the NNPFA SEI message. For example, it can be determined whether to activate the basic NNPF or the updated NNPF based on the NNPFA SEI message. As a specific example, it can be determined whether to activate the basic NNPF or the updated NNPF based on the target base flag information included in the NNPFA SEI message.
[0424] For example, the target identification information may be information corresponding to nnpfa_target_id, which indicates the target neural network post - processing filter specified by one or more NNPFC SEI messages. The one or more NNPFC SEI messages are related to the current picture and have the same identification information nnpfc_id as the target identification information. The base flag information may be information corresponding to nnpfc_base_flag indicating whether the NNPF specified by the NNPFC SEI message is a base NNPF, and the target base flag information may be information corresponding to nnpfa_target_base_flag indicating whether the target NNPF is a base NNPF with the same identification information as the target identification information.
[0425] Reference Figure 10 , the SEI messages of the NNPF to be applied to the current picture can be obtained from the bitstream. The SEI messages for the NNPF may include NNPFC SEI messages, NNPFA SEI messages, and / or post - filter hint SEI messages.
[0426] When the SEI message for the NNPF is applied to the current picture, at least one neural network that can be used as a post - processing filter can be determined based on at least one NNPFC SEI message included in the SEI message for the NNPF (S710).
[0427] Whether to activate the target NNPF that can be applied to the current picture can be determined based on at least one NNPFA SEI message obtained from the bitstream (S720). The process S720 of determining whether to activate the target NNPF may include a process of determining the target NNPF, a process of determining whether to cancel the persistence of the target NNPF, a process of determining whether the target NNPF persists, a process of determining whether to apply the base NNPF, etc.
[0428] When the target NNPF is activated (by determination or persistence), the target neural network post - processing filter can be applied to the current picture (S730).
[0429] Various embodiments of the present invention can be used to determine or activate the target NNPF.
[0430] For example, it can be determined based on the NNPFA SEI message which NNPF will be activated. Here, the target NNPF to be activated can be identified from the target identification information and the target base flag information included in the NNPFA SEI message. For example, it can be determined based on the NNPFA SEI message whether to activate the base NNPF or the updated NNPF. As a specific example, it can be determined based on the target base flag information included in the NNPFA SEI message whether to activate the base NNPF or the updated NNPF.
[0431] Figure 11 shows a process of determining a target NNPF according to an embodiment of the present disclosure. For example, Figure 11 the process shown in can be executed by an image decoding device.
[0432] Refer to Figure 11 , and obtain the target identification information and target base flag information of the NNPFA SEI message (S1110). For example, the target identification information and target base flag information of the NNPFA SEI message can be obtained from the bitstream.
[0433] For example, the target identification information of the NNPFA SEI message can be information indicating a target neural network post-processing filter specified by one or more NNPFC SEI messages related to the current picture and having the same identification information as the target identification information. The target base flag information can be information indicating whether the target NNPF is a base NNPF having the same identification information as the target identification information.
[0434] In addition, for example, when the target base flag information of the NNPFA SEI message has a value of 1, this can indicate that the target NNPF is a base NNPF. When the target base flag information of the NNPFA SEI message has a value of 0, this can indicate that the target NNPF is not a base NNPF. For example, when the value of the target base flag information is 0, this can indicate that the target NNPF is an updated NNPF.
[0435] Based on the target base flag information indicating a base NNPF (yes in S1120), that is, based on the value of the target base flag information being equal to 1, the base NNPF can be determined as the target NNPF (S1130).
[0436] Based on the target base flag information not indicating a base NNPF (no in S1120), that is, based on the value of the target base flag information being equal to 0, the target NNPF can be determined by an NNPFC SEI message that is before the NNPFA SEI message in the decoding order, has the same identification information as the NNPFA SEI message, and has a base flag information that does not indicate a base NNPF (S1140).
[0437] For this reason, the target base flag information based on the NNPFA SEI message does not indicate the base NNPF. There is an NNPFC SEI message that has the same identification information as the NNPFA SEI message and has base flag information that does not indicate the base NNPF and is before the NNPFA SEI message in the decoding order, that is, the value of the base flag information is equal to 0. For example, before the NNPFA SEI message where the value of the target base flag information is equal to 0 (in decoding order), the image encoding device 200 can encode an NNPFC SEI message that has the same identification information as the NNPFA SEI message and has base flag information that does not indicate the base NNPF (that is, the value of the base flag information is equal to 0).
[0438] Figure 12 The process of determining the target NNPF according to an embodiment of the present disclosure is shown. For example, Figure 12 The process shown in can be executed by an image decoding device.
[0439] Refer to Figure 12 , and obtain the target identification information and target base flag information of the NNPFA SEI message (S1210). For example, the target identification information and target base flag information of the NNPFA SEI message can be obtained from the bitstream.
[0440] For example, the target identification information of the NNPFA SEI message can be information indicating a target neural network post - processing filter specified by one or more NNPFC SEI messages related to the current picture and having the same identification information as the target identification information. The target base flag information can be information indicating whether the target NNPF is the base NNPF having the same identification information as the target identification information.
[0441] In addition, for example, when the target base flag information of the NNPFA SEI message has a value of 1, this can indicate that the target NNPF is the base NNPF. When the target base flag information of the NNPFA SEI message has a value of 0, this can indicate that the target NNPF is not the base NNPF. As a specific example, this can indicate that the target NNPF is an updated NNPF.
[0442] Based on the target base flag information indicating the base NNPF (yes in S1220), that is, based on the value of the target base flag information being equal to 1, the base NNPF can be determined as the target NNPF (S1240).
[0443] Based on the target base flag information not indicating the base NNPF (No in S1120), that is, based on the value of the target base flag information being equal to 0, determine whether there is an NNPFC SEI message that is before the NNPFA SEI message in the decoding order, has the same identification information as the NNPFA SEI message, and has base flag information that does not indicate the base NNPF (S1230).
[0444] Based on the presence of the NNPFA SEI message satisfying the condition (Yes in S1230), the NNPF specified by the NNPFC SEI message can be determined as the target NNPF (S1230).
[0445] Based on the absence of the NNPFA SEI message satisfying the condition (No in S1230), the base NNPF can be determined as the target NNPF (S1240).
[0446] According to the above embodiments, it is possible to clearly determine the filter to be activated between the base filter and the updated filter.
[0447] In addition, according to the above embodiments, it is possible to prevent a situation where it is not clear whether to activate the base filter or the updated filter.
[0448] Figure 13 FIG. is an illustration of a content streaming system to which embodiments according to the present disclosure can be applied.
[0449] Reference Figure 13 , a content streaming system to which the embodiments of this document are applied may mainly include an encoding server, a streaming server, a web server, a media memory, a user device, and a multimedia input device.
[0450] The encoding server compresses the content input from a multimedia input device such as a smart phone, a camera, a video camera, etc. into digital data to generate a bitstream and sends the bitstream to the streaming server. As another example, when a multimedia input device such as a smart phone, a camera, a video camera, etc. directly generates a bitstream, the encoding server can be omitted.
[0451] A bitstream can be generated by applying the encoding method or bitstream generation method of the embodiments of this document, and the streaming server can temporarily store the bitstream during the process of sending or receiving the bitstream.
[0452] The streaming server sends multimedia data to the user device based on a request from the user via the web server, and the web server serves as a medium for notifying the user of the service. When the user requests a desired service from the web server, the web server delivers it to the streaming server, and the streaming server transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server is used to control commands / responses between devices in the content streaming system.
[0453] The streaming server may receive content from a media storage and / or encoding server. For example, when content is received from the encoding server, the content may be received in real time. In this case, in order to provide a smooth streaming service, the streaming server may store the bitstream for a predetermined time.
[0454] Examples of user devices may include mobile phones, smartphones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays), digital TVs, desktop computers, digital signage, etc.
[0455] Each server in the content streaming system may operate as a distributed server, and in this case, the data received from each server may be distributed.
[0456] The scope of the present disclosure includes software or machine-executable instructions (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to various embodiments of the present disclosure to be executed on a device or computer, and non-transitory computer-readable media having such software or instructions stored thereon and executable on the device or computer.
[0457] [Industrial Applicability]
[0458] Embodiments according to the present disclosure may be used for encoding / decoding images.
Claims
1. An image decoding method performed by an image decoding apparatus, the method comprising: Obtaining a neural network post-filter characteristic (NNPFC) supplementary enhancement information (SEI) message and a neural network post-filter activation (NNPFA) SEI message; Determining at least one neural network to be used as a neural network post-filter (NNPF) based on the NNPFC SEI message; And Determining whether to activate a target NNPF to be applied to a current picture based on the NNPFA SEI message, wherein the NNPFA SEI message includes target identification information and target base flag information, wherein the target NNPF is determined based on the target identification information and the target base flag information, wherein based on the target base flag information not indicating that the target NNPF is a base NNPF, there is at least one NNPFC SEI message in decoding order before the NNPFA SEI message, wherein the identification information is equal to the target identification information and the value of the base flag information is equal to the value of the target base flag information.
2. The image decoding method according to claim 1, wherein, The target base flag information having a value of 1 indicates that the target NNPF is the base NNPF having the identification information equal to the target identification information.
3. The image decoding method according to claim 1, wherein, The target base flag information having a value of 0 indicates that the target NNPF is an NNPF specified by the NNPFC SEI message having the identification information equal to the target identification information, wherein the NNPF SEI message having the identification information equal to the target identification information is the last NNPF SEI message in decoding order before the first VCL NAL unit of the current picture and is not a repetition of the NNPFC SEI message including the base NNPF.
4. An image encoding method performed by an encoding apparatus, the image encoding method comprising: Encoding at least one neural network that can be used as a post-processing filter into a neural network post-filter characteristic (NNPFC) SEI message; And Encoding whether a target NNPF that can be applied to the current picture is activated into a neural network post-filter activation (NNPFA) SEI message, wherein the NNPFA SEI message includes target identification information and target base flag information, wherein based on the target base flag information not indicating that the target NNPF is a base NNPF, there is at least one NNPFC SEI message in decoding order before the NNPFA SEI message, wherein the identification information is equal to the target identification information and the value of the base flag information is equal to the value of the target base flag information.
5. The image encoding method according to claim 4, wherein, The target base flag information having a value of 1 indicates that the target NNPF is the base NNPF having the identification information equal to the target identification information.
6. The image encoding method according to claim 4, wherein, The target base flag information having a value of 0 indicates that the target NNPF is an NNPF specified by the NNPFC SEI message having the identification information equal to the target identification information. Among them, the NNPF SEI message having the identification information equal to the target identification information is the last NNPF SEI message before the first VCL NAL unit of the current picture in the decoding order, and is not a duplicate of the NNPFC SEI message including the basic NNPF.
7. A computer-readable digital storage medium for storing a bitstream generated using an image coding method, the bitstream including: A neural network post-filter characteristic (NNPFC) SEI message indicating at least one neural network that can be used as a post-processing filter; And A neural network post-filter activation (NNPFA) SEI message indicating whether a target NNPF that can be applied to the current picture is activated, Wherein, the NNPFA SEI message includes target identification information and target basic flag information, Wherein, based on that the target basic flag information does not indicate that the target NNPF is a basic NNPF, there is at least one NNPFC SEI message before the NNPFA SEI message in the decoding order, wherein the identification information is equal to the target identification information and the value of the basic flag information is equal to the value of the target basic flag information.
8. A method for transmitting data of an image, the method including: Generating a bitstream of the image, the bitstream being generated by: encoding at least one neural network that can be used as a post-processing filter into a neural network post-filter characteristic (NNPFC) SEI message, and encoding whether a target NNPF that can be applied to the current picture is activated into a neural network post-filter activation (NNPFA) SEI message; And Sending the data including the bitstream, Wherein, the NNPFA SEI message includes target identification information and target basic flag information, Wherein, based on that the target basic flag information does not indicate that the target NNPF is a basic NNPF, there is at least one NNPFC SEI message before the NNPFA SEI message in the decoding order, wherein the identification information is equal to the target identification information and the value of the basic flag information is equal to the value of the target basic flag information.