Image encoding / decoding method, bit stream transmission method, and recording medium storing bit stream

By restricting NNPFC SEI messages and processing NNPFA SEI messages, the problems of low efficiency and image list mismatch in high-resolution image encoding/decoding are solved, achieving efficient image encoding/decoding and accurate image reconstruction.

CN121002884APending Publication Date: 2025-11-21LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480024138.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-06
Filing Date
2024-04-05
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies suffer from low encoding/decoding efficiency when processing high-resolution and high-quality images, and the supplementary enhancement information (SEI) messages associated with post-filters of neural networks are easily associated with discardable or non-output images, resulting in a mismatch between the input image lists of the encoder and decoder.

Method used

By restricting the Neural Network Post-Filter Feature (NNPFC) SEI message to include only valid image information, avoiding association with discardable or non-output images, and processing the NNPFA SEI message during image encoding/decoding, the consistency of the input image is ensured.

Benefits of technology

It improves the efficiency of image encoding/decoding, prevents mismatches in the input image lists between the encoder and decoder, and ensures the accuracy and efficiency of image reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121002884A_ABST
    Figure CN121002884A_ABST
Patent Text Reader

Abstract

Provided are an image encoding / decoding method, a bitstream transmission method, and a computer-readable recording medium storing a bitstream. An image decoding method according to the present disclosure may comprise the steps of: acquiring unit information including a current picture; and decoding the current picture based on the unit information, in which the type of the current picture may be limited based on a neural network post filter (NNPF) related supplemental enhancement information (SEI) message included in the unit information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The disclosure relates to an image encoding / decoding method, a method of transmitting a bitstream, and a recording medium storing a bitstream, and more particularly, to a method of processing a neural network post-filter. BACKGROUND

[0002] Recently, in various fields, the demand for high-resolution and high-quality images such as high-definition (HD) images and ultra-high-definition (UHD) images is increasing. As the resolution and quality of image data increase, the amount of information or bits transmitted relatively increases compared to existing image data. The increase in the amount of information or bits transmitted results in an increase in transmission and storage costs.

[0003] Therefore, an efficient image compression technique is needed to efficiently transmit, store, and reproduce information about high-resolution and high-quality images. SUMMARY

[0004] TECHNICAL PROBLEM

[0005] An object of the disclosure is to provide an image encoding / decoding method and apparatus having improved encoding / decoding efficiency.

[0006] Further, an object of the disclosure is to provide an image encoding / decoding method and apparatus capable of processing a neural network post-filter (NNPF) related supplemental enhancement information (SEI) message without error.

[0007] Further, an object of the disclosure is to provide an image encoding / decoding method and apparatus in which unit information including a neural network post-filter characteristic (NNPFC) SEI message is restricted not to include a discardable picture.

[0008] Further, an object of the disclosure is to provide an image encoding / decoding method and apparatus in which a NNPFC SEI message is not associated with a discardable picture.

[0009] Further, an object of the disclosure is to provide an image encoding / decoding method and apparatus in which unit information including a neural network post-filter activation (NNPFA) SEI message is restricted not to include a non-output picture.

[0010] Further, an object of the disclosure is to provide an image encoding / decoding method and apparatus in which a NNPFA SEI message is not associated with a non-output picture.

[0011] Further, an object of the disclosure is to provide an image encoding / decoding method and apparatus in which an input picture of a NNPF does not include a discardable picture and / or a non-output picture, thereby preventing mismatch between an input picture list of an encoder and an input picture list of a decoder.

[0012] In addition, an object of the disclosure is to provide a non-transitory computer-readable recording medium storing a bitstream generated by the image encoding method according to the disclosure.

[0013] In addition, an object of the disclosure is to provide a non-transitory computer-readable recording medium storing a bitstream received, decoded, and used for reconstructing an image by the image decoding apparatus according to the disclosure.

[0014] In addition, an object of the disclosure is to provide a method of transmitting a bitstream generated by the image encoding method according to the disclosure.

[0015] The technical problems addressed by the disclosure are not limited to the above-mentioned technical problems, and those skilled in the art will clearly understand other technical problems not described herein from the following description.

[0016] Technical solutions

[0017] The image decoding method according to one aspect of the disclosure can be performed by an image decoding apparatus. The image decoding method can include obtaining unit information including a current picture, and encoding the current picture based on the unit information. The type of the current picture can be restricted based on a neural network post filter (NNPF) related supplemental enhancement information (SEI) message included in the unit information.

[0018] The image encoding method according to another aspect of the disclosure can be performed by an image encoding apparatus. The image encoding method can include encoding a current picture, and configuring unit information including the encoded current picture. The type of the current picture can be restricted based on a neural network post filter (NNPF) related supplemental enhancement information (SEI) message included in the unit information.

[0019] The computer-readable recording medium according to another aspect of the disclosure can store a bitstream generated by the image encoding method or apparatus of the disclosure.

[0020] The transmission method according to another aspect of the disclosure can transmit a bitstream generated by the image encoding method or apparatus of the disclosure.

[0021] The features briefly described above with respect to the disclosure are merely exemplary aspects of the following detailed description of the disclosure, and do not limit the scope of the disclosure.

[0022] Advantageous effects

[0023] According to the disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.

[0024] Also, according to the disclosure, an image encoding / decoding method and apparatus can be provided, which is capable of processing a neural network post-filter (NNPF) related supplemental enhancement information (SEI) message without error.

[0025] Also, according to the disclosure, an image encoding / decoding method and apparatus can be provided, in which unit information including a neural network post-filter characteristic (NNPFC) SEI message is restricted not to include a discardable picture.

[0026] Also, according to the disclosure, an image encoding / decoding method and apparatus can be provided, in which a NNPFC SEI message is not associated with a discardable picture.

[0027] Also, according to the disclosure, an image encoding / decoding method and apparatus can be provided, in which unit information including a neural network post-filter activation (NNPFA) SEI message is restricted not to include a non-output picture.

[0028] Also, according to the disclosure, an image encoding / decoding method and apparatus can be provided, in which a NNPFA SEI message is not associated with a non-output picture.

[0029] Also, according to the disclosure, an image encoding / decoding method and apparatus can be provided, in which an input picture of a NNPF does not include a discardable picture and / or a non-output picture, thereby preventing mismatch between an input picture list of an encoder and an input picture list of a decoder.

[0030] In addition, according to the disclosure, a non-transitory computer-readable recording medium storing a bitstream generated by an image encoding method according to the disclosure can be provided.

[0031] In addition, according to the disclosure, a non-transitory computer-readable recording medium storing a bitstream generated by an image encoding method according to the disclosure can be provided.

[0032] In addition, according to the disclosure, a method of transmitting a bitstream generated by an image encoding method according to the disclosure can be provided.

[0033] Those skilled in the art will appreciate that the effects realized by the disclosure are not limited to what has been particularly described hereinabove and other advantages of the disclosure will be more clearly understood from the detailed description. BRIEF DESCRIPTION OF DRAWINGS

[0034] Figure 1 is a view schematically showing a video encoding system to which embodiments of the disclosure can be applied.

[0035] Figure 2is a view schematically illustrating an image encoding apparatus to which embodiments of the present disclosure can be applied.

[0036] Figure 3 is a view schematically illustrating an image decoding apparatus to which embodiments of the present disclosure can be applied.

[0037] Figure 4 is a view illustrating an interleaving method for deriving a luminance channel.

[0038] Figure 5 is a flowchart for explaining an image encoding method to which embodiments of the present disclosure can be applied.

[0039] Figure 6 is a flowchart for explaining an image decoding method to which embodiments of the present disclosure can be applied.

[0040] Figure 7 is a flowchart for explaining another image encoding method to which embodiments of the present disclosure can be applied.

[0041] Figure 8 is a flowchart for explaining another image decoding method to which embodiments of the present disclosure can be applied.

[0042] Figure 9 is a view illustrating a content streaming system to which embodiments of the present disclosure can be applied. DETAILED DESCRIPTION

[0043] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so as to be easily implemented by those skilled in the art. However, the present disclosure can be implemented in various different forms and is not limited to the embodiments described herein.

[0044] In describing the present disclosure, if it is determined that a detailed description of related known functions or configurations will unnecessarily obscure the gist of the present disclosure, a detailed description thereof will be omitted. In the drawings, portions unrelated to the description of the present disclosure are omitted, and like reference numerals are assigned to like parts.

[0045] In the present disclosure, when one component is "connected", "coupled", or "linked" to another component, it can include not only a direct connection relationship but also an indirect connection relationship in which a middle component exists. In addition, when one component "includes" or "has" another component, it means that a further component can be further included, rather than excluding the other component, unless otherwise specified.

[0046] In the disclosure, the terms first, second, and the like can be used to distinguish one component from another component, unless otherwise specified, and do not limit the order or importance of the components. Accordingly, within the scope of the disclosure, a first component in one embodiment can be referred to as a second component in another embodiment, and likewise, a second component in one embodiment can be referred to as a first component in another embodiment.

[0047] In the disclosure, components are distinguished in order to clearly describe each feature, and it does not mean that the components must be separate. That is, a plurality of components can be integrated and implemented in one hardware or software unit, or one component can be distributed and implemented in a plurality of hardware or software units. Therefore, even if not otherwise specified, embodiments in which such components are integrated or distributed are included within the scope of the disclosure.

[0048] In the disclosure, the components described in various embodiments are not necessarily essential components, and some components can be optional components. Therefore, embodiments consisting of a subset of the components described in one embodiment are also included within the scope of the disclosure. Furthermore, embodiments including other components in addition to the components described in various embodiments are also included within the scope of the disclosure.

[0049] The disclosure relates to encoding and decoding of an image, and unless newly defined in the disclosure, the terms used in the disclosure can have a general meaning commonly used in the technical field to which the disclosure belongs.

[0050] In the disclosure, a "picture" generally refers to a basis representing one image for a specific time period, and a slice / tile is a coding basis constituting a part of a picture. One picture can consist of one or more slices / tiles. Furthermore, a slice / tile can include one or more coding tree units (CTUs).

[0051] In the disclosure, a "pixel" or a "pel" can refer to the smallest unit constituting one picture (or image). Furthermore, a "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or a value of a pixel, and can represent only a pixel / value of a pixel of a luminance component or only a pixel / value of a pixel of a chrominance component.

[0052] In the disclosure, a "unit" can refer to a basic unit of image processing. The unit can include at least one of a specific region of a picture and information related to the region. One unit can include one luminance block and two chrominance (e.g., Cb, Cr) blocks. In some cases, the unit can be used interchangeably with terms such as "sample array", "block", or "region". In general, an MxN block can include a set (or array) of M columns and N rows of samples (or sample array) or transform coefficients.

[0053] In the disclosure, the "current block" can refer to one of a "current coding block", a "current coding unit", a "coding target block", a "decoding target block", or a "processing target block". When performing prediction, the "current block" can refer to a "current prediction block" or a "prediction target block". When performing transform (inverse transform) / quantization (dequantization), the "current block" can refer to a "current transform block" or a "transform target block". When performing filtering, the "current block" can refer to a "filtering target block".

[0054] In addition, in the disclosure, unless explicitly stated as a chroma block, the "current block" can refer to a block including both a luma component block and a chroma component block or a "luma block of the current block". The luma component block of the current block can be explicitly indicated by a term such as "luma block" or "current luma block", which clearly indicates that it is a luma component block. Additionally, the chroma component block of the current block can be explicitly indicated by a term such as "chroma block" or "current chroma block", which clearly indicates that it is a chroma component block.

[0055] In the disclosure, the terms " / " and "and" are to be interpreted to mean "and / or". For example, the expressions "A / B" and "A, B" can mean "A and / or B". In addition, "A / B / C" and "A, B, C" can mean "at least one of: A, B, and / or C".

[0056] In the disclosure, the term "or" is to be interpreted to mean "and / or". For example, the expression "A or B" can include 1) only "A", 2) only "B", and / or 3) both "A and B". In other words, in the disclosure, the term "or" is to be interpreted to mean "additionally or alternatively".

[0057] Overview of a video encoding system

[0058] Figure 1 is a view showing a video coding system to which embodiments of the disclosure can be applied.

[0059] The video coding system according to the embodiments can include an encoding apparatus 10 and a decoding apparatus 20. The encoding apparatus 10 can deliver encoded video and / or image information or data in the form of a file or a stream to the decoding apparatus 20 via a digital storage medium or a network.

[0060] The encoding apparatus 10 according to the embodiments can include a video source generator 11, an encoding unit (encoder) 12, and a transmitter 13. The decoding apparatus 20 according to the embodiments can include a receiver 21, a decoding unit (decoder) 22, and a renderer 23. The encoding unit 12 can be referred to as a video / image encoding apparatus, and the decoding unit 22 can be referred to as a video / image decoding apparatus. The transmitter 13 can be included in the encoding unit 12. The receiver 21 can be included in the decoding unit 22. The renderer 23 can include a display, and the display can be configured as a separate device or an external component.

[0061] The video source generator 11 can obtain a video / image through a process of capturing, synthesizing, or generating a video / image. The video source generator 11 can include a video / image capturing device and / or a video / image generating device. The video / image capturing device can include, for example, one or more cameras, a video / image archive containing previously captured videos / images, and the like. The video / image generating device can include, for example, a computer, a tablet, and a smartphone, and can generate a video / image (electronically). For example, a virtual video / image can be generated by a computer or the like. In this case, the video / image capturing process can be replaced by a process of generating relevant data.

[0062] The encoding unit 12 can encode an input video / image. The encoding unit 12 can perform a series of processes such as prediction, transformation, and quantization to achieve compression and coding efficiency. The encoding unit 12 can output encoded data (encoded video / image information) in the form of a bitstream.

[0063] The transmitter 13 can obtain the encoded video / image information or data output in the form of a bitstream, and forward it to the receiver 21 of the decoding apparatus 20 or another external object in the form of a file or a stream via a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, and the like. The transmitter 13 can include an element for generating a media file through a predetermined file format, and can include an element for transmission through a broadcasting / communication network. The transmitter 13 can be provided as a separate transmission apparatus from the encoding device 12, in which case the transmission apparatus can include at least one processor that obtains the encoded video / image information or data output in the form of a bitstream, and a transmission unit for transmitting the information or data in the form of a file or a stream. The receiver 21 can extract / receive a bitstream from a storage medium or a network, and transmit the bitstream to the decoding unit 22.

[0064] The decoding unit 22 can decode a video / image by performing a series of processes corresponding to the operations of the encoding unit 12, such as dequantization, inverse transformation, and prediction.

[0065] Renderer 23 can render the decoded video / images. The rendered video / images can be displayed on a monitor.

[0066] Overview of an image encoding apparatus

[0067] Figure 2 This is a schematic view illustrating an image encoding device to which embodiments of the present disclosure can be applied.

[0068] like Figure 2 As shown, the image encoding device 100 may include an image segmenter 110, a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter-frame prediction unit 180, an intra-frame prediction unit 185, and an entropy encoder 190. The inter-frame prediction unit 180 and the intra-frame prediction unit 185 may be collectively referred to as "prediction units". The transformer 120, quantizer 130, dequantizer 140, and inverse transformer 150 may be included in a residual processor. The residual processor may further include a subtractor 115.

[0069] In some implementations, all or at least some of the components constituting the image encoding device 100 may be configured by a single hardware component (e.g., an encoder or a processor). Furthermore, the memory 170 may include a decoded image buffer (DPB) and may be configured by a digital storage medium.

[0070] The image partitioner 110 can partition an input image (or picture or frame) input to the image encoding apparatus 100 into one or more processing units. For example, the processing units can be referred to as coding units (CUs). The coding units can be obtained by recursively partitioning a coding tree unit (CTU) or a largest coding unit (LCU) according to a quad-tree-binary-tree-triple-tree (QT / BT / TT) structure. For example, one coding unit can be partitioned into coding units of a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a triple-tree structure. For the partitioning of the coding units, the quad-tree structure can be applied first, and then the binary-tree structure and / or the triple-tree structure can be applied. The encoding process of the disclosure can be performed based on a final coding unit that is no longer partitioned. A largest coding unit can be used as the final coding unit, or a coding unit of a deeper depth obtained by partitioning the largest coding unit can be used as the final coding unit. Here, the encoding process can include prediction, transform, and reconstruction processes that will be described later. As another example, the processing unit of the encoding process can be a prediction unit (PU) or a transform unit (TU). The prediction unit and the transform unit can be divided or partitioned from the final coding unit. The prediction unit can be a unit of sample prediction, and the transform unit can be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0071] The prediction unit (inter prediction unit 180 or intra prediction unit 185) can perform prediction on a block (current block) to be processed and generate a prediction block containing predicted samples for the current block. The prediction unit can determine whether to apply intra prediction or inter prediction based on the current block or coding unit. The prediction unit can generate various information related to the prediction of the current block and transmit the generated information to the entropy encoder 190. The information about the prediction can be encoded in the entropy encoder 190 and output in the form of a bitstream.

[0072] The intra prediction unit (intra predictor) 185 can predict the current block by referring to samples in the current picture. The referred samples can be located in a neighboring region of the current block, or can be located at a distant position according to an intra prediction mode and / or an intra prediction technique. The intra prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, a DC mode and a planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the level of detail of the prediction direction. However, this is merely an example, and more or less directional prediction modes can be used according to the setting. The intra prediction unit 185 can determine a prediction mode applied to the current block by using a prediction mode applied to a neighboring block.

[0073] The inter prediction unit (inter predictor) 180 can derive a prediction block of the current block based on a reference block (a reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of a block, a sub-block, or a sample based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in a reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block can be the same or different. The temporal neighboring block can be referred to as a collocated reference block, a collocated coding unit (colCU), or the like. The reference picture containing the temporal neighboring block can be referred to as a collocated picture (colPic). For example, the inter prediction unit 180 can construct a motion information candidate list based on the neighboring blocks and generate information indicating which candidate to use to derive the motion vector and / or the reference picture index of the current block. The inter prediction can be performed based on various prediction modes. For example, in the case of a skip mode and a merge mode, the inter prediction unit 180 can use the motion information of the neighboring blocks as the motion information of the current block. In the case of the skip mode, unlike the merge mode, a residual signal can not be transmitted. In the case of a motion vector prediction (MVP) mode, the motion vector of the neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be signaled by encoding a motion vector difference and an indicator for the motion vector predictor. The motion vector difference can refer to a difference between the motion vector of the current block and the motion vector predictor.

[0074] The prediction unit can generate a prediction signal based on various prediction methods and prediction techniques described below. For example, the prediction unit can not only apply intra prediction or inter prediction for prediction of the current block, but also simultaneously apply both intra prediction and inter prediction. The prediction method of simultaneously applying both intra prediction and inter prediction for prediction of the current block can be referred to as combined inter-intra prediction (CIIP). In addition, the prediction unit can perform intra block copy (IBC) for prediction of the current block. Intra block copy can be used for content image / video coding, i.e., screen content coding (SCC), of, for example, a game. IBC is a method of predicting a current picture using a previously reconstructed reference block located at a predetermined distance from the current block in the current picture. When IBC is applied, the position of the reference block in the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction in the current picture, but since the reference block is derived within the current picture, it can be performed in a manner similar to inter prediction. That is, IBC can use at least one of the inter prediction techniques described in the present disclosure.

[0075] The prediction signal generated by the prediction unit can be used to generate a reconstructed signal or to generate a residual signal. The subtractor 115 can generate a residual signal (a residual block or a residual sample array) by subtracting the prediction signal (a prediction block or a prediction sample array) output from the prediction unit from the input image signal (an original block or an original sample array). The generated residual signal can be transmitted to the transformer 120.

[0076] The transformer 120 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, the GBT refers to a transform obtained from a graph when relationship information between pixels is represented by the graph. The CNT refers to a transform obtained based on a prediction signal generated using all previous reconstructed pixels. In addition, the transform process can be applied to a square pixel block having the same size, or can be applied to a non-square variable size block.

[0077] The quantizer 130 can quantize the transform coefficients and transmit them to the entropy encoder 190. The entropy encoder 190 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 130 can rearrange the quantized transform coefficients of the block type into a one-dimensional vector form based on a coefficient scan order, and generate information about the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients.

[0078] The entropy encoder 190 can perform various encoding methods, such as exponential Golomb encoding, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), and the like. In addition to the quantized transform coefficients, the entropy encoder 190 can also encode information (e.g., values of syntax elements, etc.) required for video / image reconstruction together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in a network abstraction layer (NAL) unit. The video / image information can also include information about various parameter sets, such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information can also include general constraint information. The information signaled in the disclosure, the transmitted information, and / or the syntax elements can be encoded through the above-described encoding process and included in the bitstream.

[0079] The bitstream can be transmitted through a network or can be stored in a digital storage medium. The network can include a broadcasting network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting a signal output from the entropy encoder 190 and / or a storage unit (not shown) for storing the signal can be included as an internal / external element of the image encoding apparatus 100. Alternatively, the transmitter can be provided as a component of the entropy encoder 190.

[0080] The quantized transform coefficients output from the quantizer 130 can be used to generate a residual signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients through the dequantizer 140 and the inverse transformer 150.

[0081] The adder 155 adds the reconstructed residual signal to a prediction signal output from the inter prediction unit 180 or the intra prediction unit 185 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). If there is no residual for a block to be processed, such as the case where a skip mode is applied, the prediction block can be used as the reconstructed block. The adder 155 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra prediction of a next block to be processed in the current picture, and can be used for inter prediction of a next picture through filtering as described below.

[0082] The filter 160 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 160 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. As will be described later in the description of each filtering method, the filter 160 can generate various information related to filtering and transmit the generated information to the entropy encoder 190. The information related to filtering can be encoded by the entropy encoder 190 and output in the form of a bitstream.

[0083] The modified reconstructed picture transmitted to the memory 170 can be used as a reference picture in the inter prediction unit 180. When inter prediction is applied through the image encoding apparatus 100, prediction mismatch between the image encoding apparatus 100 and the image decoding apparatus can be avoided, and coding efficiency can be improved.

[0084] The DPB of the memory 170 can store the modified reconstructed picture to be used as a reference picture in the inter prediction unit 180. The memory 170 can store motion information of blocks for which motion information has been derived (or encoded) in the current picture and / or motion information of blocks in the reconstructed picture. The stored motion information can be transmitted to the inter prediction unit 180 and used as motion information of spatial neighboring blocks or motion information of temporal neighboring blocks. The memory 170 can store reconstructed samples of reconstructed blocks in the current picture and can transfer the reconstructed samples to the intra prediction unit 185.

[0085] Overview of an image decoding apparatus

[0086] Figure 3 FIG. 1 is a view schematically illustrating an image encoding apparatus to which embodiments of the disclosure can be applied.

[0087] As Figure 3 shown, the image decoding apparatus 200 can include an entropy decoder 210, a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter predictor 260, and an intra prediction unit 265. The inter predictor (inter prediction unit) 260 and the intra predictor (intra prediction unit) 265 can be collectively referred to as a "prediction unit (predictor)". The dequantizer 220 and the inverse transformer 230 can be included in a residual processor.

[0088] According to embodiments, all or at least some of the plurality of components constituting the image decoding apparatus 200 can be configured by hardware components (e.g., decoders or processors). Also, the memory 170 can include a decoded picture buffer (DPB) or can be configured by a digital storage medium.

[0089] The image decoding apparatus 200 that has received a bitstream containing video / image information can reconstruct a picture by performing a process corresponding to a process performed by the image encoding apparatus 100. Figure 2 For example, the image decoding apparatus 200 can perform decoding using a processing unit applied in the image encoding apparatus. Accordingly, the processing unit for decoding can be, for example, a coding unit. The coding unit can be obtained by partitioning a coding tree unit or a largest coding unit. The reconstructed picture signal decoded and output by the image decoding apparatus 200 can be reproduced by a reproduction apparatus (not shown).

[0090] The image decoding apparatus 200 can receive a bitstream from Figure 2The signal outputted in the form of a bitstream by the image encoding apparatus can be decoded by the entropy decoder 210. For example, the entropy decoder 210 can parse the bitstream to derive information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information can further include information on various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information can further include general constraint information. The image decoding apparatus can further decode the image based on the information on the parameter sets and / or the general constraint information. The signaling information / received information and / or the syntax elements described in the disclosure can be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 210 decodes information in the bitstream based on an encoding method such as exponential Golomb coding, CAVLC, or CABAC, and outputs values of syntax elements required for image reconstruction and quantized values of transform coefficients for a residual. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using decoded information of a neighboring block, a decoding target block, or a symbol / bin decoded at a previous stage, and a decoding target syntax element, predict a probability of occurrence of the bin according to the determined context model, perform arithmetic decoding on the bin, and generate a symbol corresponding to a value of each syntax element. In this case, the CABAC entropy decoding method can update the context model of the next symbol / bin using information of the decoded symbol / bin after determining the context model. Information related to prediction among the information decoded by the entropy decoder 210 can be provided to the prediction unit (inter-predictor 260 and intra-predictor 265), and the residual (i.e., quantized transform coefficients and related parameter information) on which the entropy decoding is performed in the entropy decoder 210 can be input to the dequantizer 220. In addition, information related to filtering among the information decoded by the entropy decoder 210 can be provided to the filter 240. Meanwhile, a receiver (not shown) for receiving a signal outputted from the image encoding apparatus can be further configured as an internal / external element of the image decoding apparatus 200, or the receiver can be a component of the entropy decoder 210.

[0091] Meanwhile, the image decoding apparatus according to the disclosure can be referred to as a video / image / picture decoding apparatus. The image decoding apparatus can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 210. The sample decoder can include at least one of the dequantizer 220, the inverse transformer 230, the adder 235, the filter 240, the memory 250, the inter-predictor 260, or the intra-predictor 265.

[0092] The dequantizer 220 can dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 220 can rearrange the quantized transform coefficients in the form of a two-dimensional block. In this case, the rearrangement can be performed based on a coefficient scan order performed in the image encoding apparatus. The dequantizer 220 can obtain the transform coefficients by performing dequantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step length information).

[0093] The inverse transformer 230 can perform inverse transform on the transform coefficients to obtain a residual signal (a residual block, a residual sample array).

[0094] The prediction unit can perform prediction on the current block and generate a prediction block including prediction samples for the current block. The prediction unit can determine whether to apply intra prediction or inter prediction to the current block based on information about prediction output from the entropy decoder 210, and can determine a specific intra / inter prediction mode (prediction technique).

[0095] The prediction unit can generate a prediction signal based on various prediction methods (techniques) to be described below, as is the case with the prediction unit of the image encoding apparatus 100.

[0096] The intra predictor 265 can predict the current block by referring to samples in the current picture. The description about the intra prediction unit 185 is equally applicable to the intra prediction unit 265.

[0097] The inter predictor 260 can derive a prediction block of the current block based on a reference block (a reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of a block, a sub-block, or a sample based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter predictor 260 can configure a motion information candidate list based on neighboring blocks and derive a motion vector and / or a reference picture index of the current block based on received candidate selection information. Inter prediction can be performed based on various prediction modes, and information about prediction can include information indicating an inter prediction mode of the current block.

[0098] The adder 235 can generate a reconstructed signal (a reconstructed picture, a reconstructed block, a reconstructed sample array) by adding the obtained residual signal to a prediction signal (a prediction block, a prediction sample array) output from the prediction unit (including the inter-predictor 260 and / or the intra-predictor 265). If there is no residual for a block to be processed, such as in the case of applying a skip mode, the prediction block can be used as the reconstructed block. The description regarding the adder 155 is equally applicable to the adder 235. The adder 235 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-prediction of a next block to be processed in the current picture, and can be used for inter-prediction of a next picture by filtering as described below.

[0099] The filter 240 can improve subjective / objective picture quality by applying filtering to the reconstructed signal. For example, the filter 240 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 250, specifically, in the DPB of the memory 250. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc.

[0100] The (modified) reconstructed picture stored in the DPB of the memory 250 can be used as a reference picture in the inter-predictor 260. The memory 250 can store motion information of a block for which motion information has been derived (or decoded) in the current picture and / or motion information of a block in a picture that has been reconstructed. The stored motion information can be transmitted to the inter-predictor 260 to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory 250 can store reconstructed samples of a reconstructed block in the current picture, and transfer the reconstructed samples to the intra-predictor 265.

[0101] In the present disclosure, the embodiments described in the filter 160, the inter-predictor 180, and the intra-predictor 185 of the image encoding apparatus 100 can be equally or correspondingly applied to the filter 240, the inter-predictor 260, and the intra-predictor 265 of the image decoding apparatus 200.

[0102] Neural network post filter characteristic (NNPFC)

[0103] The combination of Table 1 to Table 3 represents the NNPFC syntax structure.

[0104] [Table 1]

[0105]

[0106] [Table 2]

[0107]

[0108] [Table 3]

[0109]

[0110] The NNPFC syntax structure of Table 1 to Table 3 can be signaled in the form of a supplemental enhancement information (SEI) message. The SEI message that signals the NNPFC syntax structure of Table 1 to Table 3 can be referred to as a neural network post filter characteristics (NNPFC) SEI message.

[0111] The neural network post filter characteristics (NNPFC) SEI message specifies a neural network that can be used as a post-processing filter. The neural network post filter activation (NNPFA) SEI message instructs to use the specified neural network post-processing filter (NNPF) for a particular picture. Here, “post-processing filter” and “post filter” can have the same meaning.

[0112] The following variables need to be defined for the use of this SEI message:

[0113] - The input picture width and height in luma samples, denoted CroppedWidth and CroppedHeight, respectively, in this document.

[0114] - The luma sample array CroppedYPic[idx] and the chroma sample arrays CroppedCbPic[idx] and CroppedCrPic[idx] (when present) of the input pictures, with index idx in the range of 0 to numInputPics - 1, inclusive, are used as input to the NNPF.

[0115] - BitDepthY can denote the bit depth of the luma sample array of the input pictures.

[0116] - BitDepthC can denote the bit depth of the chroma sample array (if any) of the input pictures.

[0117] - ChromaFormatIdc can denote the chroma format identifier.

[0118] - When the value of nnpfc_auxiliary_inp_idc is equal to 1, the filter strength control value StrengthControlVal shall be a real number in the range of 0 to 1, inclusive.

[0119] An input picture with index 0 can correspond to a picture that activates the NNPF defined by this NNPFC SEI message through the NNPFA SEI message. An input picture with index i in the range of 1 to numInputPics - 1, inclusive, can precede the input picture with index i - 1 in output order.

[0120] If nnpfc_purpose & 0x08 is not equal to 0 and the input picture with index 0 is associated with a frame packing arrangement SEI message with fp_arrangement_type equal to 5, all input pictures can be associated with a frame packing arrangement SEI message with fp_arrangement_type equal to 5 and have the same value as fp_current_frame_is_frame0_flag.

[0121] More than one NNPFC SEI message can exist for the same picture. If more than one NNPFC SEI message with different nnpfc_id values for the same picture exists or is active, the more than one NNPFC SEI message can have the same or different nnpfc_purpose and nnpfc_mode_idc values.

[0122] nnpfc_purpose indicates the purpose of the NNPF as specified in Table 4. In the bitstream, the value of nnpfc_purpose shall be in the range of 0 to 63, inclusive. Values of nnpfc_purpose from 64 to 65535, inclusive, are reserved for future use. Decoders shall ignore NNPFC SEI messages with nnpfc_purpose in the range of 64 to 65535, inclusive. If a value of nnpfc_purpose is reserved for future use, the syntax elements of this SEI message can extend the existing syntax elements, provided that nnpfc_purpose is equal to the corresponding value. If ChromaFormatldc is equal to 3, nnpfc_purpose & 0x02 shall be equal to 0. If ChromaFormatldc or nnpfc_purpose & 0x02 is not equal to 0, nnpfc_purpose & 0x20 shall be equal to 0.

[0123] [Table 4]

[0124]

[0125] nnpfc_id can contain an identification number that can be used to identify the NNPF. In the bitstream, the value of nnpfc_id shall be in the range of 0 to 232 - 2 (inclusive) 32 - 2) inclusive. Values of nnpfc_id from 256 to 511 (inclusive) and 2 31 to 2 32 - 2 (inclusive) 31 and 2 32 - 2) can be reserved for future use. Decoders shall ignore an NNPFC SEI message when the nnpfc_id is in the range of 256 to 511 (inclusive) or in the range of 2 31 to 2 32 - 2 (inclusive) 31 and 2 32 - 2) inclusive.

[0126] The following applies when the -NNPFC SEI message is the first NNPFC SEI message with a particular nnpfc_id value within the current CLVS in decoding order:

[0127] - The SEI message specifies the reference NNPF.

[0128] - The SEI message applies to the current decoded picture of the current layer and all subsequent decoded pictures in output order until the end of the current CLVS.

[0129] The NNPFC SEI message can be a repetition of a previous NNPFC SEI message within the current CLVS in decoding order, and the following semantics can apply under the assumption that the SEI message is the only NNPFC SEI message within the current CLVS with the same content.

[0130] nnpfc_mode_idc equal to 0 can specify that the SEI message can contain a bitstream representing the reference NNPF, or can represent an update relative to the reference NNPF with the same nnpfc_id value.

[0131] When the NNPFC SEI message is the first NNPFC SEI message with a particular nnpfc_id value within the current CLVS in decoding order, nnpfc_mode_idc equal to 1 can specify that the reference NNPF associated with the nnpfc_id value is a neural network, and the neural network can be a neural network identified by the format using nnpfc_tag_uri, identified by the URI represented by nnpfc_uri.

[0132] If the NNPFC SEI message is neither the first NNPFC SEI message with a particular nnpfc id value in the current CLVS in decoding order, nor a repetition of the first NNPFC SEI message, nnpfc mode idc equal to 1 can specify that the update is to be used with respect to the reference NNPF with the same nnpfc id value using the format identified by nnpfc tag uri, the URI represented by nnpfc uri.

[0133] The value of nnpfc mode idc shall be constrained to be in the range of 0 to 1, inclusive, in the bitstream. Values of nnpfc mode idc in the range of 2 to 255, inclusive, can be reserved for future use and can not be present in the bitstream. Decoders shall ignore NNPFC SEI messages with nnpfc mode idc in the range of 2 to 255, inclusive. Values of nnpfc mode idc greater than 255 shall not be present in the bitstream and shall not be reserved for future use.

[0134] If the SEI message is the first NNPFC SEI message with a particular nnpfc id value in the current CLVS in decoding order, the NNPF post-processing filter (PostProcessingFilter()) can be assigned to be the same as the reference NNPF.

[0135] If the SEI message is not the first NNPFC SEI message with a particular nnpfc id value in the current CLVS in decoding order, and is not a repetition of the first NNPFC SEI message, the NNPF post-processing filter (PostProcessingFilter()) can be obtained by applying the update defined by the SEI message to the reference NNPF.

[0136] The updates are not cumulative; rather, each update is applied to the reference NNPF, which is the NNPF specified by the first NNPFC SEI message with a particular nnpfc id value in the current CLVS in decoding order.

[0137] In the bitstream, nnpfc reserved zero bit a shall be constrained to be equal to 0. Decoders shall be constrained to ignore NNPFC SEI messages with nnpfc reserved zero bit a not equal to 0.

[0138] The nnpfc_tag_uri can contain a tag URI with syntax and semantics as specified in IETF RFC 4151 that identifies the neural network used as the reference NNPFC or an update relative to the reference NNPFC with the same nnpfc_id value as specified by the nnpfc_uri. The nnpfc_tag_uri can uniquely identify the format of the neural network data specified by the nnpfc_uri without the need for a central registry. The nnpfc_tag_uri equal to "tag:iso.org,2023:15938-17" can specify that the neural network data identified by the nnpfc_uri conforms to ISO / IEC 15938-17.

[0139] The nnpfc_uri can contain a URI with syntax and semantics as specified in IETF Internet Standard 66 that identifies the neural network used as the reference NNPFC or an update relative to the reference NNPFC with the same nnpfc_id value.

[0140] The nnpfc_property_present_flag equal to 1 can specify that syntax elements related to filter usage, input format, output format, and complexity are present. The nnpfc_property_present_flag equal to 0 can specify that syntax elements related to filter usage, input format, output format, and complexity are not present. If the SEI message is the first NNPFC SEI message with the particular nnpfc_id value in the current CLVS in decoding order, the value of nnpfc_property_present_flag shall be constrained to be equal to 1. If the value of nnpfc_property_present_flag is equal to 0, all syntax elements that exist only when the value of nnpfc_property_present_flag is equal to 1 and for which no inferred value is specified can be inferred to be equal to the corresponding syntax elements in the NNPFC SEI message of the reference NNPFC that contains the update provided by this SEI message.

[0141] The nnpfc_base_flag equal to 1 can specify that this SEI message represents a reference NNPFC. The nnpfc_base_flag value equal to 0 can specify that this SEI message represents an update related to a reference NNPFC. If nnpfc_base_flag is not present, the nnpfc_base_flag value can be inferred to be 0.

[0142] The value of nnpfc_base_flag is subject to the following constraint:

[0143] - When the NNPFC SEI message is the first NNPFC SEI message with a particular nnpfc id value in the current CLVS in decoding order, the value of nnpfc base flag shall be equal to 1.

[0144] - When the NNPFC SEI message nnpfcB is not the first NNPFC SEI message with a particular nnpfc id value in the current CLVS in decoding order, and the value of nnpfc base flag is equal to 1, the NNPFC SEI message shall be a repetition of the first NNPFC SEI message nnpfcA with the same nnpfc id value in decoding order, i.e. the payload content of nnpfcB shall be the same as the payload content of nnpfcA.

[0145] If the NNPFC SEI message is not the first NNPFC SEI message with a particular nnpfc id value in the current CLVS in decoding order, and is not a repetition of the first NNPFC SEI message with that particular nnpfc id value, the following applies:

[0146] - An SEI message can define updates with respect to the reference NNPFC preceding them in decoding order and having the same nnpfc id value.

[0147] - An SEI message is only associated with the current reconstructed picture of the current layer and all subsequent reconstructed pictures in output order up to the end of the current CLVS or the next reconstructed picture within the current CLVS, and with subsequent NNPFC SEI messages with earlier values of a particular nnpfc id value in decoding order within the current CLVS.

[0148] When the NNPFC SEI message nnpfcCurr is not the first NNPFC SEI message with a particular nnpfc id value in decoding order within the current CLVS, is not a repetition of the first NNPFC SEI message with that particular nnpfc id value (i.e. the value of nnpfc base flag is equal to 0), and the value of nnpfc property present flag is equal to 1, the following constraint applies:

[0149] - The value of nnpfc purpose in this NNPFC SEI message shall be the same as the value of nnpfc purpose in the first NNPFC SEI message with that particular nnpfc id value in decoding order within the current CLVS.

[0150] - The value of the syntax element in the NNPFC SEI message that is located in decoding order after nnpfc_property_present_flag and before nnpfc_complexity_info_present_flag shall be the same as the value of the corresponding syntax element in the first NNPFC SEI message with that particular nnpfc id value in decoding order within the current CLVS.

[0151] - nnpfc_complexity_info_present_flag is equal to 0; or nnpfc_complexity_info_present_flag is equal to 1 in the first NNPFC SEI message with that particular nnpfc id value in decoding order within the current CLVS, and all of the following apply:

[0152] (1) nnpfc_parameter_type_idc in nnpfcCurr shall be equal to nnpfc_parameter_type_idc in nnpfcBase.

[0153] (2) When nnpfc_log2_parameter_bit_length_minus3 is present in nnpfcCurr, its value shall be less than or equal to nnpfc_log2_parameter_bit_length_minus3 in nnpfcBase.

[0154] (3) If nnpfc_num_parameters_idc in nnpfcBase is equal to 0, nnpfc_num_parameters_idc in nnpfcCurr shall be equal to 0.

[0155] (4) Otherwise (nnpfc_num_parameters_idc in nnpfcBase is greater than 0), nnpfc_num_parameters_idc in nnpfcCurr shall be greater than 0 and less than or equal to nnpfc_num_parameters_idc in nnpfcBase.

[0156] (5) If nnpfc_num_kmac_operations_idc in nnpfcBase is equal to 0, nnpfc_num_kmac_operations_idc in nnpfcCurr shall be equal to 0.

[0157] (6) Otherwise (nnpfc num kmac operations idc in nnpfcBase is greater than 0), nnpfc num kmac operations idc in nnpfcCurr shall be greater than 0 and less than or equal to nnpfc num kmac operations idc in nnpfcBase.

[0158] (7) If nnpfc total kilobyte size in nnpfcBase is equal to 0, nnpfc total kilobyte size in nnpfcCurr shall be equal to 0.

[0159] (8) Otherwise (nnpfc total kilobyte size in nnpfcBase is greater than 0), nnpfc total kilobyte size in nnpfcCurr shall be greater than 0 and less than or equal to nnpfc total kilobyte size in nnpfcBase.

[0160] When nnpfc purpose & 0x02 is not equal to 0, nnpfc out sub c flag can specify the values of variables outSubWidthC and outSubHeightC. nnpfc out sub c flag equal to 1 can specify outSubWidthC equal to 1 and outSubHeightC equal to 1. nnpfc out sub c flag equal to 0 can specify outSubWidthC equal to 2 and outSubHeightC equal to 1. When ChromaFormatldc is equal to 2 and nnpfc out sub c flag is present, the value of nnpfc out sub c flag shall be equal to 1.

[0161] When nnpfc_purpose & 0x20 is not equal to 0, nnpfc_out_colour_format_idc can specify the colour format of the NNPFC output, and thus the values of the variables outSubWidthC and outSubHeightC. nnpfc_out_colour_format_idc equal to 1 can specify that the colour format of the NNPFC output is 4:2:0 format, and both outSubWidthC and outSubHeightC are equal to 2. nnpfc_out_colour_format_idc equal to 2 can specify that the colour format of the NNPFC output is 4:2:2 format, and outSubWidthC is equal to 2, outSubHeightC is equal to 1. nnpfc_out_colour_format_idc equal to 3 can specify that the colour format of the NNPFC output is 4:2:4 format, and both outSubWidthC and outSubHeightC are equal to 1. The value of nnpfc_out_colour_format_idc can not be restricted to be equal to 0.

[0162] When both nnpfc_purpose & 0x02 and nnpfc_purpose & 0x20 are equal to 0, outSubWidthC and outSubHeightC can be inferred to be equal to SubWidthC and SubHeightC, respectively.

[0163] nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples can specify the width and height, respectively, of the luma sample array of the picture resulting from applying the NNPF identified by nnpfc_id to the cropped decoded output picture. If nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples are not present, they can be inferred to be equal to CroppedWidth and CroppedHeight, respectively. The value of nnpfc_pic_width_in_luma_samples shall be in the range of CroppedWidth to CroppedWidth x 16 - 1, inclusive. The value of nnpfc_pic_height_in_luma_samples shall be in the range of CroppedHeight to CroppedHeight x 16 - 1, inclusive.

[0164] nnpfc_num_input_pics_minus1 + 1 can specify the number of decoded output pictures used as input to the NNPF. The value of nnpfc_num_input_pics_minus1 can not be constrained to be in the range of 0 to 63, inclusive.

[0165] nnpfc_interpolated_pics[ i ] can specify the number of interpolated pictures generated by the NNPF between the i-th picture and the (i+1)-th picture used as input to the NNPF. The value of nnpfc_interpolated_pics[ i ] can not be constrained to be in the range of 0 to 63, inclusive. The value of nnpfc_interpolated_pics[ i ] can not be constrained to be greater than 0 for at least one i value in the range of 0 to nnpfc_num_input_pics_minus1 - 1, inclusive.

[0166] nnpfc_input_pic_output_flag[ i ] equal to 1 can specify that the NNPF generates a corresponding output picture for the i-th input picture. nnpfc_input_pic_output_flag[ i ] equal to 0 can specify that the NNPF does not generate a corresponding output picture for the i-th input picture.

[0167] The variable numInputPics indicating the number of pictures used as input to the NNPF, and the variable numOutputPics indicating the total number of pictures generated as a result of the NNPF can be derived as shown in Table 5.

[0168] [Table 5]

[0169]

[0170] nnpfc_component_last_flag equal to 1 can specify that the last dimension in the input tensor inputTensor of the NNPF and the output tensor outputTensor resulting from the NNPF is used for the current channel. nnpfc_component_last_flag equal to 0 can specify that the third dimension in the input tensor inputTensor of the NNPF and the output tensor outputTensor resulting from the NNPF is used for the current channel.

[0171] The first dimension in the input tensor and the output tensor can be used for batch index, which is a common practice in some neural network frameworks. Although the formula in the semantics of this SEI message uses the batch size corresponding to the batch index equal to 0, the batch size used for the neural network inference input is determined by the post-processing implementation.

[0172] For example, when nnpfc_inp_order_idc is equal to 3 and nnpfc_auxiliary_inp_idc is equal to 1, there are 7 channels in the input tensor, including 4 luma matrices, 2 chroma matrices, and 1 auxiliary input matrix. In this case, the DeriveInputTensors() process derives each of the 7 channels of the input tensor one by one, and when a particular channel among these channels is being processed, that channel can be referred to as the current channel during the processing.

[0173] nnpfc_inp_format_idc can specify the method of converting the sample values of the cropped decoded output picture into the input values of the NNPF. When nnpfc_inp_format_idc is equal to 0, the input values of the NNPF are real numbers, and the functions InpY() and InpC() can be specified as shown in Equation 1.

[0174] [Equation 1]

[0175] InpY(x) = x ÷ ((1 « BitDepth Y ) - 1)

[0176] InpC(x) = x ÷ ((1 « BitDepth C ) - 1)

[0177] When nnpfc_inp_format_idc is equal to 1, the input values of the NNPF are unsigned integers, and the functions InpY() and InpC() are specified as indicated in Table 6.

[0178] [Table 6]

[0179]

[0180] The variable inpTensorBitDepthY can be derived from the syntax element nnpfc_inp_tensor_luma_bitdepth_minus8 as specified below. The variable inpTensorBitDepthC can be derived from the syntax element nnpfc_inp_tensor_chroma_bitdepth_minus8 as specified below.

[0181] Values of nnpfc_inp_format_idc greater than 1 are reserved for future use and can not be present in the bitstream. Decoders must ignore NNPFC SEI messages that contain a reserved value of nnpfc_inp_format_idc.

[0182] nnpfc_inp_tensor_luma_bitdepth_minus8 + 8 can specify the bit depth of the luma sample values in the input integer tensor. The value of inpTensorBitDepthY is derived as indicated in Equation 2.

[0183] [Equation 2]

[0184] inpTensorBitDepth Y = nnpfc_inp_tensor_luma_bitdepth_minus8 + 8

[0185] Bitstream conformance requires that the value of nnpfc_inp_tensor_luma_bitdepth_minus8 can be constrained to be in the range of 0 to 24, inclusive.

[0186] nnpfc_inp_tensor_chroma_bitdepth_minus8 + 8 can specify the bit depth of the chroma sample values in the input integer tensor. The value of inpTensorBitDepthC is derived as indicated in Equation 3.

[0187] [Equation 3]

[0188] inpTensorBitDepth C= nnpfc inp tensor luma bitdepth minus8 + 8

[0189] Bitstream conformance requires that the value of nnpfc inp tensor luma bitdepth minus8 can be restricted to the range of 0 to 24, inclusive.

[0190] nnpfc inp order idc can specify the method to order the sample arrays of the cropped decoded output pictures as one of the NNPF input pictures.

[0191] In the bitstream, the value of nnpfc inp order idc shall be in the range of 0 to 3, inclusive. Values of nnpfc inp order idc in the range of 4 to 255, inclusive, can not be present in the bitstream. Decoders shall ignore NNPF SEI messages with nnpfc inp order idc in the range of 4 to 255, inclusive. Values of nnpfc inp order idc greater than 255 can not be present in the bitstream and are not reserved for future use.

[0192] When ChromaFormatIdc is not equal to 1, the value of nnpfc inp order idc shall not be 3.

[0193] Table 7 gives a referential description of the values of nnpfc inp order idc.

[0194] [Table 7]

[0195]

[0196]

[0197] A patch refers to a rectangular array of samples extracted from a component (e.g., luma or chroma component) of a picture.

[0198] nnpfc auxiliary inp idc greater than 0 can specify that auxiliary input data is present in the input tensors of the NNPF. nnpfc auxiliary inp idc equal to 0 can specify that no auxiliary input data is present in the input tensors. nnpfc auxiliary inp idc equal to 1 can specify that the auxiliary input data is derived by the methods shown in Table 8 to Table 10.

[0199] The value of nnpfc auxiliary inp idc shall be in the range of 0 to 1, inclusive, in the bitstream. Values of nnpfc inp order idc in the range of 2 to 255, inclusive, shall not be present in the bitstream. Decoders shall ignore NNPFC SEI messages for which nnpfc inp order idc is in the range of 2 to 255, inclusive. Values of nnpfc inp order idc greater than 255 shall not be present in the bitstream and are not reserved for future use.

[0200] When the value of nnpfc auxiliary inp idc is equal to 1, the variable strengthControlScaledVal can be derived as shown in equation 4.

[0201] [Equation 4]

[0202]

[0203] The process DeriveInputTensors() for deriving an input tensor inputTensor for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft indicating the top-left sample position of a sample patch included in the input tensor can be represented as a combination of Table 8 to Table 10.

[0204] [Table 8]

[0205]

[0206] [Table 9]

[0207]

[0208] [Table 10]

[0209]

[0210] nnpfc separate colour description present flag equal to 1 can specify that a unique combination of primary colours, transfer characteristics and matrix coefficients for pictures resulting from the NNPF is specified in the SEI message syntax structure. nnpfc separate colour description present flag equal to 0 can specify that the combination of primary colours, transfer characteristics and matrix coefficients for pictures resulting from the NNPF is the same as indicated in the VUI parameters of the CLVS.

[0211] The semantics of nnpfc_colour_primaries are the same as those defined for the vui_colour_primaries syntax element, except that:

[0212] - nnpfc_colour_primaries can specify the colour primaries of the pictures resulting from the application of the NNPF specified in the SEI message, other than the colour primaries used in the CLVS.

[0213] - If nnpfc_colour_primaries is not present in the NNPFC SEI message, its value can be inferred to be the same as the value of vui_colour_primaries.

[0214] The semantics of nnpfc_transfer_characteristics are the same as those defined for the vui_transfer_characteristics syntax element, except that:

[0215] - nnpfc_transfer_characteristics can specify the transfer characteristics of the pictures resulting from the application of the NNPF specified in the SEI message, other than the transfer characteristics used in the CLVS.

[0216] - If nnpfc_transfer_characteristics is not present in the NNPFC SEI message, its value can be inferred to be the same as the value of vui_transfer_characteristics.

[0217] The semantics of nnpfc_matrix_coeffs are the same as those specified for the vui_matrix_coeffs syntax element, except that:

[0218] - nnpfc_matrix_coeffs can specify the matrix coefficients of the pictures resulting from the application of the NNPF specified in the SEI message, other than the matrix coefficients used in the CLVS.

[0219] - If nnpfc_matrix_coeffs is not present in the NNPFC SEI message, the value of nnpfc_matrix_coeffs can be inferred to be the same as the value of vui_matrix_coeffs.

[0220] - The allowed values of nnpfc_matrix_coeffs can not be restricted to the decoded video picture chroma format indicated by the ChromaFormatldc values in the VUI parameters semantics.

[0221] - If the value of nnpfc_matrix_coeffs is equal to 0, the value of nnpfc_out_order_idc shall not be equal to 1 or 3.

[0222] nnpfc_out_format_idc equal to 0 can specify that the sample values of the NNPF output are real numbers with values in the range of 0 to 1, which are linearly mapped to unsigned integers with values in the range of 0 to (1 « bitDepth) - 1, where bitDepth is the required bit depth for subsequent post-processing or display. nnpfc_out_format_idc equal to 1 can specify that the luma sample values of the NNPF output are unsigned integers with values in the range of 0 to (1 « (nnpfc_out_tensor_luma_bitlength_minus8 + 8)) - 1, inclusive, and the chroma sample values of the NNPF output are unsigned integers with values in the range of 0 to (1 « (nnpfc_out_tensor_chroma_bitlength_minus8 + 8)) - 1, inclusive.

[0223] Values of nnpfc_out_format_idc greater than 1 are reserved for future use and can not be present in the bitstream. Decoders must ignore NNPFC SEI messages that contain reserved values of nnpfc_out_format_idc.

[0224] nnpfc_out_tensor_luma_bitdepth_minus8 + 8 can specify the bit depth of the luma sample values in the output integer tensor. The value of nnpfc_out_tensor_luma_bitdepth_minus8 shall be in the range of 0 to 24, inclusive.

[0225] nnpfc_out_tensor_chroma_bitdepth_minus8 + 8 can specify the bit depth of the chroma sample values in the output integer tensor. The value of nnpfc_out_tensor_chroma_bitdepth_minus8 shall be in the range of 0 to 24, inclusive.

[0226] If nnpfc_purpose & 0x10 is not equal to 0, the value of nnpfc_out_format_idc shall be equal to 1 and at least one of the following constraints is satisfied:

[0227] -nnpfc_out_tensor_luma_bitdepth_minus8+8 is greater than BitDepth Y

[0228] -nnpfc_out_tensor_chroma_bitdepth_minus8+8 is greater than BitDepth C

[0229] `nnpfc_out_order_idc` specifies the output order of NNPF output samples. In the bitstream, the value of `nnpfc_out_order_idc` should be in the range of 0 to 3. Values ​​of `nnpfc_out_order_idc` in the range of 4 to 255 may not exist in the bitstream. The decoder should ignore NNPFC SEI messages with `nnpfc_out_order_idc` in the range of 4 to 255. Values ​​of `nnpfc_out_order_idc` greater than 255 may not exist in the bitstream and are not reserved for future use. When the value of `nnpfc_purpose&0x02` is 0, the value of `nnpfc_out_order_idc` should not be equal to 3.

[0230] Table 11 describes the values ​​of nnpfc_out_order_idc.

[0231] [Table 11]

[0232]

[0233] Given the vertical sample coordinates cTop and horizontal sample coordinates cLeft of the top-left sample position of the patch of samples included in the input tensor, the processing of sample values ​​in the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor outputTensor, StoreOutputTensors(), can be represented as a combination of Tables 12 and 13.

[0234] [Table 12]

[0235]

[0236] [Table 13]

[0237]

[0238] nnpfc overlap can specify the number of horizontal and vertical sample overlaps for the NNPF neighboring input tensors. The value of nnpfc overlap shall be in the range of 0 to 16383, inclusive.

[0239] nnpfc_constant_patch_size_flag equal to 1 can specify that the NNPF accepts only the patch size indicated by nnpfc patch width minusl and nnpfc patch height minusl as input. nnpfc_constant_patch_size_flag equal to 0 can specify that the NNPF accepts any patch size (width inpPatchWidth, height inpPatchHeight) as input, such that the width of the extended patch (i.e., the patch plus the overlap region) (equal to inpPatchWidth + 2 x nnpfc overlap) is a positive integer multiple of nnpfc extended patch width cd delta minusl + 1 + 2 x nnpfc overlap, and the height of the extended patch (equal to inpPatchHeight + 2 x nnpfc overlap) is a positive integer multiple of nnpfc extended patch height cd delta minusl + 1 + 2 x nnpfc overlap.

[0240] When nnpfc_constant_patch_size_flag is equal to 1, nnpfc patch width minusl + 1 can specify the horizontal sample count of the required patch size for the NNPF input. The value of nnpfc patch width minusl shall be in the range of 0 to Min(32766, CroppedWidth - 1), inclusive.

[0241] When nnpfc_constant_patch_size_flag is equal to 1, nnpfc patch height minusl + 1 can specify the vertical sample count of the required patch size for the NNPF input. The value of nnpfc patch height minusl shall be in the range of 0 to Min(32766, CroppedHeight - 1), inclusive.

[0242] When nnpfc_constant_patch_size_flag is equal to 0, nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 x nnpfc_overlap can specify the greatest common divisor of all allowed extended patch width values required for the NNPF input. The value of nnpfc_extended_patch_width_cd_delta_minus1 shall be in the range of 0 to Min(32766, CroppedWidth - 1), inclusive.

[0243] When nnpfc_constant_patch_size_flag is equal to 0, nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 x nnpfc_overlap can specify the greatest common divisor of all allowed extended patch height values required for the NNPF input. The value of nnpfc_extended_patch_height_cd_delta_minus1 shall be in the range of 0 to Min(32766, CroppedHeight - 1), inclusive.

[0244] The variables inpPatchWidth and inpPatchHeight can be set to the width and height of the patch size, respectively.

[0245] If the value of nnpfc_constant_patch_size_flag is equal to 0, the following applies:

[0246] The values of inpPatchWidth and inpPatchHeight can be provided by external means or set by the post-processor itself.

[0247] The value of inpPatchWidth + 2 x nnpfc overlap shall be a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1 + 1 + 2 x nnpfc overlap and inpPatchWidth shall be less than or equal to CroppedWidth. The value of inpPatchHeight + 2 x nnpfc overlap shall be a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1 + 1 + 2 x nnpfc overlap and inpPatchHeight shall be less than or equal to CroppedHeight.

[0248] Otherwise (nnpfc_constant_patch_size_flag is equal to 1), the value of inpPatchWidth can be set to nnpfc_patch_width_minus1 + 1 and the value of inpPatchHeight can be set to nnpfc_patch_height_minus1 + 1.

[0249] The variables outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth and outPatchCHeight can be derived as shown in Table 14.

[0250] [Table 14]

[0251]

[0252] Bitstream conformance requires that outPatchWidth x CroppedWidth is equal to nnpfc_pic_width_in_luma_samples x inpPatchWidth and outPatchHeight x CroppedHeight is equal to nnpfc_pic_height_in_luma_samples x inpPatchHeight.

[0253] nnpfc_padding_type can specify the padding process employed when referring to sample positions outside the boundaries of the cropped decoded output picture, as specified in Table 15. The value of nnpfc_padding_type shall be in the range of 0 to 15, inclusive.

[0254] [Table 15]

[0255] nnpfc_padding_type Description 0 Zero padding 1 Copy padding 2 Reflect padding 3 Wrap padding 4 Fixed padding 5…15 Preserve

[0256] nnpfc_luma_padding_val can specify the luma value used for padding when the value of nnpfc_padding_type is 4.

[0257] nnpfc_cb_padding_val can specify the Cb value used for padding when the value of nnpfc_padding_type is 4.

[0258] nnpfc_cr_padding_val can specify the Cr value used for padding when the value of nnpfc_padding_type is 4.

[0259] The input to the InpSampleVal(y, x, picHeight, picWidth, CroppedPic) function is the vertical sample position y, the horizontal sample position x, the picture height picHeight, the picture width picWidth, and the sample array CroppedPic, and the function can return the SampleVal value derived as shown in Table 16.

[0260] For the input to the InpSampleVal() function, to be compatible with the input tensor rules of some inference engines, the vertical position can be listed before the horizontal position.

[0261] [Table 16]

[0262]

[0263] The process shown in Table 17 can be used in conjunction with the NNPF PostProcessingFilter() to generate filtered and / or interpolated pictures on a slice-by-slice basis. These pictures can include the Y, Cb, and Cr sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic, respectively, as indicated by nnpfc_out_order_idc:

[0264] [Table 17]

[0265]

[0266] The order in which the pictures stored in the output tensor can be the output order, and the output order generated by applying the NNPF to the output order can be interpreted as an output order that does not conflict with the output order of the input pictures.

[0267] nnpfc_complexity_info_present_flag equal to 1 can specify that one or more syntax elements indicating the complexity of the NNPF associated with nnpfc id are present. nnpfc_complexity_info_present_flag equal to 0 can specify that no syntax elements indicating the complexity of the NNPF associated with nnpfc id are present.

[0268] nnpfc_parameter_type_idc equal to 0 can specify that the neural network uses only integer parameters. nnpfc_parameter_type_flag equal to 1 can specify that the neural network can use floating-point or integer parameters. nnpfc_parameter_type_idc equal to 2 can specify that the neural network uses only binary parameters. nnpfc_parameter_type_idc equal to 3 can be reserved for future use and can not be present in the bitstream. Decoders shall ignore NNPFC SEI messages with nnpfc_parameter_type_idc equal to 3.

[0269] nnpfc_log2_parameter_bit_length_minus3 equal to 0, 1, 2, and 3 can specify that the neural network does not use parameters with bit length greater than 8, 16, 32, and 64, respectively. When nnpfc_parameter_type_idc is present and nnpfc_log2_parameter_bit_length_minus3 is not present, the neural network can not use parameters with bit length greater than 1.

[0270] nnpfc_num_parameters_idc can specify the maximum number of neural network parameters of the NNPF in units of powers of 2048. nnpfc_num_parameters_idc equal to 0 can specify that the maximum number of neural network parameters is unknown. The value of nnpfc_num_parameters_idc shall be in the range of 0 to 52, inclusive. Values of nnpfc_num_parameters_idc greater than 52 are reserved for future use and can not be present in the bitstream. Decoders shall ignore NNPFC SEI messages with nnpfc_num_parameters_idc greater than 52.

[0271] If the value of nnpfc_num_parameters_idc is greater than 0, the variable maxNumParameters is derived as indicated in Equation 5.

[0272] [Formula 5]

[0273] maxNumParameters = (2048 « nnpfc_num_parameters_idc) - 1

[0274] The requirement for bitstream conformance is that the number of neural network parameters of a NNPF shall be constrained to be less than or equal to maxNumParameters.

[0275] nnpfc_num_kmac_operations_idc greater than 0 can specify that the maximum number of multiply-accumulate operations per sample of the NNPF is less than or equal to nnpfc_num_kmac_operations_idc x 1000. nnpfc_num_kmac_operations_idc equal to 0 can specify that the maximum number of multiply-accumulate operations of the network is unknown. The value of nnpfc_num_kmac_operations_idc shall be in the range of 0 to 2 32 - 2, inclusive. 32 - 2, inclusive.

[0276] nnpfc_total_kilobyte_size greater than 0 can specify the total size in kilobytes required to store the uncompressed parameters of the neural network. The total size in bits is the number equal to or greater than the sum of bits used to store each parameter. nnpfc_total_kilobyte_size is the total size in bits rounded up to the nearest integer divided by 8000. nnpfc_total_kilobyte_size equal to 0 can specify that the total size required to store the parameters of the neural network is unknown. The value of nnpfc_total_kilobyte_size shall be in the range of 0 to 2 32 - 2, inclusive. 32 - 2, inclusive.

[0277] In the bitstream, nnpfc_reserved_zero_bit_b shall be equal to 0. Decoders shall ignore NNPFC SEI messages for which nnpfc_reserved_zero_bit_b is not equal to 0.

[0278] nnpfc_payload_byte[ i ] can contain the i-th byte of the bitstream. The sequence of bytes nnpfc_payload_byte[ i ] for all i values that exist shall be a conforming bitstream to ISO / IEC 15938-17.

[0279] Neural network post filter activation (NNFPA)

[0280] The syntax structure of NNFPA is shown in Table 18.

[0281] [Table 18]

[0282]

[0283] The NNPFA syntax structure of Table 18 can be signaled in the form of an SEI message. The SEI message signaling the NNPFA syntax structure of Table 18 can be referred to as an NNPFA SEI message.

[0284] The neural network post-filter activation (NNPFA) SEI message can activate or deactivate the possible use of a target neural network post-filter (NNPF) identified by nnpfa_target_id for post-processing filtering of a set of pictures. For a particular picture for which the NNPF is activated, the target NNPF can be the NNPF specified by the last NNPFC SEI message with nnpfc_id equal to nnpfa_target_id, which can precede, in decoding order, the first VCL NAL unit of the current picture and can not correspond to a repetition of an NNPFC SEI message containing the underlying NNPF.

[0285] There can be several NNPFA SEI messages for the same picture, e.g., when the NNPF is used for different purposes or for filtering different color components.

[0286] nnpfa_target_id can specify the target NNPF, which is specified by one or more NNPFC SEI messages related to the current picture and having nnpfc_id equal to nnpfa_target_id.

[0287] The value of nnpfa_target_id shall be in the range of 0 to 2 32 - 2, inclusive. The nnpfa_target_id values in the range of 256 to 511, inclusive, and 2 32 - 2 can be reserved for future use. Decoders shall ignore NNPFA SEI messages with nnpfa_target_id in the range of 256 to 511, inclusive, or 2 31 - 2, inclusive. The nnpfa_target_id values in the range of 256 to 511, inclusive, and 2 32 - 2 can be reserved for future use. Decoders shall ignore NNPFA SEI messages with nnpfa_target_id in the range of 256 to 511, inclusive, or 2 31 - 2, inclusive. The nnpfa_target_id values in the range of 256 to 511, inclusive, and 2 32 - 2 can be reserved for future use. Decoders shall ignore NNPFA SEI messages with nnpfa_target_id in the range of 256 to 511, inclusive, or 2 31 - 2, inclusive. The nnpfa_target_id values in the range of 256 to 511, inclusive, and 2 32 - 2 can be reserved for future use. Decoders shall ignore NNPFA SEI messages with nnpfa_target_id in the range of 256 to 511, inclusive, or 2 31 - 2, inclusive. The nnpfa_target_id values in the range of 256 to 511, inclusive, and 2 32 - 2 can be reserved for future use. Decoders shall ignore NNPFA SEI messages with nnpfa_target_id in the range of 256 to 511, inclusive, or 2

[0288] A NNPFA SEI message with a particular nnpfa_target_id value shall not exist in the current PU unless one or both of the following conditions are true:

[0289] - Within the current CLVS, there exists a NNPFC SEI message with nnpfc_id equal to the particular nnpfa_target_id value, which message exists in a PU that precedes the current PU in decoding order.

[0290] - There exists in the current PU a NNPFC SEI message with nnpfc_id equal to the particular nnpfa_target_id value.

[0291] When a PU contains both a NNPFC SEI message with a particular nnpfc_id value and a NNPFA SEI message with nnpfa_target_id equal to the particular nnpfc_id value, in decoding order, the NNPFC SEI message shall precede the NNPFA SEI message.

[0292] nnpfa_cancel_flag equal to 1 can specify that the persistence of the target NNPF established by any previous NNPFA SEI message with the same nnpfa_target_id as the current SEI message is cancelled, i.e., the target NNPF is no longer used, unless it is activated by another NNPFA SEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 0. nnpfa_cancel_flag equal to 0 can specify that nnpfa_persistence_flag is followed.

[0293] nnpfa_persistence_flag can specify the persistence of the target NNPF for the current layer. nnpfa_persistence_flag equal to 0 can specify that the target NNPF can only be used for post-processing filtering of the current picture. nnpfa_persistence_flag equal to 1 can specify that the target NNPF can be used for post-processing filtering of the current picture and all subsequent pictures of the current layer in output order until one or more of the following conditions are true:

[0294] - A new CLVS of the current layer starts.

[0295] - The bitstream ends.

[0296] - output the pictures in the current layer associated with the NNPFA SEI messages having the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1 after the current picture in output order.

[0297] For the subsequent pictures in the current layer associated with the NNPFA SEI messages having the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1, the target NNPFA is not applied.

[0298] nnpfcTargetPictures can be the set of pictures referred to by the last NNPFC SEI message with nnpfc_id equal to nnpfa_target_id that precede the current NNPFA SEI message in decoding order. nnpfaTargetPictures can be the set of pictures for which the current NNPFA SEI message activates the target NNPFA. One requirement for bitstream conformance is that any picture included in nnpfaTargetPictures shall also be included in nnpfcTargetPictures.

[0299] Post filter hint

[0300] The syntax structure of the post-filter hint is shown in Table 19.

[0301] [Table 19]

[0302]

[0303] The post-filter hint syntax structure of Table 19 can be signaled in the form of an SEI message. The SEI message that signals the post-filter hint syntax structure of Table 19 can be referred to as the post-filter hint SEI message.

[0304] The post-filter hint SEI message can provide the correlation information of the post-filter coefficients or the post-filter design in order to potentially use a set of pictures that are decoded and output in the post-processing in order to obtain an improved display quality.

[0305] filter_hint_cancel_flag equal to 1 can specify that the SEI message cancels the persistence of any previous post-filter hint SEI message to be applied to the current layer in output order. filter_hint_cancel_flag equal to 0 can specify that the post-filter hint information is followed.

[0306] filter_hint_persistence_flag can specify the persistence of the post filter hint SEI message for the current layer. A filter_hint_persistence_flag equal to 0 can specify that the post filter hint is only applicable to the current decoded picture. A filter_hint_persistence_flag equal to 1 can specify that the post filter hint SEI message is applicable to the current decoded picture and remains valid for all subsequent pictures in output order in the current layer until one or more of the following conditions are true:

[0307] - a new CLVS of the current layer starts.

[0308] - the end of the bitstream is reached.

[0309] - a picture in the AU associated with the post filter hint SEI message in the current layer is output after the current picture in output order.

[0310] filter_hint_size_y can specify the vertical dimension of the filter coefficients or the related array. The value of filter_hint_size_y shall be in the range of 1 to 15, inclusive.

[0311] filter_hint_size_x can specify the horizontal dimension of the filter coefficients or the related array. The value of filter_hint_size_x shall be in the range of 1 to 15, inclusive.

[0312] filter_hint_type can identify the type of filter hint sent as specified in Table 20. The value of filter_hint_type shall be in the range of 0 to 2, inclusive. A filter_hint_type value equal to 3 can not be present in the bitstream. Decoders shall ignore post filter hint SEI messages with filter_hint_type equal to 3.

[0313] [Table 20]

[0314] Value Description 0 Coefficients of 2D-FIR filter 1 Coefficients of 1D-FIR filter 2 Cross-correlation matrix

[0315] A filter_hint_chroma_coeff_present_flag equal to 1 can specify that chroma filter coefficients are present. A filter_hint_chroma_coeff_present_flag equal to 0 can specify that chroma filter coefficients are not present.

[0316] filter_hint_value[cldx][cy][cx] can specify in 16-bit precision the elements of a filter coefficient or cross-correlation matrix between the original signal and the decoded signal. The value of filter_hint_value[cldx][cy][cx] shall be in the range of -2 31 +1 to 2 31 -1 (including -2 31 +1 and 2 31 -1). cldx can specify the related color component, cy denotes the counter in vertical direction, and cx denotes the counter in horizontal direction. Depending on the value of filter_hint_type, the following can apply:

[0317] - If filter_hint_type is equal to 0, the coefficients of a two-dimensional Finite Impulse Response (FIR) filter of size filter_hint_size_y x filter_hint_size_x can be signaled.

[0318] - Otherwise, if filter_hint_type is equal to 1, the filter coefficients of two one-dimensional FIR filters can be signaled. In this case, filter_hint_size_y shall be equal to 2. The cy index equal to 0 specifies the filter coefficients of the horizontal filter and the cy index equal to 1 specifies the filter coefficients of the vertical filter. In the filtering process, the horizontal filter is applied first and the result is filtered by the vertical filter.

[0319] - Otherwise (filter_hint_type is equal to 2), the signaled hint can specify the cross-correlation matrix between the original signal s and the decoded signal s'.

[0320] The normalized cross-correlation matrix for the related color component identified by cldx of size filter_hint_size_y x filter_hint_size_x can be defined as in Equation 6.

[0321] [Equation 6]

[0322]

[0323] In Formula 6, s denotes a sample array of a color component cIdx of an original picture, s' denotes a corresponding array of a decoded picture, h denotes a vertical height of a relevant color component, w denotes a horizontal width of the relevant color component, bitDepth denotes a bit depth of the color component, OffsetY is equal to (filter_hint_size_y >> 1), OffsetX is equal to (filter_hint_size_x >> 1), cy ranges from 0 <= cy < filter_hint_size_y, and cx ranges from 0 <= cx < filter_hint_size_x.

[0324] The decoder can derive the Wiener post-filter from the cross-correlation matrix of the original signal and the decoded signal and the auto-correlation matrix of the decoded signal.

[0325] Problems of related art

[0326] According to the present disclosure, the problems of the prior art are as follows:

[0327] The conventional neural network post-filter characteristic (NNPFC) SEI message does not consider discardable pictures and non-output pictures. If this is not described, the following cases are possible or allowed.

[0328] Case 1) An access unit (AU) containing a picture marked as a discardable picture can include a NNPFC SEI message.

[0329] Case 2) An AU containing a picture marked as a discardable picture or a non-output picture can include a neural network post-filter activation (NNPFA) SEI message.

[0330] Case 3) When a NNPF is activated, the NNPF can take input pictures containing one or more discardable pictures and / or non-output pictures.

[0331] However, if the NNPFC SEI message and / or the NNPFA SEI message is sent in an AU including a discardable picture, it can cause serious problems. Specifically, if the picture included in the AU is discarded, the SEI message sent in the AU can also be discarded. In addition, the NNPFA SEI message should not be sent in an AU including a non-output picture. This is because the SEI message design is broken when the NNPFA SEI message activates the NNPF, the current picture (i.e., the picture in the same AU as the NNPFA SEI message) becomes the first input picture of the NNPF. In addition, if one or more input pictures of the NNPF are discardable pictures and / or non-output pictures, the list of input pictures in the encoder and the list of input pictures in the decoder can not be consistent with each other.

[0332] To solve the above-described problems of the prior art, improvements are needed for the NNPFC SEI message and / or the NNPFA SEI message.

[0333] Embodiments

[0334] Hereinafter, the NNPFC can be signaled in the form of an SEI message as the NNPFC syntax structure of Table 1 to Table 3, in which case the NNPFC can be an NNPFC SEI message. The NNPFA can be signaled in the form of an SEI message as the NNPFA syntax structure of Table 18, in which case the NNPFA can be an NNPFA SEI message. The post-filter hint can be signaled in the form of an SEI message as the post-filter hint syntax structure of Table 19, in which case the post-filter hint can be a post-filter hint SEI message.

[0335] Embodiments according to the present disclosure can include various aspects to improve some or all of the above-described problems. These aspects can be applied alone or in combination with two or more.

[0336] Aspect 1) can specify a constraint such that the NNPFC SEI message cannot be present within an AU containing a discardable picture. The discardable picture can refer to a picture that can be removed from the bitstream without affecting the decoding of other pictures in the coded bitstream.

[0337] Aspect 2) alternatively, a constraint can be specified such that the NNPFC SEI message cannot be associated with a discardable picture.

[0338] Aspect 3) can specify a constraint such that the NNPFA SEI message cannot be present within an AU containing a non-output picture. The non-output picture can refer to a coded picture that can be decoded by a decoder but the decoder can not generate a reconstructed picture therefrom.

[0339] Aspect 4) alternatively, a constraint can be specified such that the NNPFA SEI message cannot be associated with a non-output picture.

[0340] Aspect 5) can specify a constraint such that the input picture of an activated NNPF should not be a discardable picture.

[0341] Referring to Table 1 to Table 3 and Table 18, the above-described NNPFC syntax structure and semantics, NNPFA syntax structure and semantics have been described.

[0342] According to the present disclosure, at least some of the above-described problems can be solved by improving at least some of the above-described NNPFC syntax structure and semantics, NNPFA syntax structure and semantics by considering at least some of the above-described aspects 1 to 5.

[0343] In the present disclosure, for example, an AU will be described as a unit including an NNPFC SEI message (NNPFC SEI message and / or NNPFA SEI message) and a picture. However, it will be apparent to those skilled in the art that the unit including the NNPFC SEI message and the picture is not limited to an AU, and any unit such as a picture unit (PU) can be used. In addition, in the present disclosure, the unit can be collectively referred to as "unit information".

[0344] Embodiment 1

[0345] Embodiment 1 of the present disclosure relates to the above-described aspect 1, aspect 3, and / or aspect 5. According to Embodiment 1 of the present disclosure, the NNPFC SEI message and / or the NNPFA SEI message is restricted not to be included in the unit information including a discardable picture and / or a non-output picture, thereby solving the above-described problem of the prior art.

[0346] According to Embodiment 1 of the present disclosure, the semantics of the NNPFC SEI message can be restricted, modified, and / or changed.

[0347] Specifically, the NNPFC SEI message can be restricted not to be in an AU containing a coded picture that is discardable (discardable picture). As described above, the discardable picture can refer to a picture that can be removed / dropped from a bitstream without affecting the decoding of other pictures in the coded bitstream.

[0348] In addition, when the NNPFC SEI message defines an NNPF that is activated by the NNPFA SEI message, the list of input pictures of the NNPF can be restricted not to include a discardable picture.

[0349] According to Embodiment 1 of the present disclosure, the restriction on the NNPFC SEI message can be implemented by specifying a restriction on the semantics of the NNPFC SEI message.

[0350] More specifically, the semantics of the NNPFC SEI message can be specified with the following Table 21.

[0351] [Table 21]

[0352]

[0353] In addition, according to Embodiment 1 of the present disclosure, the semantics of the NNPFA SEI message can be restricted, modified, and / or changed. Specifically, the NNPFA SEI message can be restricted not to exist in an AU containing a coded picture that is not an output picture (non-output picture).

[0354] For example, if the AU includes the NNPFA SEI message, the pictures included in the AU can be restricted to have to be output. In this case, the AU can be replaced with the term unit information as described above. This constraint that the pictures are output can also be implemented by constraining the value (e.g., flag) of information specifying whether the picture is output to a value indicating "output".

[0355] According to Embodiment 1 of the present disclosure, the constraint on the NNPFA SEI message can be implemented by specifying a constraint on semantics of the NNPFA SEI message.

[0356] More specifically, the semantics of the NNPFA SEI message can be specified with the constraint of Table 22 below.

[0357] [Table 22]

[0358] The NNPFA SEI message shall not be present in an access unit that also contains coded pictures that are not output pictures.

[0359] According to Embodiment 1 of the present disclosure, when the discardable picture is discarded, effects of preventing cases where the NNPFC SEI message included in the same AU is discarded together can be expected. Further, effects of preventing cases where the picture activated by the NNPFA SEI message is not output can be expected. In addition, effects of preventing cases where the input picture list in the encoder does not match the input picture list in the decoder can be expected. According to Embodiment 1 of the present disclosure, effects of solving the above-described prior art problems, including the above-described effects, can be expected.

[0360] Embodiment 2

[0361] Embodiment 2 of the present disclosure relates to the above-described Aspect 2, Aspect 4, and / or Aspect 5. According to Embodiment 2 of the present disclosure, the NNPFC SEI message and / or the NNPFA SEI message can be restricted not to be associated with a discardable picture and / or a non-output picture, thereby solving the above-described prior art problems.

[0362] According to Embodiment 2 of the present disclosure, the semantics of the NNPFC SEI message can be constrained, modified, and / or changed.

[0363] Specifically, the NNPFC SEI message can be restricted not to be associated with a discardable coded picture (discardable picture). As described above, the discardable picture can refer to a picture that can be removed / discarded from the bitstream without affecting the decoding of other pictures in the coded bitstream.

[0364] Further, when the NNPF defined by the NNPFC SEI message is activated by the NNPFA SEI message, the input picture list of the NNPF can be restricted not to include the discardable picture.

[0365] According to Embodiment 2 of the present disclosure, the constraint on the NNPFC SEI message can be implemented by specifying a constraint on semantics of the NNPFC SEI message.

[0366] More specifically, the following Table 23 can be specified for semantics of the NNPFC SEI message.

[0367] [Table 23]

[0368]

[0369] Further, according to Embodiment 2 of the present disclosure, semantics of the NNPFA SEI message can be constrained, modified, and / or changed. Specifically, the NNPFA SEI message can be limited to not be associated with an encoded picture that is not an output picture (non-output picture).

[0370] According to Embodiment 2 of the present disclosure, the constraint on the NNPFA SEI message can be implemented by specifying a constraint on semantics of the NNPFA SEI message.

[0371] More specifically, the following Table 24 can be specified for semantics of the NNPFA SEI message.

[0372] [Table 24]

[0373] The NNPFA SEI message shall not be associated with coded pictures that are not output pictures.

[0374] According to Embodiment 2 of the present disclosure, when a discardable picture is discarded, effects of preventing a case where the NNPFC SEI message included in the same AU is discarded together can be expected. Further, effects of preventing a case where a picture activated by the NNPFA SEI message is not output can be expected. In addition, effects of preventing a case where an input picture list in an encoder and an input picture list in a decoder do not match can be expected. According to Embodiment 2 of the present disclosure, effects of solving the above-described prior art problems, including the above-described effects, can be expected.

[0375] Hereinafter, an image encoding method and an image decoding method according to various embodiments of the present disclosure will be described.

[0376] Figure 5 is a flowchart for explaining an image encoding method to which an embodiment of the present disclosure is applicable.

[0377] Figure 6 is a flowchart for explaining an image decoding method to which an embodiment of the present disclosure is applicable.

[0378] Figure 5 The image encoding method of can be performed by the image encoding apparatus 100, and Figure 6 The image decoding method of can be performed by the image decoding apparatus 200.

[0379] Referring to Figure 5 , the image encoding apparatus can generate NNPF (Neural Network Post Filter) related information about a neural network post-processing filter to be applied to a current picture (S501). The image encoding apparatus can encode the generated NNPF related information to generate an NNPF (Neural Network Post Filter) related SEI message (S502). The image encoding apparatus can transmit the generated NNPF related SEI message to, for example, an image decoding apparatus (S503). The image encoding apparatus performs steps S501 and S502, and step S503 can be a part of a transmission method performed by a separate transmission apparatus.

[0380] Referring to Figure 6 , the image decoding apparatus can receive an NNPF related SEI message about a neural network post-processing filter to be applied to a current picture (S601). The image decoding apparatus can decode the received NNPF related SEI message to reconstruct NNPF related information (S602). The image decoding apparatus can apply an NNPF to the current picture based on the reconstructed NNPF related information (S603). The above-described steps S602 and S603 can be performed on the condition that the NNPF is applied to the current picture.

[0381] In the descriptions referring to Figure 5 and Figure 6 , the NNPF related SEI message can include an NNPFC SEI message, an NNPFA SEI message, and / or an SEI message about a post-filtering hint according to the disclosure. In addition, the NNPF related information can refer to information signaled through syntax elements included in the NNPF related SEI message.

[0382] The NNPF related SEI message can be transmitted by being included in a predetermined unit. The predetermined unit can include one or more pictures. The predetermined unit can be an AU, but is not limited thereto, and can be expressed as "unit information" in the disclosure.

[0383] Figure 7 is a flowchart for explaining another image encoding method to which embodiments of the disclosure are applicable.

[0384] Figure 8 is a flowchart for explaining another image decoding method to which embodiments of the disclosure are applicable.

[0385] Figure 7 The image encoding method of Figure 8 may be performed by the image encoding apparatus 100, and The image decoding method of may be performed by the image decoding apparatus 200.

[0386] Referring toFigure 7 The image encoding apparatus can perform the step S701 of encoding the current picture and the step S702 of configuring unit information including the encoded current picture.

[0387] The image encoding apparatus can configure a bitstream including the unit information and transmit the bitstream to the image decoding apparatus. The image decoding apparatus can receive the bitstream and obtain the unit information included in the bitstream.

[0388] In particular, with reference to Figure 8 The image decoding apparatus can perform the step S801 of obtaining the unit information including the current picture and the step S802 of decoding the current picture based on the unit information.

[0389] As to the NNPFC SEI message and / or the NNPFA SEI message, the image encoding method and / or the image decoding method according to the embodiments of the disclosure and according to the aspects of the disclosure can be applied to the image encoding method and / or the image decoding method described with reference to Figure 7 and Figure 8 of the disclosure.

[0390] In particular, in the image encoding method and / or the image decoding method according to the disclosure, the type of the current picture can be restricted based on a NNPF (Neural Network Post Filter) related SEI (Supplemental Enhancement Information) message included in the unit information.

[0391] Further, in the image encoding method and / or the image decoding method according to the disclosure, based on the unit information including a NNPFA (Neural Network Post Filter Activation) SEI message, the current picture can be restricted to an output picture among the output picture and the non-output picture.

[0392] Further, in the image encoding method and / or the image decoding method according to the disclosure, based on the current picture being a non-output picture, the unit information can be restricted to not including the NNPFA SEI message.

[0393] Further, in the image encoding method and / or the image decoding method according to the disclosure, based on the current picture being a discardable picture, the unit information should not include a NNPFC (Neural Network Post Filter Characteristics) SEI message.

[0394] Further, in the image encoding method and / or the image decoding method according to the disclosure, the NNPFC SEI message included in the unit information can be restricted to not being associated with the discardable picture.

[0395] Further, in the image encoding method and / or the image decoding method according to the disclosure, the NNPFA SEI message included in the unit information can be restricted to not being associated with the non-output picture. Further, in the image encoding method and / or the image decoding method according to the disclosure, the NNPFA SEI message included in the unit information can be restricted to not being associated with the non-output picture.

[0396] Further, in the image encoding method and / or the image decoding method according to the present disclosure, an input picture of the NNPF activated by the NNPF-related SEI message can be restricted to not include a discardable picture.

[0397] Referring to Figure 7 and Figure 8 The image encoding method and the image decoding method described can be combined with the image encoding method and the image decoding method described with reference to Figure 5 and Figure 6 For example, the image encoding apparatus that performs the image encoding method of Figure 7 may generate and transmit the NNPF-related SEI message regarding the current picture to the image decoding apparatus by performing the image encoding method of Figure 5 Further, the image decoding apparatus that performs the image decoding method of Figure 8 may reconstruct the NNPF-related information from the NNPF-related SEI message and then apply the NNPF to the current picture by performing the image decoding method of Figure 6

[0398] According to the embodiments of the present disclosure, when a discardable picture is discarded, effects of preventing a case where the NNPFC SEI message included in the same AU is discarded together, effects of preventing a case where a picture activated by the NNPFA SEI message is not output, and effects of preventing a case where an input picture list in an encoder does not match an input picture list in a decoder can be expected. Including the above effects, effects of solving the above-described prior art problems can be expected.

[0399] Figure 9 is a diagram illustrating a content streaming system to which embodiments of the present disclosure can be applied.

[0400] As illustrated in Figure 9 , the content streaming system to which embodiments of the present disclosure can be applied can generally include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0401] The encoding server compresses content input from a multimedia input device such as a smartphone, a camera, a camcorder, or the like into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, a camcorder, or the like directly generates a bitstream, the encoding server can be omitted.

[0402] The bitstream can be generated by the image encoding method or the image encoding apparatus to which embodiments of the present disclosure can be applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream. ​

[0403] The streaming server transmits multimedia data to the user equipment based on a request of the user through the web server, and the web server serves as a medium to inform the user of the service. When the user requests a desired service to the web server, the web server can deliver the request to the streaming server, and the streaming server can transmit multimedia data to the user. In this case, the content streaming system can include a separate control server. At this time, the control server is used to control commands / responses between devices in the content streaming system.

[0404] The streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store a bitstream for a predetermined time.

[0405] Examples of the user equipment can include a mobile phone, a smart phone, a laptop computer, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation, a tablet personal PC, a tablet PC, an ultrabook, a wearable device (e.g., a smart watch, smart glasses, a head-mounted display), a digital TV, a desktop computer, a digital signage, etc.

[0406] Each server in the content streaming system can operate as a distributed server, in which case, distributed processing can be performed on data received from each server.

[0407] The scope of the present disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) for enabling operations of the methods according to various embodiments to be performed on devices or computers, non-transitory computer-readable media on which such software or commands are stored and executable on devices or computers.

[0408] Industrial applicability

[0409] Embodiments of the present disclosure can be used to encode or decode an image.

Claims

1. A picture decoding method performed by a picture decoding apparatus, the picture decoding method comprising: obtaining unit information including a current picture; and decoding the current picture based on the unit information, wherein a type of the current picture is restricted based on a neural network post filter (NNPF) related supplemental enhancement information (SEI) message included in the unit information. the current picture is restricted to an output picture among the output picture and a non-output picture based on the unit information including a neural network post filter activation (NNPFA) SEI message.

2. The image decoding method of claim 1, wherein, the unit information is restricted to not including the NNPFA SEI message based on the current picture being the non-output picture.

3. The image decoding method of claim 1, wherein, the unit information is restricted to not including a neural network post filter characteristic (NNPFC) SEI message based on the current picture being a discardable picture.

4. The image decoding method of claim 1, wherein, the NNPFC SEI message included in the unit information is restricted to not being associated with the discardable picture.

5. The image decoding method of claim 1, wherein, the NNPFC SEI message included in the unit information is restricted to not being associated with the non-output picture.

6. The image decoding method of claim 1, wherein, an input picture of a NNPF activated by the NNPF related SEI message is restricted to not including the discardable picture.

7. The image decoding method of claim 1, wherein, 8.A picture encoding method performed by a picture encoding apparatus, the picture encoding method comprising: encoding a current picture; and configuring unit information including the encoded current picture, wherein a type of the current picture is restricted based on a neural network post filter (NNPF) related supplemental enhancement information (SEI) message included in the unit information. the current picture is restricted to an output picture among the output picture and a non-output picture based on the unit information including a neural network post filter activation (NNPFA) SEI message. the unit information is restricted to not including the NNPFA SEI message based on the current picture being the non-output picture.

9. The image coding method of claim 8, wherein, the unit information is restricted to not including a neural network post filter characteristic (NNPFC) SEI message based on the current picture being a discardable picture.

10. The image coding method of claim 8, wherein, the NNPFC SEI message included in the unit information is restricted to not being associated with the discardable picture.

11. The image coding method of claim 8, wherein, the NNPFC SEI message included in the unit information is restricted to not being associated with the non-output picture.

12. The image coding method of claim 8, wherein, an input picture of a NNPF activated by the NNPF related SEI message is restricted to not including the discardable picture.

13. The image coding method of claim 8, wherein, 15.A computer readable recording medium storing a bitstream generated by a picture encoding method, the picture encoding method comprising:

14. The image coding method of claim 8, wherein, encoding a current picture; and configuring unit information including the encoded current picture, wherein a type of the current picture is restricted based on a neural network post filter (NNPF) related supplemental enhancement information (SEI) message included in the unit information. 16.A method of transmitting a bitstream generated by a picture encoding method, the picture encoding method comprising: encoding a current picture; and configuring unit information including the encoded current picture, ​ ​ ​ ​ Among them, the type of the current picture is limited based on a neural network post-filter (NNPF) related supplemental enhancement information (SEI) message included in the unit information.