Image encoding / decoding method, method for transmitting bit stream, and recording medium for storing bit stream

By processing NNPFC SEI messages, determining and activating neural network post-processing filters, the problem of low efficiency in encoding/decoding of high-resolution and high-quality images is solved, and transmission and storage costs are reduced.

CN120604516APending Publication Date: 2025-09-05LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380091758.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-19
Filing Date
2023-11-20
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing technologies have low encoding/decoding efficiency when processing high-resolution and high-quality images, resulting in increased transmission and storage costs.

Method used

By processing an NNPFC SEI message including a base neural network post-processing filter, a neural network that can be used as a post-processing filter is determined, and based on the format and usage information included in the NNPFC SEI message, it is determined whether to activate the target neural network post-processing filter.

Benefits of technology

Improves the efficiency of image encoding/decoding and reduces transmission and storage costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120604516A_ABST
    Figure CN120604516A_ABST
Patent Text Reader

Abstract

An image encoding / decoding method, a method for transmitting a bitstream, and a computer-readable recording medium for storing the bitstream are provided. An image decoding method according to the present disclosure comprises the steps of: acquiring a supplemental enhancement information (SEI) message for a post neural network filter (NNPF) to be applied to a current picture; determining at least one neural network capable of serving as a post-processing filter based on at least one neural network post filter characteristic (NNPFC) SEI message included in the SEI message for the NNPF, based on the SEI message for the NNPF being applied to the current picture; and determining whether to activate a target neural network post-processing filter applicable to the current picture based on at least one neural network post-filter activation (NNPFA) SEI message included in the SEI message for the NNPF, where format information and usage information included in the NNPFC SEI message may be determined based on the NNPFC SEI message including a base neural network post-processing filter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method for image encoding / decoding, a method for transmitting a bitstream, and a recording medium storing the bitstream, and more particularly, to a method for processing a neural network post-processing filter. Background Art

[0002] In recent years, the demand for high-resolution and high-quality images, such as high-definition (HD) and ultra-high-definition (UHD), has been growing in various fields. As image data becomes higher in resolution and higher in quality, the amount of information transmitted, or the bit rate, increases compared to conventional image data. This increase in the amount of information transmitted, or the bit rate, leads to increased transmission and storage costs.

[0003] Therefore, efficient image compression technology is needed to effectively transmit, store, and reproduce information of high-resolution and high-quality images. Summary of the Invention

[0004] Technical issues

[0005] The present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0006] The present disclosure is to provide a method for processing NNPFC SEI messages including a basic neural network post-processing filter.

[0007] The present disclosure provides a method for processing NNPFC SEI messages including updating neural network post-processing filters.

[0008] This disclosure aims to clarify the meaning of the NNPFC SEI message in specific situations.

[0009] This disclosure aims to clearly provide the format, usage, and complexity information of the NNPFC SEI message.

[0010] This disclosure is intended to clarify the process when format and usage information does not exist.

[0011] The present disclosure is to provide a non-transitory computer-readable recording medium for storing a bit stream generated by the image encoding method according to the present disclosure.

[0012] The present disclosure is to provide a non-transitory computer-readable recording medium for storing a bit stream received and decoded by the image decoding apparatus according to the present disclosure and used for image reconstruction.

[0013] The present disclosure is to provide a method for transmitting a bitstream generated by the image encoding method according to the present disclosure.

[0014] The technical problems to be achieved by the present disclosure are not limited to the above-mentioned technical problems, and other technical problems not mentioned can be clearly understood by those skilled in the art from the following description.

[0015] Technical Solution

[0016] According to an embodiment of the present disclosure, an image decoding method performed by an image decoding device may include: obtaining a supplemental enhancement information (SEI) message for a neural network post filter (NNPF) to be applied to a current picture; based on the SEI message for the NNPF being applied to the current picture, determining at least one neural network that can be used as a post-processing filter based on at least one neural network post filter characteristic (NNPFC) SEI message included in the SEI message for the NNPF; and determining whether to activate a target neural network post-processing filter that can be applied to the current picture based on at least one neural network post filter activation (NNPFA) SEI message included in the SEI message for the NNPF, wherein format information and usage information included in the NNPFC SEI message can be determined based on the NNPFC SEI message including a basic neural network post-processing filter.

[0017] According to an embodiment of the present disclosure, an image encoding method performed by an image encoding device may include: encoding at least one neural network that can be used as a post-processing filter into at least one neural network post-filter characteristic (NNPFC) supplemental enhancement information (SEI) message; and encoding whether to activate a target neural network post-processing filter that can be applied to a current picture into at least one neural network post-filter activation (NNPFA) SEI message, wherein, based on the SEI message for the neural network post-filter (NNPF) being applied to the current picture in the image decoding device, format information and usage information included in the NNPFC SEI message can be determined based on the NNPFC SEI message including a basic neural network post-processing filter.

[0018] According to another embodiment of the present disclosure, a computer-readable recording medium may store a bitstream generated by the image encoding method or apparatus of the present disclosure.

[0019] According to another embodiment of the present disclosure, a transmission method may transmit a bit stream generated by the image encoding method or apparatus of the present disclosure.

[0020] The features of the present disclosure briefly summarized above are merely exemplary embodiments of the detailed description below and are not intended to limit the scope of the present disclosure.

[0021] Beneficial effects

[0022] According to the present disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.

[0023] According to the present disclosure, the type of information present in the message may be determined based on the type of the NNPFC SEI message.

[0024] According to the present disclosure, a method of processing format and usage information may be determined based on the type of the NNPFC SEI message.

[0025] According to the present disclosure, format and usage information may be inferred based on the NNPFC SEI message including the underlying neural network post-processing filter.

[0026] According to the present disclosure, a non-transitory computer-readable recording medium storing a bit stream generated by the image encoding method according to the present disclosure may be provided.

[0027] According to the present disclosure, a non-transitory computer-readable recording medium for storing a bit stream received and decoded by the image decoding apparatus according to the present disclosure and used for image reconstruction may be provided.

[0028] According to the present disclosure, a method for transmitting a bitstream generated by an image encoding method can be provided.

[0029] Effects obtainable from the present disclosure are not limited to the above-described effects, and other effects that are not described can be clearly understood by those having ordinary skill in the art from the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 A schematic diagram of a video encoding system to which embodiments of the present disclosure may be applied is shown.

[0031] Figure 2 A schematic diagram showing an image encoding device to which an embodiment of the present disclosure can be applied is shown.

[0032] Figure 3 A schematic diagram showing an image decoding device to which an embodiment of the present disclosure can be applied is shown.

[0033] Figure 4 is a diagram illustrating an interleaving scheme for deriving a luma channel.

[0034] Figure 5 is a flowchart for illustrating an image encoding method to which an embodiment of the present disclosure can be applied.

[0035] Figure 6 is a flowchart for illustrating an image decoding method to which an embodiment of the present disclosure can be applied.

[0036] Figure 7 An exemplary diagram showing a content streaming system to which embodiments of the present disclosure may be applied is shown. DETAILED DESCRIPTION

[0037] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement them. However, the present disclosure can be implemented in various forms and is not limited to the embodiments described herein.

[0038] When describing the embodiments of the present disclosure, when it is considered that a detailed description of a well-known configuration or function will obscure the main points of the present disclosure, its detailed description is omitted. Additionally, parts not related to the description of the present disclosure are omitted from the drawings, and similar reference numerals are assigned to similar components.

[0039] In the present disclosure, when a component is described as being "connected," "coupled," or "linked" to another component, this may include not only a direct connection but also an indirect connection with another component interposed therebetween. Additionally, when a component is described as "including" or "having" another component, unless explicitly stated otherwise, this does not exclude other components but may further include additional components.

[0040] In the present disclosure, the terms first, second, etc. are used only to distinguish one component from another component and do not limit the order or importance of the components unless otherwise explicitly stated. Therefore, within the scope of the present disclosure, the first component in one embodiment may be referred to as the second component in another embodiment, and similarly, the second component in one embodiment may be referred to as the first component in another embodiment.

[0041] In this disclosure, distinguishable components are described for the purpose of clearly illustrating their respective characteristics and do not necessarily imply that the components are separate. In other words, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed across multiple hardware or software units. Therefore, such integrated or distributed implementations are also included within the scope of this disclosure, even if they are not explicitly described.

[0042] In the present disclosure, the components described in the various embodiments do not necessarily mean essential components, and some may be optional components. Therefore, an embodiment consisting of a subset of the components described in one embodiment is also included in the scope of the present disclosure. Additionally, an embodiment including additional components in addition to the components described in the various embodiments is also included in the scope of the present disclosure.

[0043] The present disclosure relates to encoding and decoding of images, and unless otherwise defined in the present disclosure, terms used herein may have ordinary meanings commonly used in the technical field to which the present disclosure belongs.

[0044] In this disclosure, a "picture" generally refers to a unit representing a single image at a specific point in time. A slice / tile is a coding unit that constitutes a portion of a picture, and a picture may be composed of one or more slices / tiles. Additionally, a slice / tile may include one or more coding tree units (CTUs).

[0045] In the present disclosure, "pixel" or "picture element" may refer to the smallest unit constituting a picture (or image). Additionally, the term "sample" may be used as a corresponding term for a pixel. A sample may generally represent a pixel or a pixel value, and may indicate only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component.

[0046] In the present disclosure, a "unit" may refer to a basic unit of image processing. A unit may include at least one of a specific area of ​​a picture or information related to the area. Depending on the context, the term "unit" may be used interchangeably with "sample array," "block," "area," and the like. Typically, an M×N block may include a set (or array) of samples (or sample array) consisting of M columns and N rows, or a set (or array) of transform coefficients.

[0047] In the present disclosure, the term "current block" may refer to one of a "current coding block," a "current coding unit," a "coding target block," a "decoding target block," or a "processing target block." When prediction is performed, the "current block" may refer to a "current prediction block" or a "prediction target block." When transform (inverse transform) / quantization (dequantization) is performed, the "current block" may refer to a "current transform block" or a "transform target block." When filtering is performed, the "current block" may refer to a "filtering target block."

[0048] In the present disclosure, unless explicitly stated as a chroma block, the term "current block" may refer to a block including both a luma component block and a chroma component block, or may refer to a "luma block of the current block." The luma component block of the current block may be explicitly expressed with terms such as "luma block" or "current luma block," which clearly indicates that it is a luma component block. Additionally, the chroma component block of the current block may be explicitly expressed with terms such as "chroma block" or "current chroma block," which clearly indicates that it is a chroma component block.

[0049] In the present disclosure, " / " and "," may refer to "and / or". For example, "A / B" and "A, B" may refer to "A and / or B". Additionally, "A / B / C" and "A, B, C" may refer to "at least one of A, B and / or C".

[0050] In the present disclosure, "or" may refer to "and / or". For example, "A or B" may mean 1) only "A", 2) only "B", or 3) "A and B". Alternatively, in the present disclosure, "or" may also mean "additionally or alternatively".

[0051] Overview of Video Coding Systems

[0052] Figure 1 A schematic diagram of a video encoding system to which embodiments of the present disclosure may be applied is shown.

[0053] The video encoding system according to an embodiment may include an encoder device 10 and a decoder device 20. The encoder device 10 may transmit encoded video and / or image information or data to the decoder device 20 in the form of a file or stream via a digital storage medium or a network.

[0054] The encoder device 10 according to the embodiment may include a video source generator 11, an encoder 12, and a transmitter 13. The decoder device 20 according to the embodiment may include a receiver 21, a decoder 22, and a renderer 23. The encoder 12 may be referred to as a video / image encoder, and the decoder 22 may be referred to as a video / image decoder. The transmitter 13 may be included in the encoder 12. The receiver 21 may be included in the decoder 22. The renderer 23 may include a display, and the display may be configured as a separate device or an external component.

[0055] The video source generator 11 can obtain a video / image by capturing, synthesizing, or generating a video / image. The video source generator 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured videos / images, etc. The video / image generation device may include, for example, a computer, a tablet, or a smartphone, and may (electronically) generate the video / image. For example, a virtual video / image may be generated by a computer, etc., and in this case, the video / image capture process may be replaced by a process that generates relevant data.

[0056] The encoder 12 may encode the input video / image. The encoder 12 may perform a series of processes such as prediction, transformation, and quantization for compression and coding efficiency. The encoder 12 may output encoded data (encoded video / image information) in the form of a bitstream.

[0057] Transmitter 13 can obtain the encoded video / image information or data output in the form of a bitstream and transmit it in the form of a file or stream to receiver 21 of decoder device 20 or another external device via a digital storage medium or network. Digital storage media may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. Transmitter 13 may include components for generating media files using a predetermined file format and components for transmission via a broadcast / communication network. Transmitter 13 may be provided as a transmission device separate from encoding device 10. In this case, the transmission device may include at least one processor for obtaining the encoded video / image information or data in the form of a bitstream and a transmitter for delivering the data in the form of a file or stream. Receiver 21 may extract / receive the bitstream from the storage medium or network and transmit it to decoder 22.

[0058] The decoder 22 may decode the video / image by performing a series of processes corresponding to the operations of the encoder 12 , such as dequantization, inverse transformation, prediction, and the like.

[0059] The renderer 23 may render the decoded video / image. The rendered video / image may be displayed via a display unit.

[0060] Overview of Image Coding Devices

[0061] Figure 2 A schematic diagram showing an image encoding device to which an embodiment of the present disclosure can be applied is shown.

[0062] like Figure 2 As described above, the image encoding apparatus 100 may include an image segmenter 110, a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter-frame predictor 180, an intra-frame predictor 185, and an entropy encoder 190. The inter-frame predictor 180 and the intra-frame predictor 185 may be collectively referred to as a "predictor." The transformer 120, the quantizer 130, the dequantizer 140, and the inverse transformer 150 may be included in a residual processor. The residual processor may further include a subtractor 115.

[0063] All or at least some of the components constituting the image encoding apparatus 100 may be implemented as a single hardware component (ie, an encoder or a processor) depending on the embodiment. Additionally, the memory 170 may include a decoded picture buffer (DPB) and may be implemented by a digital storage medium.

[0064] The image splitter 110 may split the input image (or picture, frame) input to the image encoding device 100 into at least one processing unit. For example, a processing unit may be referred to as a coding unit (CU). A coding unit may be obtained by recursively splitting a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree, binary tree, or ternary tree (QT / BT / TT) structure. For example, a coding unit may be divided into coding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. To split a coding unit, a quadtree structure may be applied first, followed by a binary tree structure and / or a ternary tree structure. The encoding process according to the present disclosure may be performed based on a final coding unit that is not further split. A maximum coding unit may be used directly as the final coding unit, or a coding unit of greater depth obtained by splitting the maximum coding unit may be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and / or reconstruction, which will be described later. As another example, the processing unit used in the encoding process may be a prediction unit (PU) or a transform unit (TU). The prediction unit and the transform unit may each be divided or partitioned from the final coding unit. A prediction unit may be a unit for sample prediction, and a transform unit may be a unit for deriving a transform coefficient and / or a residual signal from the transform coefficient.

[0065] The predictor (inter-frame predictor 180 or intra-frame predictor 185) can perform prediction on the target block (current block) and generate a prediction block including prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block or coding unit (CU). The predictor can generate various information related to the prediction of the current block and send it to the entropy encoder 190. The prediction-related information can be encoded by the entropy encoder 190 and can be output in the form of a bitstream.

[0066] The intra-frame predictor 185 can predict the current block by referencing samples within the current picture. The referenced samples may be located in an adjacent area of ​​the current block, or may be located at a farther position depending on the intra-frame prediction mode and / or intra-frame prediction method. The intra-frame prediction mode may include multiple non-directional modes and multiple directional modes. The non-directional mode may include, for example, a DC mode and a planar mode. The directional mode may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the granularity of the prediction direction. However, this is an example, and depending on the configuration, a greater or lesser number of directional prediction modes may be used. The intra-frame predictor 185 may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.

[0067] The inter-frame predictor 180 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted at the block, sub-block, or sample level based on the correlation of motion information between neighboring blocks and the current block. Motion information may include a motion vector and a reference picture index. Motion information may also include information regarding the inter-frame prediction direction (i.e., L0 prediction, L1 prediction, Bi prediction, etc.). In inter-frame prediction, neighboring blocks may include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a collocated reference block or collocated coding unit (colCU), and the reference picture including the temporally neighboring block may be referred to as a collocated picture (colPic). For example, the inter-frame predictor 180 may construct a motion information candidate list based on the neighboring blocks and generate information indicating which candidate to use to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 180 can use the motion information of the neighboring blocks as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be sent. In motion vector prediction (MVP) mode, the motion vector of the neighboring block can be used as a motion vector predictor, and the motion vector of the current block can be signaled by encoding a motion vector difference and an indicator for the motion vector predictor. The motion vector difference may refer to the difference between the motion vector of the current block and the motion vector predictor.

[0068] The predictor can generate a prediction signal based on various prediction methods and / or prediction techniques described below. For example, the predictor can apply intra-frame prediction or inter-frame prediction for the prediction of the current block, and can also apply both intra-frame and inter-frame prediction simultaneously. A prediction method that applies both intra-frame and inter-frame prediction for the prediction of the current block is referred to as combined intra-frame inter-frame prediction (CIIP). Additionally, the predictor can perform intra-block copying (IBC) for the prediction of the current block. For example, intra-block copying can be used in applications such as picture content coding (SCC) in gaming content image / video coding. IBC is a method that predicts the current block using a pre-reconstructed reference block within the current picture that is located at a predetermined distance from the current block. When IBC is applied, the position of the reference block within the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC essentially performs prediction within the current picture, but because it derives the reference block within the current picture, it can operate similarly to inter-frame prediction. In other words, IBC can use at least one of the inter-frame prediction methods described in this disclosure.

[0069] The prediction signal generated by the predictor can be used to generate a reconstructed signal or a residual signal. The subtractor 115 can generate a residual signal (residual block, residual sample array) by subtracting the prediction signal (prediction block, prediction sample array) output from the predictor from the input image signal (original block, original sample array). The generated residual signal can be sent to the transformer 120.

[0070] The transformer 120 may generate transform coefficients by applying a transform method to the residual signal. For example, the transform method may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT refers to a transform obtained from a graph when relationship information between pixels is represented as a graph. CNT refers to a transform obtained based on a prediction signal generated by using all previously reconstructed pixels. The transform process may be applied to pixel blocks of the same square size or non-square, variable-sized blocks.

[0071] The quantizer 130 may quantize the transform coefficients and transmit the quantized transform coefficients to the entropy encoder 190. The entropy encoder 190 may encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 130 may rearrange the block-shaped quantized transform coefficients into a one-dimensional vector based on a coefficient scanning order and may generate information about the quantized transform coefficients based on the one-dimensional vector of the quantized transform coefficients.

[0072] The entropy encoder 190 can perform various encoding methods, such as Exponential Golomb, context-adaptive variable length coding (CAVLC), or context-adaptive binary arithmetic coding (CABAC). The entropy encoder 190 can not only encode the quantized transform coefficients, but also encode the information required for video / image reconstruction (i.e., the values ​​of syntax elements) together with the quantized transform coefficients or separately. The encoded information (i.e., the encoded video / image information) can be transmitted or stored in the form of a bitstream in a network abstraction layer (NAL) unit. The video / image information may also include information about various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may also include general constraint information. The signaling information, transmitted information, and / or syntax elements described in this disclosure may be encoded through the above-mentioned encoding process and included in the bitstream.

[0073] The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, etc., and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting a signal output from the entropy encoder 190 and / or a storage unit (not shown) for storing the signal may be provided as an internal / external element of the image encoding apparatus 100, or the transmitter may be configured as a component of the entropy encoder 190.

[0074] The quantized transform coefficients output from the quantizer 130 may be used to generate a residual signal. For example, the residual signal (residual block or residual sample) may be reconstructed by dequantizing and inverse-transforming the quantized transform coefficients through the dequantizer 140 and the inverse transformer 150.

[0075] The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, or reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-frame predictor 180 or the intra-frame predictor 185. When the target block has no residual (such as when skip mode is applied), the prediction block can be used as a reconstructed block. The adder 155 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next target block in the current picture, and as described later, after filtering, it can also be used for inter-frame prediction of the next picture.

[0076] The filter 160 may apply filtering to the reconstructed signal to enhance the subjective / objective quality. For example, the filter 160 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and the modified reconstructed picture may be stored in the memory 170, specifically in the DPB of the memory 170. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. The filter 160 may generate various filter-related information as described later in the description of each filtering method and send it to the entropy encoder 190. The filter-related information may be encoded by the entropy encoder 190 and output in the form of a bitstream.

[0077] The modified reconstructed picture transmitted to the memory 170 may be used as a reference picture in the inter-frame predictor 180. In this case, when applying inter-frame prediction, the image encoding apparatus 100 may avoid prediction mismatch between the image encoding apparatus 100 and the image decoding apparatus, and may improve encoding efficiency.

[0078] The DPB in the memory 170 can store the modified reconstructed picture for use as a reference picture in the inter-frame predictor 180. The memory 170 can store motion information of blocks in the current picture for which motion information has been derived (or encoded) and / or motion information of blocks in the reconstructed image. The stored motion information can be sent to the inter-frame predictor 180 to be used as motion information for spatially adjacent blocks or temporally adjacent blocks. The memory 170 can store reconstructed samples of the reconstructed blocks in the current picture and send them to the intra-frame predictor 185.

[0079] Overview of Image Decoding Equipment

[0080] Figure 3 A schematic diagram showing an image decoding device to which an embodiment of the present disclosure can be applied is shown.

[0081] like Figure 3 As shown, the image decoding apparatus 200 may include an entropy decoder 210, a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame predictor 260, and an intra-frame predictor 265. The inter-frame predictor 260 and the intra-frame predictor 265 may be collectively referred to as a "predictor." The dequantizer 220 and the inverse transformer 230 may be included in a residual processor.

[0082] All or at least some of the plurality of components constituting the image decoding apparatus 200 may be implemented as a single hardware component (ie, a decoder or a processor) depending on an embodiment. Additionally, the memory 170 may include a DPB and may be implemented by a digital storage medium.

[0083] The image decoding apparatus 200 receiving a bit stream including video / image information may perform the same operation as that performed by Figure 2 The image is reconstructed by processing corresponding to the processing performed by the image encoding device 100 in the image decoding device 200. For example, the image decoding device 200 can use the processing unit applied in the image encoding device to perform decoding. Therefore, the processing unit used for decoding can be, for example, a coding unit. The coding unit can be a coding tree unit, or can be obtained by dividing the maximum coding unit. In addition, the reconstructed image signal decoded and output by the image decoding device 200 can be played back by a playback device (not shown).

[0084] The image decoding device 200 may receive Figure 2The received signal is output by the image encoding device in the form of a bitstream. The entropy decoder 210 can be decoded. For example, the entropy decoder 210 can parse the bitstream to extract information required for image reconstruction (or picture reconstruction) (i.e., video / image information). The video / image information may also include information about various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may also include general constraint information. The image decoding device may additionally use the information about the parameter sets and / or the general constraint information to decode the image. The signaling information, reception information, and / or syntax elements described in the present disclosure may be obtained from the bitstream by decoding through a decoding process. For example, the entropy decoder 210 may decode the information in the bitstream based on a coding method such as Exponential Golomb, CAVLC, or CABAC, and may output syntax element values ​​required for image reconstruction and quantized values ​​of transform coefficients related to the residual. More specifically, the CABAC entropy decoding method can receive a bin corresponding to a syntax element in a bitstream, can use information about the decoded target syntax element, decoded information about the decoded target block and adjacent blocks, or information about previously decoded symbols / bins to determine a context model, can predict the probability of a bin appearing based on the determined context model, and can perform arithmetic decoding of the bin to generate a symbol corresponding to each syntax element. In this case, the CABAC entropy decoding method can use the decoded symbol / bin information to update the context model for the next symbol / bin after determining the context model. Among the decoded information from the entropy decoder 210, prediction-related information can be provided to the predictor (inter-frame predictor 260 and intra-frame predictor 265), and the residual value entropy-decoded by the entropy decoder 210 (in other words, quantized transform coefficients and related parameter information) can be input to the dequantizer 220. Additionally, among the decoded information from the entropy decoder 210, filtering-related information can be provided to the filter 240. In addition, a receiver (not shown) that receives a signal output from the image encoding apparatus may be additionally configured as an internal / external element of the image decoding apparatus 200 , or the receiver may be configured as a component of the entropy decoder 210 .

[0085] In addition, the image decoding device according to the present disclosure may also be referred to as a video / image / picture decoding device. The image decoding device may include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoder 210, and the sample decoder may include at least one of a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame predictor 260, or an intra-frame predictor 265.

[0086] The dequantizer 220 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 220 may rearrange the quantized transform coefficients into two-dimensional blocks. In this case, the rearrangement may be performed based on the coefficient scanning order used in the image encoding apparatus. The dequantizer 220 may dequantize the quantized transform coefficients using a quantization parameter (i.e., quantization step size information) and obtain the transform coefficients.

[0087] The inverse transformer 230 may perform inverse transform on the transform coefficients to obtain a residual signal (a residual block or a residual sample array).

[0088] The predictor may perform prediction on the current block and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on the prediction-related information output from the entropy decoder 210, and may determine a specific intra / inter prediction mode (prediction method).

[0089] The predictor can generate a prediction signal based on various prediction methods (techniques) to be described later, which is the same as the description of the predictor in the image encoding device 100 .

[0090] The intra predictor 265 may predict the current block by referring to samples within the current picture. The description of the intra predictor 185 may also be applied to the intra predictor 265 in the same manner.

[0091] The inter-frame predictor 260 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted at the block, sub-block, or sample level based on the correlation of motion information between neighboring blocks and the current block. Motion information may include a motion vector and a reference picture index. Motion information may also include information regarding the inter-frame prediction direction (i.e., L0 prediction, L1 prediction, Bi prediction, etc.). In inter-frame prediction, neighboring blocks may include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. For example, the inter-frame predictor 260 may construct a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index of the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes (methods), and prediction-related information may include information indicating the inter-frame prediction mode (method) applied to the current block.

[0092] The adder 235 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter-frame predictor 260 and / or the intra-frame predictor 265). When the target block has no residual (such as when skip mode is applied), the prediction block can be used as the reconstructed block. The description of the adder 155 can also be applied to the adder 235 in the same manner. The adder 235 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next target block in the current picture, and as described later, after filtering, it can also be used for inter-frame prediction of the next picture.

[0093] The filter 240 may apply filtering to the reconstructed signal to enhance the subjective / objective quality. For example, the filter 240 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and the modified reconstructed picture may be stored in the memory 250, specifically in the DPB of the memory 250. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc.

[0094] The (modified) reconstructed picture stored in the DPB of the memory 250 can be used as a reference picture in the inter-frame predictor 260. The memory 250 can store motion information of blocks in the current picture for which motion information has been derived (or decoded) and / or motion information of blocks in the reconstructed image. The stored motion information can be sent to the inter-frame predictor 260 for use as motion information of spatially or temporally neighboring blocks. The memory 250 can store reconstructed samples of the reconstructed blocks in the current picture and send them to the intra-frame predictor 265.

[0095] In this specification, the embodiments described for the filter 160, the inter-frame predictor 180, and the intra-frame predictor 185 of the image encoding device 100 may be applied to the filter 240, the inter-frame predictor 260, and the intra-frame predictor 265 of the image decoding device 200 in the same or corresponding manner.

[0096] Neural Network Post-Filter Characteristic (NNPFC)

[0097] The combination of Table 1 and Table 2 represents the NNPFC grammatical structure.

[0098] [Table 1]

[0099]

[0100] [Table 2]

[0101]

[0102] The NNPFC syntax structures of Table 1 and Table 2 may be signaled in the form of a Supplemental Enhancement Information (SEI) message. The SEI message that signals the NNPFC syntax structures of Table 1 and Table 2 may be referred to as an NNPFC SEI message.

[0103] The NNPFC SEI message may specify a neural network that may be used as a post-processing filter. The use of a specified post-processing filter for a particular picture may be indicated using a Neural Network Post-Filter Activation SEI message. Here, "post-processing filter" and "post-filter" may have the same meaning.

[0104] To use these SEI messages, you may need to define the following variables:

[0105] -The width and height of the decoded output picture can be cropped in units of luminance samples. The width and height can be expressed as CroppedWidth and CroppedHeight respectively.

[0106] - CroppedYPic[idx], which is the array of luma samples of the cropped decoded output picture, and CroppedCbPic[idx] and CroppedCrPic[idx] (if present), which are arrays of chroma samples, may be used as input to the post-processing filter, where idx may be in the range of 0 to numInputPics-1.

[0107] -BitDepthY may represent the bit depth of the luma sample array of the cropped decoded output picture.

[0108] -BitDepthC may represent the bit depth of the chroma sample array (if present) of the cropped decoded output picture.

[0109] -ChromaFormatIdc may represent a chroma format identifier.

[0110] - When the value of nnpfc_auxiliary_inp_idc is 1, the filter strength control value StrengthControlVal should be a real number in the range of 0 to 1.

[0111] The variables SubWidthC and SubHeightC can be derived from ChromaFormatIdc. Two or more NNPFC SEI messages can exist for the same picture. When two or more NNPFC SEI messages with different nnpfc_id values ​​exist or are activated for the same picture, the NNPFC SEI messages can have the same or different nnpfc_purpose and nnpfc_mode_idx values.

[0112] nnpfc_id may contain an identification number that can be used to identify the post-processing filter. The value of nnpfc_id should be between 0 and 2. 32 -2. 256 to 511 and 2 31 to 2 32 nnpfc_id values ​​in the range of -2 are reserved for future use. Decoders should ignore nnpfc_id values ​​in the range of 256 to 511 or 2. 31 to 2 32 NNPFC SEI message in the range of -2.

[0113] When the NNPFC SEI message is the first NNPFC SEI message with a particular nnpfc_id value in decoding order within the current coding layer video sequence (CLVS), the following applies:

[0114] -SEI messages may represent basic post-processing filters.

[0115] - A SEI message may be associated with the current decoded picture and all subsequent decoded pictures of the current layer in output order until the end of the current CLVS.

[0116] The NNPFC SEI message may be a repetition of a previous NNPFC SEI message in decoding order within the current CLVS, and subsequent semantics may apply under the assumption that this SEI message is the only NNPFC SEI message with the same content within the current CLVS.

[0117] When the NNPFC SEI message is not the first NNPFC SEI message with a particular nnpfc_id value in decoding order within the current CLVS, the following may apply.

[0118] - The SEI message can be associated with the current decoded picture and all subsequent decoded pictures of the current CLVS or current layer in output order until the end of the current CLVS, or can be associated with the next NNPFC SEI message with a specific nnpfc_id value in output order within the current CLVS.

[0119] When the SEI message is the first NNPFC SEI message with a particular nnpfc_id value in decoding order within the current CLVS, a value of 1 for nnpfc_mode_idc may indicate that the base post-processing filter associated with the nnpfc_id value is a neural network, and the neural network may be a neural network identified by a URI indicated by nnpfc_uri in the format identified by the tag URI nnpfc_tag_uri.

[0120] When the NNPFC SEI message is not the first NNPFC SEI message with a particular nnpfc_id value in decoding order within the current CLVS, a value of 1 for nnpfc_mode_idc may indicate that updates to base post-processing filters with the same nnpfc_id value are defined by the URI indicated by nnpfc_uri in the format identified by the tag URI nnpfc_tag_uri.

[0121] In the bitstream, the value of nnpfc_mode_idc may be constrained to be in the range of 0 to 1. nnpfc_mode_idc values ​​in the range of 2 to 255 are reserved for future use and may not be present in the bitstream. The decoder shall ignore NNPFC SEI messages with nnpfc_mode_idc values ​​in the range of 2 to 255. nnpfc_mode_idc values ​​greater than 255 shall not be present in the bitstream and may not be reserved for future use.

[0122] When the SEI message is the first NNPFC SEI message with a particular nnpfc_id value in decoding order within the current CLVS, the post-processing filter PostProcessingFilter() may be assigned to be the same as the base post-processing filter.

[0123] When the SEI message is not the first NNPFC SEI message with a particular nnpfc_id value in decoding order within the current CLVS, the post-processing filter PostProcessingFilter() may be obtained by applying the updates defined by the SEI message to the base post-processing filter.

[0124] Updates are not cumulative, but instead each update may be applied to the base post-processing filter specified by the first NNPFC SEI message with a particular nnpfc_id value in decoding order within the current CLVS.

[0125] nnpfc_reserved_zero_bit_a may be constrained to be equal to 0 due to bitstream constraints. A decoder may be constrained to ignore NNPFC SEI messages for which the value of nnpfc_reserved_zero_bit_a is not 0.

[0126] nnpfc_tag_uri may include a tag URI with syntax and semantics as specified in IETF RFC 4151 that identifies a neural network used as a base post-processing filter or an update related to a base post-processing filter using the nnpfc_id value specified by nnpfc_uri. By using nnpfc_tag_uri, the format of the neural network data specified by nnrpf_uri can be uniquely identified without a central registration authority. nnpfc_tag_uri equal to "tag:iso.org,2023:15938-17" may indicate that the neural network data identified by nnpfc_uri conforms to ISO / IEC 15938-17.

[0127] nnpfc_uri may comprise a URI having the syntax and semantics specified in IETF Internet Standard 66 that identifies a neural network used as a base post-processing filter or an update to a base post-processing filter having the same nnpfc_id value.

[0128] A value of 1 for nnfc_formatting_and_purpose_flag may indicate the presence of syntax elements related to filter purpose, input format, output format, and complexity. A value of 0 for nnfc_formatting_and_purpose_flag may indicate the absence of syntax elements related to filter purpose, input format, output format, and complexity.

[0129] When the SEI message is the first NNPFC SEI message with a particular nnpfc_id value in decoding order within the current CLVS, the value of nnpfc_formatting_and_purpose_flag may need to be equal to 1. When the SEI message is not the first NNPFC SEI message with a particular nnpfc_id value in decoding order within the current CLVS, the value of nnpfc_formatting_and_purpose_flag may need to be equal to 0.

[0130] nnpfc_purpose may indicate the purpose of the post-processing filter as specified in Table 3.

[0131] Due to bitstream constraints, nnpfc_purpose values ​​shall be in the range of 0 to 5. nnpfc_purpose values ​​from 6 to 1023 shall not be present in the bitstream and are reserved for future use. The decoder shall ignore NNPFC SEI messages with nnpfc_purpose in the range of 6 to 1023. nnpfc_purpose values ​​greater than 1023 shall not be present in the bitstream and are not reserved for future use.

[0132] [Table 3]

[0133]

[0134] When the reserved value of npfc_purpose is used in the future, the syntax of this SEI message may be extended to include a syntax element that exists conditionally on nnpfc_purpose being equal to the corresponding value.

[0135] When the value of SubWidthC is 1 and the value of SubHeightC is 1, the value of nnpfc_purpose should not be 2 or 4.

[0136] A value of 1 for nnpfc_out_sub_c_flag may indicate that the value of outSubWidthC is 1 and the value of outSubHeightC is 1. A value of 0 for nnpfc_out_sub_c_flag may indicate that the value of outSubWidthC is 2 and the value of outSubHeightC is 1. When nnpfc_out_sub_c_flag is not present, outSubWidthC may be inferred to be equal to SubWidthC, and outSubHeightC may be inferred to be equal to SubHeightC. When the value of ChromaFormatIdc is 2 and nnpfc_out_sub_c_flag is present, the value of nnpfc_out_sub_c_flag shall be equal to 1.

[0137] nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples may indicate the width and height, respectively, of the luma sample array of the picture resulting from applying the post-processing filter identified by nnpfc_id to the cropped decoded output picture. When nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples are not present, they may be inferred to be equal to CroppedWidth and CroppedHeight, respectively. The value of nnpfc_pic_width_in_luma_samples shall be in the range of CroppedWidth to CroppedWidth*16-1. The value of nnpfc_pic_height_in_luma_samples shall be in the range of CroppedHeight to CroppedHeight*16-1.

[0138] nnpfc_num_input_pics_minus2+2 may indicate the number of decoded output pictures used as input to the post-processing filter.

[0139] nnpfc_interpolated_pics[i] may represent the number of interpolated pictures generated by a post-processing filter between an i-th picture and an (i+1)-th picture used as an input of the post-processing filter.

[0140] A variable numInputPics indicating the number of pictures used as input of the post-processing filter and a variable numOutputPics indicating the number of pictures generated as a result of the post-processing filter may be derived as shown in Table 4.

[0141] [Table 4]

[0142]

[0143] The value 1 of nnpfc_component_last_flag may indicate that the last dimension of the input tensor inputTensor of the post-processing filter and the output tensor outputTensor obtained from the post-processing filter is used for the current channel. The value 0 of nnpfc_component_last_flag may indicate that the third dimension of the input tensor inputTensor of the post-processing filter and the output tensor outputTensor obtained from the post-processing filter is used for the current channel.

[0144] The first dimension of the input and output tensors can be used as a batch index used in some neural network frameworks. Although the formulas in the semantics of this SEI message use a batch size corresponding to batch index 0, the batch size used as input for neural network inference can be determined by the post-processing implementation.

[0145] For example, when the value of nnpfc_inp_order_idc is equal to 3 and the value of nnpfc_auxiliary_inp_idc is equal to 1, the input tensor may include 7 channels, including 4 luma matrices, 2 chroma matrices, and 1 auxiliary input matrix. In this case, the DeriveInputTensors() process may derive each of the 7 channels of the input tensor one by one, and when processing a particular channel among these channels, the channel may be referred to as the current channel during the process.

[0146] nnpfc_inp_format_idc may indicate a method of converting the sample values ​​of the cropped decoded output picture into input values ​​of the post-processing filter. When nnpfc_inp_format_idc is 0, the input value of the post-processing filter is a real number, and the InpY() and InpC() functions may be defined as shown in Equation 1.

[0147] [Formula 1]

[0148] InpY(x)=x÷((1<<BitDepthY)-1)

[0149] InpC(x)=x÷((1< <BitDepthC)-1)

[0150] When the value of nnpfc_inp_format_idc is 1, the input value of the post-processing filter is an unsigned integer, and the InpY() and InpC() functions can be derived as shown in Table 5.

[0151] [Table 5]

[0152]

[0153] The variable inpTensorBitDepth can be derived from the following syntax element nnpfc_inp_tensor_bitlength_minus8.

[0154] Values ​​of nnpfc_inp_format_idc greater than 1 may be reserved for future use and may not be present in the bitstream.A decoder shall ignore NNPFC SEI messages containing reserved values ​​of nnpfc_inp_format_idc.

[0155] The value of nnpfc_inp_tensor_bitlength_minus8+8 may represent the bit depth of the luma sample values ​​in the input integer tensor. The value of inpTensorBitDepth may be derived as shown in Equation 2.

[0156] [Formula 2]

[0157] inpTensorBitDepthnnpfc_inp_tensor_bitdepth_minus8+8

[0158] The value of nnpfc_inp_tensor_bitlength_minus8 should be constrained to exist in the range of 0 to 24.

[0159] nnpfc_inp_order_idc may indicate a method of arranging the sample array of the cropped decoded output picture as one of the input pictures of the post-processing filter.

[0160] The value of nnpfc_inp_order_idc shall be present in the bitstream in the range of 0 to 3. nnpfc_inp_order_idc values ​​between 4 and 255 shall not be present in the bitstream. The decoder shall ignore NNPFC SEI messages with nnpfc_inp_order_idc in the range of 4 to 255. nnpfc_inp_order_idc values ​​greater than 255 shall not be present in the bitstream and are not reserved for future use.

[0161] When the value of ChromaFormatIdc is not 1, the value of nnpfc_inp_order_idc should not be 3.

[0162] Table 6 includes a description of the values ​​of nnpfc_inp_order_idc.

[0163] [Table 6]

[0164]

[0165] A patch may be a rectangular array of samples of a component (eg, a luma component or a chroma component) from a picture.

[0166] nnpfc_auxiliary_inp_idc greater than 0 may indicate the presence of auxiliary input data in the input tensor of the neural network postfilter. A nnpfc_auxiliary_inp_idc value of 0 may indicate the absence of auxiliary input data in the input tensor. A nnpfc_auxiliary_inp_idc value of 1 may indicate that the auxiliary input data is derived using the methods described in Tables 7 to 9.

[0167] The value of nnpfc_auxiliary_inp_idc shall be present in the bitstream in the range of 0 to 1. nnpfc_auxiliary_inp_idc values ​​between 2 and 255 shall not be present in the bitstream. The decoder shall ignore NNPFC SEI messages with nnpfc_inp_order_idc values ​​between 2 and 255. nnpfc_inp_order_idc values ​​greater than 255 shall not be present in the bitstream and shall not be reserved for future use.

[0168] The procedure DeriveInputTensors() for deriving the input tensor inputTensor for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft of the top left sample position of a sample patch included in the specified input tensor can be described as a combination of Tables 7 to 9.

[0169] [Table 7]

[0170]

[0171] [Table 8]

[0172]

[0173] [Table 9]

[0174]

[0175] A value of 1 for nnpfc_separate_colour_description_present_flag may indicate that a unique combination of primary colors, transform characteristics, and matrix coefficients for the picture resulting from the post-processing filter is specified in the SEI message syntax structure. A value of 0 for nnfpc_separate_colour_description_present_flag may indicate that the combination of primary colors, transform characteristics, and matrix coefficients for the picture resulting from the post-processing filter is the same as indicated by the VUI parameters of the CLVS.

[0176] The nnpfc_colour_primaries may have the same semantics as the vui_colour_primaries syntax element, except for the following:

[0177] -nnpfc_colour_primaries may indicate the primaries of the picture resulting from applying the neural network postfilter specified in the SEI message, instead of the primaries used for CLVS.

[0178] - When nnpfc_colour_primaries is not present in the NNPFC SEI message, the value of nnpfc_colour_primaries may be inferred to be equal to the value of vui_colour_primaries.

[0179] The nnpfc_transfer_characteristics may have the same semantics as the vui_transfer_characteristics syntax element, except for the following:

[0180] -nnpfc_transfer_characteristics may indicate the transform characteristics of the picture resulting from applying the neural network postfilter specified in the SEI message, instead of the transform characteristics used for CLVS.

[0181] - When nnpfc_transfer_characteristics is not present in the NNPFC SEI message, the value of nnpfc_transfer_characteristics may be inferred to be equal to the value of vui_transfer_characteristics.

[0182] The nnpfc_matrix_coeffs may have the same semantics as the vui_matrix_coeffs syntax element, except for the following:

[0183] -nnpfc_matrix_coeffs may indicate the matrix coefficients of the picture resulting from applying the neural network postfilter specified in the SEI message, instead of the matrix coefficients used for CLVS.

[0184] - When nnpfc_matrix_coeffs is not present in the NNPFC SEI message, the value of nnpfc_matrix_coeffs may be inferred to be equal to the value of vui_matrix_coeffs.

[0185] - The allowed values ​​of nnpfc_matrix_coeffs may not be restricted by the chroma format of the decoded video picture indicated by the ChromaFormatIdc value in the semantics of the VUI parameter.

[0186] - When the value of nnpfc_matrix_coeffs is equal to 0, the value of nnpfc_out_order_idc should not be equal to 1 or 3.

[0187] A value of 0 for nnpfc_out_format_idc may indicate that for subsequent post - processing or display, the sample values output by the post - processing filter are real numbers in the range of 0 to 1, which are linearly mapped to unsigned integer values in the range of 0 to (1<<bitDepth)-1. A value of 1 for nnpfc_out_format_flag may indicate that the sample values output by the post - processing filter are unsigned integers in the range of 0 to (1<<(nnpfc_out_tensor_bitlength_minus8 + 8))-1. Values of nnpfc_out_format_idc greater than 1 should not exist in the bitstream. The decoder shall ignore the NNPFC SEI message containing the reserved value of nnpfc_out_format_idc. The “+8” may indicate the bit depth of the sample values in the output integer tensor. The value of nnpfc_out_tensor_bitlength_minus8 should exist in the range of 0 to 24. [[ID=I]]

[0188] nnpfc_out_order_idc may indicate the output order of the samples output by the post - processing filter. The value of nnpfc_out_order_idc should exist in the range of 0 to 3 in the bitstream. Values of nnpfc_out_order_idc from 4 to 255 should not exist in the bitstream. The decoder shall ignore the NNPFC SEI message with nnpfc_out_order_idc value in the range of 4 to 255. Values of nnpfc_out_order_idc greater than 255 should not exist in the bitstream and are not reserved for future use. When the value of nnpfc_purpose is 2 or 4, the value of nnpfc_out_order_idc should not be equal to 3.

[0189] Table 10 describes the values of nnpfc_out_order_idc.

[0190] [Table 10]

[0191]

[0192] [[ID=I5]]The process StoreOutputTensors() for deriving the sample values in FilteredYPic, FilteredCbPic, and FilteredCrPic as an array of filtered output samples from the output tensor outputTensor for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft indicating the top - left sample position of the patch including the samples in the input tensor can be represented as in the combination of Table 11 and Table 12.

[0193] [Table 11]

[0194]

[0195] [Table 12]

[0196]

[0197] A value of 1 for nnpfc_constant_patch_size_flag may indicate that the post-processing filter accepts exactly the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1 as input. A value of 0 for nnpfc_constant_patch_size_flag may indicate that the post-processing filter accepts any patch size that is a positive integer multiple of the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1 as input.

[0198] When the value of nnpfc_constant_patch_size_flag is 1, nnpfc_patch_width_minus1+1 may indicate the number of horizontal samples of the patch size required for the input of the post-processing filter. The value of nnpfc_patch_width_minus1 should be in the range of 0 to Min(32766, CroppedWidth-1).

[0199] When the value of nnpfc_constant_patch_size_flag is 1, nnpfc_patch_height_minus1+1 may indicate the number of vertical samples of the patch size required for the post-processing filter input. The value of nnpfc_patch_height_minus1 should be in the range of 0 to Min(32766, CroppedHeight-1).

[0200] The variables inpPatchWidth and inpPatchHeight can be set to the width and height of the patch size respectively.

[0201] When the value of nnpfc_constant_patch_size_flag is 0, the following applies:

[0202] -The values ​​of inpPatchWidth and inpPatchHeight can be provided by external means or set by the post-processor itself.

[0203] -inpPatchWidth should be a positive integer multiple of nnpfc_patch_width_minus1+1 and should be less than or equal to CroppedWidth. -inpPatchHeight should be a positive integer multiple of nnpfc_patch_height_minus1+1 and should be less than or equal to CroppedHeight.

[0204] Otherwise (ie, when the value of nnpfc_constant_patch_size_flag is 1), the value of inpPatchWidth may be set equal to nnpfc_patch_width_minus1+1, and the value of inpPatchHeight may be set equal to nnpfc_patch_height_minus1+1.

[0205] nnpfc_overlap indicates the number of horizontal and vertical samples that overlap adjacent input tensors of the post-processing filter. The value of nnpfc_overlap should be in the range of 0 to 16383.

[0206] The variables outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, outPatchCHeight, and overlapSize may be derived as shown in Table 13.

[0207] [Table 13]

[0208]

[0209] A bitstream conformance requirement may be that outPatchWidth*CroppedWidth shall be equal to nnpfc_pic_width_in_luma_samples*inpPatchWidth, and outPatchHeight*CroppedHeight shall be equal to nnpfc_pic_height_in_luma_samples*inpPatchHeight.

[0210] nnpfc_padding_type may indicate a padding process when referring to a sample position outside the cropped decoded output picture boundary, as described in Table 14. The value of nnpfc_padding_type should be within the range of 0 to 15.

[0211] [Table 14]

[0212] nnpfc_padding_type describe 0 Zero padding 1 Copy Fill 2 Reflection Fill 3 Surround Fill 4 Fixed padding 5…15 Reserve

[0213] nnpfc_luma_padding_val may indicate a luma value to be used for padding when the value of nnpfc_padding_type is 4.

[0214] nnpfc_cb_padding_val may indicate a Cb value to be used for padding when the value of nnpfc_padding_type is 4.

[0215] nnpfc_cr_padding_val may indicate a Cr value to be used for padding when a value of nnpfc_padding_type is 4.

[0216] The function InpSampleVal(y,x,picHeight,picWidth,CroppedPic) whose input includes the vertical sample position y, the horizontal sample position x, the picture height picHeight, the picture width picWidth and the sample array CroppedPic can return the sampleVal value derived as shown in Table 15.

[0217] For the input to the InpSampleVal() function, the vertical position can be listed before the horizontal position to ensure compatibility with the input tensor rules of some inference engines.

[0218] [Table 15]

[0219]

[0220] The processing in Table 16 may be used to generate a filtered picture by filtering the cropped decoded output picture in a patch-based manner using a post-processing filter PostProcessingFilter(), and the filtered picture may include a Y sample array FilteredYPic, a Cb sample array FilteredCbPic, and a Cr sample array FilteredCrPic, as indicated by nnpfc_out_order_idc.

[0221] [Table 16]

[0222]

[0223] A value of 1 for nnpfc_complexity_info_present_flag may indicate the presence of one or more syntax elements indicating the complexity of the post-processing filter associated with nnpfc_id. A value of 0 for nnpfc_complexity_info_present_flag may indicate the absence of syntax elements indicating the complexity of the post-processing filter associated with nnpfc_id.

[0224] A nnpfc_parameter_type_idc value of 0 may indicate that the neural network uses only integer parameters. A nnpfc_parameter_type_flag value of 1 may indicate that the neural network can use floating-point or integer parameters. A nnpfc_parameter_type_idc value of 2 may indicate that the neural network uses only binary parameters. A nnpfc_parameter_type_idc value of 3 may be reserved for future use and shall not be present in the bitstream. A decoder shall ignore NNPFC SEI messages with a nnpfc_parameter_type_idc value of 3.

[0225] The values ​​of 0, 1, 2, and 3 for nnpfc_log2_parameter_bit_length_minus3 may respectively indicate that the neural network does not use parameters with bit lengths greater than 8, 16, 32, and 64. When nnpfc_parameter_type_idc is present and nnpfc_log2_parameter_bit_length_minus3 is not present, the neural network may not use parameters with bit lengths greater than 1.

[0226] nnpfc_num_parameters_idc may be a power of 2048 to indicate the maximum number of neural network parameters used for the post-processing filter. A nnpfc_num_parameters_idc value of 0 may indicate that the maximum number of neural network parameters is unknown. nnpfc_num_parameters_idc values ​​shall be in the range of 0 to 52. nnpfc_num_parameters_idc values ​​greater than 52 shall not be present in the bitstream. Decoders shall ignore NNPFC SEI messages with nnpfc_num_parameters_idc values ​​greater than 52.

[0227] When the value of nnpfc_num_parameters_idc is greater than 0, the variable maxNumParameters may be derived as shown in Equation 3.

[0228] [Formula 3]

[0229] maxNumParameters=(2048<<nnpfc_num_parameters_idc)-1

[0230] The number of neural network parameters used for post-processing filters should be constrained to be less than or equal to maxNumParameters.

[0231] A value of nnpfc_num_kmac_operations_idc greater than 0 may indicate that the maximum number of multiply-accumulate operations per sample for the post-processing filter is less than or equal to nnpfc_num_kmac_operations_idc*1000. A value of nnpfc_num_kmac_operations_idc equal to 0 may indicate that the maximum number of multiply-accumulate operations of the network is unknown. The value of nnpfc_num_kmac_operations_idc should be between 0 and 2. 32 In the range of -1.

[0232] A value of nnpfc_total_kilobyte_size greater than 0 may indicate that the total size (in kilobytes) required to store the uncompressed parameters of the neural network is unknown. The total size (in bits) may be greater than or equal to the sum of the number of bits used to store each parameter. nnpfc_total_kilobyte_size may be the result of dividing the total size (in bits) by 8000 and rounding. A value of 0 for nnpfc_total_kilobyte_size may indicate that the total size required to store the parameters of the neural network is unknown. The value of nnpfc_total_kilobyte_size should be between 0 and 2. 32 In the range of -1.

[0233] In the bitstream, nnpfc_reserved_zero_bit_b shall be equal to 0. The decoder shall ignore NNPFC SEI messages with nnpfc_reserved_zero_bit_b not equal to 0.

[0234] nnpfc_payload_byte[i] may contain the i-th byte of the bitstream. The byte sequence nnpfc_payload_byte[i] for all present values ​​of i shall be a complete bitstream conforming to ISO / IEC 15938-17.

[0235] Neural Network Post-Filter Activation (NNPFA)

[0236] The syntax structure of NNPFA is shown in Table 17.

[0237] [Table 17]

[0238]

[0239] The NNPFA syntax structure in Table 17 may be signaled in the form of a SEI message. The SEI message that signals the NNPFA syntax structure in Table 17 may be referred to as an NNPFA SEI message.

[0240] The NNPFA SEI message may activate or deactivate the possible use of the target neural network post-processing filter identified by nnpfa_target_id for post-processing filtering of a set of pictures.

[0241] When the post-processing filters are used for different purposes or filter different color components, there may be multiple NNPFA SEI messages for the same picture.

[0242] nnpfa_target_id may indicate a target neural network post-processing filter associated with the current picture and specified by one or more NNPFC SEI messages having the same nnpfc_id as NNPFA_target_id.

[0243] The value of nnpfa_target_id should be between 0 and 2. 32 -2. 256 to 511 and 2 31 to 2 32 nnpfa_target_id values ​​in the range of -2 are reserved for future use. Decoders should ignore nnpfa_target_id values ​​in the range of 256 to 511 or 2 31 to 2 32 NNPFA SEI message in the -2 range.

[0244] The NNPFA SEI message with a specific value of nnpfa_target_id shall not be present in the current picture unit (PU) unless one or both of the following conditions are true: Here, a picture unit (PU) may be a set of NAL units including the VCL NAL units of a coded picture and the associated non-VCL NAL units.

[0245] - Within the current CLVS, there is an NNPFC SEI message with the same nnpfc_id as the specific value of nnpfa_target_id present in a PU present before the current PU in decoding order.

[0246] - There is an NNPFCSEI message for the current PU with the same nnpfc_id as the specific value of nnpfa_target_id.

[0247] When a PU includes both an NNPFC SEI message with a specific value of nnpfc_id and an NNPFA SEI message with nnpfa_target_id equal to the specific value of nnpfc_id, the NNPFC SEI message shall precede the NNPFC SEI message in decoding order.

[0248] A value of 1 for nnpfa_cancel_flag may indicate that the persistence of the target neural network post-processing filter set by any previous NNPFA SEI message with the same nnpfa_target_id as the current SEI message is canceled. In other words, the target neural network post-processing filter should no longer be used unless it is activated by an NNPFA SEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 0. A value of 0 for nnpfa_cancel_flag may indicate that nnpfa_persistence_flag is followed.

[0249] nnpfa_persistence_flag may indicate the persistence of the target neural network post-processing filter for the current layer. A nnpfa_persistence_flag value of 0 may indicate that the target neural network post-processing filter can be used for post-processing filtering only for the current picture. A nnpfa_persistence_flag value of 1 may indicate that the target neural network post-processing filter can be used for post-processing filtering for the current picture and all subsequent pictures in the current layer in output order until one or more of the following conditions are true:

[0250] - Start a new CLVS for the current layer.

[0251] - End of bitstream.

[0252] - Pictures in the current layer that are associated with an NNPFA SEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1 are output after the current picture in output order.

[0253] The target neural network post-processing filter shall not be applied to subsequent pictures in the current layer that are associated with an NNPFA SEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1.

[0254] Post-Filter Tips

[0255] The syntax structure for post-filter hint is shown in Table 18.

[0256] [Table 18]

[0257]

[0258] The post-filter hint syntax structure in Table 18 may be signaled in the form of a SEI message. The SEI message that signals the post-filter hint syntax structure in Table 18 may be referred to as a post-filter hint SEI message.

[0259] The post-filter hint SEI message may provide relevant information about the design of a post-filter, or post-filter coefficients, to potentially use the decoded and output picture set for post-processing to achieve enhanced display quality.

[0260] A filter_hint_cancel_flag value of 1 may indicate that the SEI message cancels the persistence of the previous post-filter hint SEI message to be applied to the current layer in output order. A filter_hint_cancel_flag value of 0 may indicate that the post-filter hint information is followed.

[0261] filter_hint_persistence_flag may indicate the persistence of the post-filter hint SEI message for the current layer. A filter_hint_persistence_flag value of 0 may indicate that the post-filter hint is applicable only to the current decoded picture. A filter_hint_persistence_flag value of 1 may indicate that the post-filter hint SEI message is applicable to the current decoded picture and remains valid for all subsequent pictures in the current layer in output order until one or more of the following conditions are true:

[0262] - Start a new CLVS for the current layer.

[0263] - End of bitstream.

[0264] - The pictures of the AU associated with the post-filter hint SEI message in the current layer are output after the current picture in output order.

[0265] filter_hint_size_y may indicate the vertical size of the filter coefficients or associated array. The value of filter_hint_size_y should be in the range of 1 to 15.

[0266] filter_hint_size_x may indicate the horizontal size of the filter coefficients or associated array. The value of filter_hint_size_x should be in the range of 1 to 15.

[0267] filter_hint_type may indicate the type of filter hint being sent, as shown in Table 19. The value of filter_hint_type shall be in the range of 0 to 2. A value of filter_hint_type equal to 3 shall not be present in the bitstream. The decoder shall ignore post-filter hint SEI messages with filter_hint_type equal to 3.

[0268] [Table 19]

[0269] value describe 0 2D-FIR filter coefficients 1 1D-FIR filter coefficients 2 Cross-correlation matrix

[0270] A filter_hint_chroma_coeff_present_flag value of 1 may indicate the presence of filter coefficients for chroma. A filter_hint_chroma_coeff_present_flag value of 0 may indicate the absence of filter coefficients for chroma.

[0271] filter_hint_value[cIdx][cy][cx] may represent the cross-correlation matrix element or filter coefficient between the original signal and the decoded signal with 16-bit precision. The value of filter_hint_value[cIdx][cy][cx] shall be between -2 31 +1 to 2 31 The variable cIdx may indicate the corresponding color component, cy may indicate the counter in the vertical direction, and cx may indicate the counter in the horizontal direction. Depending on the value of filter_hint_type, the following may apply:

[0272] - When the value of filter_hint_type is 0, the coefficients of a two-dimensional finite impulse response (FIR) filter of size filter_hint_size_y × filter_hint_size_x are sent.

[0273] Otherwise, when the value of filter_hint_type is 1, the filter coefficients of two one-dimensional FIR filters are sent. In this case, the value of filter_hint_size_y should be 2. The index cy equal to 0 indicates the filter coefficients of the horizontal filter, and cy equal to 1 indicates the filter coefficients of the vertical filter. In the filtering process, the horizontal filter may be applied first, and then the result may be filtered by the vertical filter.

[0274] - Otherwise (ie, when the value of filter_hint_type is 2), the hint sent may represent the cross-correlation matrix between the original signal s and the decoded signal s'.

[0275] The normalized cross - correlation matrix of the color component identified by cIdx (with size filter_hint_size_y×filter_hint_size_x) is defined as in Equation 4.

[0276] [Equation 4]

[0277]

[0278] In Equation 4, s indicates the sample array of the color component cIdx of the original image, s' indicates the corresponding array of the decoded picture, h indicates the vertical height of the relevant color component, w indicates the horizontal width of the relevant color component, and bitDepth indicates the bit - depth of the color component. Additionally, OffsetY is equal to (filter_hint_size_y>>1), OffsetX is equal to (filter_hint_size_x>>1), the range of cy is 0 <= cy < filter_hint_size_y, and the range of cx is 0 <= cx < filter_hint_size_x.

[0279] The decoder can derive the Wiener post - filter from the cross - correlation matrix between the original signal and the decoded signal and the auto - correlation matrix of the decoded signal.

[0280] Problems with prior art

[0281] The Neural Network Post - Filter Characteristics (NNPFC) SEI message includes the flag nnpfc_formatting_and_puropose_flag to control the presence of syntax elements associated with a specific description, and the specific description includes descriptions related to the complexity information, format, and purpose of the filter. The presence of the flag is specified as follows:

[0282] ------------Fragment start------------

[0283] When the specific SEI message is the first NNPFC SEI message with a specific nnpfc_id value in the current CLVS in decoding order, the value of nnpfc_formatting_and_puropose_flag shall be 1. On the other hand, when the specific SEI message is not the first NNPFC SEI message with a specific nnpfc_id value in the current CLVS in decoding order, the value of nnpfc_formatting_and_puropose_flag shall be 0.

[0284] ------------Fragment end------------

[0285] The above constraint may specify that the signaling of syntax elements describing the format, usage, and complexity is allowed only for SEIs that include base neural network post-processing filters, while the same method is not allowed in the case of update filters. However, the meaning of the above constraint may be problematic when the signaling of format, usage, and complexity is absent due to at least two possibilities:

[0286] 1. In the case of an NNPFC message in which the format, usage and complexity information is not present, the corresponding information is inferred from the NNPFC SEI message including the underlying neural network post-processing filters.

[0287] 2. In the case of an NNPFC message in which the format, usage and complexity information do not exist, the corresponding information is unknown.

[0288] When the flag (ie, nnpfc_formatting_and_purpose_flag) is 0, the meaning of the corresponding message shall be specified in the NNPFC SEI message semantics.

[0289] Implementation Method

[0290] The present disclosure proposes implementations that can solve the above-mentioned problems.

[0291] Furthermore, in describing the present disclosure, the terms post-processing filter, post-filter, and post-processing filter may be used as the same meaning.In addition, the following embodiments may be used alone or in combination of two or more.

[0292] Hereinafter, NNPFC refers to the NNPFC syntax structure of Table 1 and Table 2, which may be signaled in the form of an SEI message. In this case, NNPFC may be the NNPFC SEI message. NNPFA refers to the NNPFA syntax structure of Table 17, which may be signaled in the form of an SEI message. In this case, NNPFA may be the NNPFA SEI message. Postfilter hint refers to the postfilter hint syntax structure of Table 18, which may be signaled in the form of an SEI message. In this case, postfilter hint may be the postfilter hint SEI message.

[0293] 1. The signaling of format and usage information only exists in the NNPFC SEI message including the basic neural network post-processing filter

[0294] 2. In the NNPFC SEI message including the updated neural network post-processing filter, the format and usage information is inferred to be the same as the information signaled in the related NNPFC SEI message including the base neural network post-processing filter (i.e., the NNPFC SEI message with the same nnpfc_id)

[0295] 3. Alternatively, when format and usage information does not exist, information for the corresponding information is designated as unknown or externally provided.

[0296] 4. Alternatively, the following may apply:

[0297] a. In the case of NNPFC SEI messages including the underlying neural network post-processing filter, specify the signaling of the presence format and usage information

[0298] b. In the case of NNPFC SEI messages that include updates to neural network post-processing filters, signaling of format and usage information may be present. In other words, signaling may be optional, and when absent, it is specified to be inferred to be the same as what was signaled in messages related to the base filters.

[0299] Implementation Method 1

[0300] This embodiment provides a detailed description of items 1 and 2 of the above embodiment. This embodiment may be based on a standard (eg, VVC) document.

[0301] This disclosure proposes the following updates:

[0302] Neural network post-filter characteristics SEI

[0303] Table 20 below is an example of modification to a certain element of the NNPFC SEI syntax.

[0304] [Table 20]

[0305]

[0306]

[0307] The following is an example of modification of the semantics related to the above modification.

[0308] The Neural Network Post-Filter Characteristics (NNPFC) SEI message may specify a neural network that may be used as a post-processing filter. The use of a post-processing filter specified for a particular filter may be indicated by a Neural Network Post-Filter Activation SEI message.

[0309] When the specific SEI message is not the first NNPFC SEI message with a specific nnpfc_id value in decoding order within the current CLVS, the post-processing filter PostProcessingFilter() may be obtained by applying the update information defined by the SEI message to the basic post-processing filter.

[0310] NOTE—Update information is not cumulative, and each update is applicable to the base post-processing filter, which is the post-processing filter specified by the first NNPFC SEI message with a specific nnpfc_id value in decoding order within the current CLVS.

[0311] In a bitstream conforming to a revision of this document, nnpfc_reserved_zero_bit_a may be constrained to be 0. A decoding device, in other words, a decoder may be constrained to ignore NNPFC SEI messages in which the value of nnpfc_reserved_zero_bit_a is not 0.

[0312] nnpfc_tag_uri may include a tag URI for the syntax and semantics specified in IETF RFC 4151 that identifies the format and related information about a neural network used as a base post-processing filter, or an update to the base post-processing filter specified by nnpfc_uri with the same nnpfc_id value.

[0313] NOTE - nnpfc_tag_uri uniquely identifies the format of neural network data specified by nnrpf_uri without the need for a central registration authority.

[0314] When nnpfc_tag_uri is "tag:iso.org,2023:15938-17", it may indicate that the neural network data identified by nnpfc_uri complies with ISO / IEC15938-17.

[0315] nnpfc_uri may comprise a URI specific to the syntax and semantics specified in IETF Internet Standard 66 that identifies a neural network used as a base post-processing filter or an update to a base post-processing filter having the same nnpfc_id value.

[0316] When the value of formatting_and_purpose_flag is 1, it may specify that syntax elements related to filter purpose, input format, output format, and complexity are present. When the value of nnpfc_formatting_and_purpose_flag is 0, it may specify that syntax elements related to filter purpose, input format, output format, and complexity are not present.

[0317] When the SEI message is the first NNPFC SEI message with a specific nnpfc_id value in decoding order within the current CLVS, the value of nnpfc_formatting_and_purpose_flag may be restricted to 1. On the other hand, when the SEI message is not the first NNPFC SEI message with a specific nnpfc_id value in decoding order within the current CLVS, the value of nnpfc_formatting_and_purpose_flag may be restricted to 0. When the value of nnpfc_formatting_and_purpose_flag is 0, the values ​​of syntax elements related to filter purpose, input format, output format, and complexity may each be inferred as the values ​​of corresponding syntax elements in the NNPFC SEI message including the neural network post-processing filter with the same nnpfc_id.

[0318] Implementation Method 2

[0319] This embodiment provides a detailed description of items 1 and 2 of the above embodiment. This embodiment may be based on a standard (eg, VVC) document.

[0320] This disclosure proposes the following updates:

[0321] Neural network post-filter characteristics SEI

[0322] An example of modification of a certain element of the NNPFC SEI syntax is shown in Table 20 above.

[0323] The following are examples of modifications of semantics related to the modification examples in Table 20.

[0324] The Neural Network Post-Filter Characteristics (NNPFC) SEI message may specify a neural network that may be used as a post-processing filter. The use of a post-processing filter specified for a particular filter may be indicated by a Neural Network Post-Filter Activation SEI message.

[0325] When the specific SEI message is not the first NNPFC SEI message with a specific nnpfc_id value in decoding order within the current CLVS, the post-processing filter PostProcessingFilter() may be obtained by applying the update information defined by the SEI message to the basic post-processing filter.

[0326] NOTE—Update information is not cumulative, and each update is applicable to the base post-processing filter, which is the post-processing filter specified by the first NNPFC SEI message with a specific nnpfc_id value in decoding order within the current CLVS.

[0327] In a bitstream conforming to a revision of this document, the value of nnpfc_reserved_zero_bit_a may be restricted to 0. A decoding device, in other words, a decoder may be restricted to ignore NNPFC SEI messages in which the value of nnpfc_reserved_zero_bit_a is not 0.

[0328] nnpfc_tag_uri may include a tag URI for the syntax and semantics specified in IETF RFC 4151 that identifies the format and related information about a neural network used as a base post-processing filter, or an update to the base post-processing filter specified by nnpfc_uri with the same nnpfc_id value.

[0329] NOTE - nnpfc_tag_uri uniquely identifies the format of neural network data identified by nnrpf_uri without the need for a central registry.

[0330] When nnpfc_tag_uri is "tag:iso.org,2023:15938-17", it may indicate that the neural network data identified by nnpfc_uri complies with ISO / IEC15938-17.

[0331] nnpfc_uri may comprise a URI specific to the syntax and semantics specified in IETF Internet Standard 66 that identifies a neural network used as a base post-processing filter or an update to a base post-processing filter having the same nnpfc_id value.

[0332] When the value of formatting_and_purpose_flag is 1, it may specify that syntax elements related to filter purpose, input format, output format, and complexity are present. When the value of nnpfc_formatting_and_purpose_flag is 0, it may specify that syntax elements related to filter purpose, input format, output format, and complexity are not present.

[0333] When the SEI message is an NNPFC SEI message having a specific nnpfc_id value within the current CLVS and including a basic neural network post-processing filter, the value of nnpfc_formatting_and_purpose_flag may be restricted to 1. On the other hand, when the SEI message is not an NNPFC SEI message having a specific nnpfc_id value within the current CLVS and including a basic neural network post-processing filter, the value of nnpfc_formatting_and_purpose_flag may be restricted to 0. When the value of nnpfc_formatting_and_purpose_flag is 0, the values ​​of syntax elements related to filter purpose, input format, output format, and complexity may each be inferred as the values ​​of corresponding syntax elements in the NNPFC SEI message including the neural network post-processing filter having the same nnpfc_id.

[0334] Implementation 3

[0335] This embodiment provides a detailed description of items 1 and 2 of the above embodiment. This embodiment may be based on a standard (eg, VVC) document.

[0336] This disclosure proposes the following updates:

[0337] Neural network post-filter characteristics SEI

[0338] An example of modification of certain elements of the NNPFC SEI syntax is shown in Table 20 above.

[0339] The following are examples of modifications of semantics related to the modification examples in Table 20.

[0340] The Neural Network Post-Filter Characteristics (NNPFC) SEI message may specify a neural network that may be used as a post-processing filter. The use of a post-processing filter specified for a particular filter may be indicated by a Neural Network Post-Filter Activation SEI message.

[0341] When the specific SEI message is not the first NNPFC SEI message with a specific nnpfc_id value in decoding order within the current CLVS, the post-processing filter PostProcessingFilter() may be obtained by applying the update information defined by the SEI message to the basic post-processing filter.

[0342] NOTE—Update information is not cumulative, and each update is applicable to the base post-processing filter, which is the post-processing filter specified by the first NNPFC SEI message with a specific nnpfc_id value in decoding order within the current CLVS.

[0343] In a bitstream conforming to a revision of this document, the value of nnpfc_reserved_zero_bit_a may be restricted to 0. A decoding device, in other words, a decoder may be restricted to ignore NNPFC SEI messages in which the value of nnpfc_reserved_zero_bit_a is not 0.

[0344] nnpfc_tag_uri may include a tag URI for the syntax and semantics specified in IETF RFC 4151 that identifies the format and related information about a neural network used as a base post-processing filter, or an update to the base post-processing filter specified by nnpfc_uri with the same nnpfc_id value.

[0345] NOTE - nnpfc_tag_uri uniquely identifies the format of neural network data identified by nnrpf_uri without the need for a central registry.

[0346] When nnpfc_tag_uri is "tag:iso.org,2023:15938-17", it may indicate that the neural network data identified by nnpfc_uri complies with ISO / IEC15938-17.

[0347] nnpfc_uri may comprise a URI specific to the syntax and semantics specified in IETF Internet Standard 66 that identifies a neural network used as a base post-processing filter or an update to a base post-processing filter having the same nnpfc_id value.

[0348] When the value of formatting_and_purpose_flag is 1, it may specify that syntax elements related to filter purpose, input format, output format, and complexity are present. When the value of nnpfc_formatting_and_purpose_flag is 0, it may specify that syntax elements related to filter purpose, input format, output format, and complexity are not present.

[0349] When the SEI message is an NNPFC SEI message having a specific nnpfc_id value within the current CLVS and including a basic neural network post-processing filter, the value of nnpfc_formatting_and_purpose_flag may be restricted to 1. On the other hand, when the SEI message is not an NNPFC SEI message having a specific nnpfc_id value within the current CLVS and including a basic neural network post-processing filter, the value of nnpfc_formatting_and_purpose_flag may be restricted to 0. When the value of nnpfc_formatting_and_purpose_flag is 0, the values ​​of syntax elements related to filter purpose, input format, output format, and complexity may be unknown or may be provided by other external means.

[0350] Implementation 4

[0351] This embodiment provides a detailed description of items 1 and 2 of the above embodiment. This embodiment may be based on a standard (eg, VVC) document.

[0352] This disclosure proposes the following updates:

[0353] Neural network post-filter characteristics SEI

[0354] An example of modification of a certain element of the NNPFC SEI syntax is shown in Table 20 above.

[0355] The following are examples of modifications of semantics related to the modification examples in Table 20.

[0356] The Neural Network Post-Filter Characteristics (NNPFC) SEI message may specify a neural network that may be used as a post-processing filter. The use of a post-processing filter specified for a particular filter may be indicated by a Neural Network Post-Filter Activation SEI message.

[0357] When the specific SEI message is not the first NNPFC SEI message with a specific nnpfc_id value in decoding order within the current CLVS, the post-processing filter PostProcessingFilter() may be obtained by applying the update information defined by the SEI message to the basic post-processing filter.

[0358] NOTE—Update information is not cumulative, and each update is applicable to the base post-processing filter, which is the post-processing filter specified by the first NNPFC SEI message with a specific nnpfc_id value in decoding order within the current CLVS.

[0359] In a bitstream conforming to a revision of this document, the value of nnpfc_reserved_zero_bit_a may be restricted to 0. A decoding device, in other words, a decoder may be restricted to ignore NNPFC SEI messages in which the value of nnpfc_reserved_zero_bit_a is not 0.

[0360] nnpfc_tag_uri may include a tag URI for the syntax and semantics specified in IETF RFC 4151 that identifies the format and related information about a neural network used as a base post-processing filter, or an update to the base post-processing filter specified by nnpfc_uri with the same nnpfc_id value.

[0361] NOTE - nnpfc_tag_uri uniquely identifies the format of neural network data identified by nnrpf_uri without the need for a central registry.

[0362] When nnpfc_tag_uri is "tag:iso.org,2023:15938-17", it may indicate that the neural network data identified by nnpfc_uri complies with ISO / IEC15938-17.

[0363] nnpfc_uri may comprise a URI specific to the syntax and semantics specified in IETF Internet Standard 66 that identifies a neural network used as a base post-processing filter or an update to a base post-processing filter having the same nnpfc_id value.

[0364] When the value of formatting_and_purpose_flag is 1, it may specify that syntax elements related to filter purpose, input format, output format, and complexity are present. When the value of nnpfc_formatting_and_purpose_flag is 0, it may specify that syntax elements related to filter purpose, input format, output format, and complexity are not present.

[0365] When the SEI message is an NNPFC SEI message with a specific nnpfc_id value within the current CLVS and including a base neural network post-processing filter, the value of nnpfc_formatting_and_purpose_flag may be restricted to 1. On the other hand, when the SEI message is not an NNPFC SEI message with a specific nnpfc_id value within the current CLVS and including a base neural network post-processing filter, the value of nnpfc_formatting_and_purpose_flag may be 0. When the value of nnpfc_formatting_and_purpose_flag is 0, the values ​​of syntax elements related to filter purpose, input format, output format, and complexity may each be inferred as the value of the corresponding syntax element in the previous NNPFC SEI message with the same nnpfc_id in decoding order.

[0366] Hereinafter, an image encoding method and an image decoding method according to various embodiments of the present disclosure will be described. Figure 5 The image encoding method may be performed by the image encoding apparatus 100, and Figure 6 The image decoding method may be performed by the image decoding device 200. In addition, Figure 5 and Figure 6 The image encoding and decoding methods may be based on the above-mentioned embodiments.

[0367] Reference Figure 5 , at least one neural network that can be used as a post-processing filter can be determined, and information about the determined neural network can be encoded into at least one NNPFC SEI message S510.

[0368] Whether to activate a target neural network post-processing filter applicable to the current picture may be determined, and information about the determined target neural network post-processing filter may be encoded into an NNPFA SEI message S520. The process for determining activation of the target neural network post-processing filter S520 may include a process for determining the target neural network post-processing filter, a process for determining whether to cancel persistence of the target neural network post-processing filter, and a process for determining whether to maintain persistence of the target neural network post-processing filter.

[0369] Relevant information about the design of the post-processing filter or the post-processing filter coefficients may be encoded into a post-filter hint SEI message S530. The NNPFC SEI message, the NNPFA SEI message and / or the post-filter hint SEI message may be included in the SEI message for the neural network post filter (NNFP).

[0370] When the SEI message for NNPF is applied to the current picture in the image decoding device, the target neural network post-processing filter can be determined or specified through various embodiments of the present application.

[0371] For example, when an NNPFC SEI message is present, the information that may be present in the message may be determined based on whether the NNPFC SEI includes a base neural network post-processing filter. The information present in the NNPFC SEI message may be determined based on whether it has the same value as nnpfc_id. In this case, within the current CLVS, the NNPFC SEI message that includes the base neural network post-processing filter may be the first NNPFC SEI message in decoding order and may have a specific nnpfc_id value. In addition, the information that may be included in the NNPFC SEI message may include at least one of filter usage, input format, output format (collectively referred to as format information), and complexity information. Furthermore, for example, specific information (e.g., format information or usage information) may be encoded and signaled only when the NNPFC SEI message includes a base neural network post-processing filter. On the other hand, when the NNPFC SEI message includes information for updating the neural network post-processing filter (i.e., updating the neural network post-processing filter), the format and purpose information may not be encoded and may be inferred in the decoder to be the same as the specific information (e.g., format information or purpose information, etc.) included in the NNPFC SEI message including the basic neural network post-processing filter, or may be encoded in a specific manner to be inferred to be the same. As another example, when specific information (e.g., format information or purpose information, etc.) does not exist, the information may not be noticed by the decoder, etc. In other words, it may not be explicitly encoded and may be provided in an external manner. For example, the case where specific information does not exist may include a case where information indicating whether specific information exists (e.g., nnpfc_formatting_and_purpose_flag, etc.) is determined to be a specific value (e.g., 0). Furthermore, as another example, when the NNPFC SEI message includes a basic neural network post-processing filter, specific information (e.g., format information or purpose information, etc.) may be necessarily signaled, whereas in other cases (e.g., when the NNPFC SEI message includes information for updating the neural network post-processing filter), it may not be encoded and may be inferred to be the same as the information included in the NNPFC SEI message including the basic neural network post-processing filter, or may be encoded to be inferred to be the same. As an example, when information indicating whether specific information exists (e.g., nnpfc_formatting_and_purpose_flag, etc.) is a specific value (e.g., 0), the specific information may be inferred to be the same value as the value of the corresponding syntax element in the preceding NNPFC SEI message having the same nnpfc_id value in decoding order without being encoded, or may be encoded to be inferred to be the same value.

[0372] Reference Figure 6, the SEI message for NNFP to be applied to the current picture can be obtained from the bitstream. The SEI message for NNFP may include an NNPFC SEI message, an NNPFA SEI message and / or a post-filter hint SEI message.

[0373] When the SEI message for NNFP is applied to the current picture, at least one neural network that can be used as a post-processing filter can be determined based on at least one NNPFC SEI message included in the SEI message for NNFP S610.

[0374] Based on at least one NNPFA SEI message obtained from the bitstream, it is possible to determine whether to activate a target neural network post-processing filter applicable to the current picture S620. The process for determining activation of the target neural network post-processing filter S620 may include a process of determining the target neural network post-processing filter, a process of determining whether to cancel the persistence of the target neural network post-processing filter, and a process of determining whether to maintain the persistence of the target neural network post-processing filter.

[0375] When the target neural network post-processing filter is activated, the target neural network post-processing filter may be applied to the current picture S630.

[0376] Various embodiments of the present disclosure may be used to determine or specify a target neural network post-processing filter.

[0377] For example, when an NNPFC SEI message is present, the information that may be present in the message may be determined based on whether the NNPFC SEI message includes a base neural network post-processing filter. In other words, the information present in the NNPFC SEI message may be determined based on whether it has the same value as nnpfc_id. In this case, within the current CLVS, the NNPFC SEI message that includes the base neural network post-processing filter may be the first NNPFC SEI message in decoding order and may have a specific nnpfc_id value. Furthermore, the information that may be included in the NNPFC SEI message may include at least one of filter usage, input format, output format (collectively referred to as format information), and complexity information. Furthermore, for example, specific information (e.g., format information or usage information) may be signaled only when the NNPFC SEI message includes the base neural network post-processing filter. On the other hand, when the NNPFC SEI message includes information for updating the neural network post-processing filter (i.e., updating the neural network post-processing filter), the format and usage information may be inferred to be the same as the specific information (e.g., format information or usage information) included in the NNPFC SEI message that includes the base neural network post-processing filter. As another example, when specific information (e.g., format information or purpose information) is absent, the information may be unknown, may not be decoded, or may be provided externally. For example, the absence of specific information may include a case where information indicating the presence of specific information (e.g., nnpfc_formatting_and_purpose_flag) is a specific value (e.g., 0). Furthermore, as another example, when an NNPFC SEI message includes a base neural network post-processing filter, specific information (e.g., format information or purpose information) may be necessarily signaled, while in other cases (e.g., when the NNPFC SEI message includes information for updating the neural network post-processing filter), the information may be inferred to be the same as the information included in the NNPFC SEI message including the base neural network post-processing filter. For example, when information indicating the presence of specific information (e.g., nnpfc_formatting_and_purpose_flag) is a specific value (e.g., 0), the specific information may be inferred to be the same value as the corresponding syntax element in a preceding NNPFC SEI message having the same nnpfc_id value in decoding order.

[0378] According to the present disclosure, the meaning of the NNPFC SEI message may be clearly defined to improve coding efficiency.

[0379] Figure 7 An exemplary schematic diagram showing a content streaming system to which embodiments of the present disclosure may be applied is shown.

[0380] like Figure 7 As shown, a content streaming system to which the embodiments of the present disclosure are applied may generally include an encoding server, a streaming server, a Web server, a media storage, a user device, and a multimedia input device.

[0381] The encoding server compresses the content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data, generates a bitstream, and sends it to the streaming server. As another example, when a multimedia input device such as a smartphone, camera, or camcorder directly generates a bitstream, the encoding server can be omitted.

[0382] A bitstream may be generated by applying the video encoding method and / or the image encoding apparatus according to the embodiments of the present disclosure, and the streaming server may temporarily store the bitstream during a process of transmitting or receiving the bitstream.

[0383] The streaming server can transmit multimedia data to a user device via a web server based on user requests, and the web server can act as an intermediary to notify users of available services. When a user requests a desired service from the web server, the web server can transmit the request to the streaming server, and the streaming server can transmit the multimedia data to the user. In this case, the content streaming system can include a separate control server, and in this case, the control server can be used to control the command / response exchanges between devices within the content streaming system.

[0384] The streaming server may receive content from a media storage and / or encoding server. For example, when receiving content from an encoding server, the content may be received in real time. In this case, in order to provide a seamless streaming service, the streaming server may store the bitstream for a period of time.

[0385] Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet PCs, tablet PCs, ultrabooks, wearable devices (i.e., smart watches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, and digital signage.

[0386] Each server within the content streaming system may operate as a distributed server, in which case data received by each server may be processed in a distributed manner.

[0387] The scope of the present disclosure includes software or machine-executable instructions (i.e., operating systems, applications, firmware, programs, etc.) that enable methods according to various embodiments to be performed on a device or computer, as well as non-transitory computer-readable media in which such software or instructions are stored and can be executed on a device or computer.

[0388] Industrial Applicability

[0389] The embodiments of the present disclosure may be used to encode / decode images.

Claims

1. An image decoding method performed by an image decoding device, the image decoding method comprising the following steps: Obtaining a supplemental enhancement information (SEI) message for a neural network post filter (NNPF) to be applied to the current picture; determining, based on the SEI message for the NNPF being applied to the current picture, at least one neural network capable of being used as a post-processing filter based on at least one neural network post-filter characteristic (NNPFC) SEI message included in the SEI message for the NNPF; and determining whether to activate a target neural network post-processing filter applicable to the current picture based on at least one neural network post-filter activation NNPF SEI message included in the SEI message for the NNPF, The format information and usage information included in the NNPFC SEI message are determined based on the fact that the NNPFC SEI message includes a basic neural network post-processing filter.

2. An image encoding method performed by an image encoding device, the image encoding method comprising the following steps: encoding at least one neural network operable as a post-processing filter into at least one neural network post-filter characteristic NNPFC supplemental enhancement information SEI message; as well as Whether to activate the target neural network post-processing filter that can be applied to the current picture is encoded into at least one neural network post-filter activation NNPFA SEI message, In which, the SEI message for the neural network post-filter NNPF is applied to the current picture in the image decoding device, and the format information and usage information included in the NNPFC SEI message are determined based on the NNPFC SEI message including the basic neural network post-processing filter.

3. A method for transmitting a bit stream generated by an image encoding method, the image encoding method comprising the steps of: encoding at least one neural network operable as a post-processing filter into at least one neural network post-filter characteristic NNPFC supplemental enhancement information SEI message; as well as Whether to activate the target neural network post-processing filter that can be applied to the current picture is encoded into at least one neural network post-filter activation NNPFA SEI message, In which, the SEI message for the neural network post-filter NNPF is applied to the current picture at the image decoding device, and the format information and usage information included in the NNPFC SEI message are determined based on the NNPFC SEI message including the basic neural network post-processing filter.

4. A computer-readable recording medium storing a bit stream generated by an image encoding method, the image encoding method comprising the steps of: encoding at least one neural network operable as a post-processing filter into at least one neural network post-filter characteristic NNPFC supplemental enhancement information SEI message; as well as Whether to activate the target neural network post-processing filter that can be applied to the current picture is encoded into at least one neural network post-filter activation NNPFA SEI message, In which, the SEI message for the neural network post-filter NNPF is applied to the current picture at the image decoding device, and the format information and usage information included in the NNPFC SEI message are determined based on the NNPFC SEI message including the basic neural network post-processing filter.