Image encoding / decoding method, method for transmitting bit stream, and recording medium for storing bit stream

By receiving and reconstructing the filter-related information after the neural network, the image encoding and decoding process is optimized, the problem of high-resolution image transmission and storage costs is solved, the encoding and decoding efficiency is improved, and the filter information signal transmission and repeated message processing of the basic filter are realized.

CN120513638APending Publication Date: 2025-08-19LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380090074.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-18
Filing Date
2023-12-28
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

In the encoding and decoding process of high resolution and high quality images, the problem of increasing the amount of information leading to high transmission and storage costs, and the signaling efficiency of the filter attribute information is low, making it difficult to efficiently inform the existence of the filter and the information of the basic filter.

Method used

By receiving and reconstructing supplementary enhancement information related to the post-neural network filter, a post-neural network filter characteristic message including filter identification and attribute presence information is generated and sent, the image encoding and decoding process is optimized, unnecessary signaling is omitted, and messages with the same content are repeated in the current encoding unit.

Benefits of technology

It improves the efficiency of image encoding and decoding, optimizes the notification of filter attribute information, realizes efficient filter information signal transmission, reduces transmission and storage costs, and supports repeated message processing of the basic filter.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120513638A_ABST
    Figure CN120513638A_ABST
Patent Text Reader

Abstract

An image encoding / decoding method, a method for transmitting a bitstream, and a computer-readable recording medium for storing the bitstream are provided. An image decoding method according to the present disclosure may comprise the steps of: receiving a post neural network filter (NNPF) related supplemental enhancement information (SEI) message to be applied to a current picture; reconstructing NNPF related information on the basis of the NNPF related SEI message; and applying the NNPF to the current picture on the basis of the NNPF-related information. The NNPF related SEI message may include a post neural network filter characteristic (NNPFC) SEI message, the NNPF SEI message including filter identification information and filter attribute presence information. The NNPFC SEI message may also include base filter presence information indicating whether the NNPFC SEI message includes a base neural network post-filter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image encoding / decoding method, a method of transmitting a bitstream, and a recording medium storing the bitstream, and more particularly, to a method of processing a neural network post-filter. Background Art

[0002] Recently, demand for high-resolution and high-quality images, such as high-definition (HD) and ultra-high-definition (UHD), has been increasing in various fields. As the resolution and quality of image data improve, the amount of information or bits transmitted increases relative to existing image data. This increase in the amount of information or bits transmitted leads to increased transmission and storage costs.

[0003] Therefore, there is a need for efficient image compression technology for efficiently transmitting, storing, and reproducing information about high-resolution and high-quality images. Summary of the Invention

[0004] Technical issues

[0005] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0006] In addition, an object of the present disclosure is to provide an image encoding / decoding method and apparatus that can specify a time point for applying an updated neural network post-processing filter.

[0007] In addition, an object of the present disclosure is to provide an image encoding / decoding method and apparatus that can omit signaling of filter property presence information.

[0008] In addition, an object of the present disclosure is to provide an image encoding / decoding method and apparatus that efficiently signals the presence or absence of filter properties and the presence or absence of a base filter.

[0009] In addition, an object of the present disclosure is to provide an image encoding / decoding method and apparatus that can repeat an NNPFC SEI message including a basic neural network post-processing filter.

[0010] In addition, an object of the present disclosure is to provide an image encoding / decoding method and apparatus, wherein a repeated NNPFC SEI message has the same content as a first NNPFC SEI message having the same filter identification information (e.g., nnpfc_id) value within a current CLVS in decoding order.

[0011] In addition, an object of the present disclosure is to provide an image encoding / decoding method and apparatus that efficiently signals base filter existence information.

[0012] Furthermore, an object of the present disclosure is to provide a non-transitory computer-readable recording medium storing a bit stream generated by the image encoding method according to the present disclosure.

[0013] In addition, an object of the present disclosure is to provide a non-transitory computer-readable recording medium storing a bit stream received and decoded by the image decoding apparatus according to the present disclosure and used to reconstruct an image.

[0014] Furthermore, an object of the present disclosure is to provide a method of transmitting a bit stream generated by the image encoding method according to the present disclosure.

[0015] The technical problems solved by the present disclosure are not limited to the above-mentioned technical problems, and other technical problems not described herein will become apparent to those skilled in the art from the following description.

[0016] Technical Solution

[0017] An image decoding method according to one aspect of the present disclosure can be performed by an image decoding device and can include the following steps: receiving a neural network post filter (NNPF) related supplemental enhancement information (SEI) message to be applied to a current picture; reconstructing NNPF related information based on the NNPF related SEI message; and applying NNPF to the current picture based on the NNPF related information.

[0018] In the image decoding method according to the present disclosure, the NNPF-related SEI message may include a neural network post-filter characteristic (NNPFC) SEI message, which NNPFC SEI message includes filter identification information and filter property existence information, and the NNPFC SEI message may also include basic filter existence information specifying whether the NNPFC SEI message includes a basic neural network post-processing filter.

[0019] According to another aspect of the present disclosure, an image encoding method can be performed by an image encoding device and can include the following steps: generating neural network post filter (NNPF) related information to be applied to the current picture; and generating NNPF related supplemental enhancement information (SEI) message based on the NNPF related information.

[0020] In the image encoding method according to the present disclosure, the NNPF-related SEI message may include a neural network post-filter characteristics (NNPFC) SEI message, which NNPFC SEI message includes filter identification information and filter property existence information, and the NNPFC SEI message may also include basic filter existence information specifying whether the NNPFC SEI message includes a basic neural network post-processing filter.

[0021] A computer-readable recording medium according to another aspect of the present disclosure may store a bitstream generated by the image encoding method or apparatus of the present disclosure.

[0022] A transmitting method according to another aspect of the present disclosure may transmit a bit stream generated by the image encoding method or apparatus of the present disclosure.

[0023] The features briefly summarized above with respect to the present disclosure are merely exemplary aspects of the following detailed description of the present disclosure and do not limit the scope of the present disclosure.

[0024] Beneficial effects

[0025] According to the present disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.

[0026] Furthermore, according to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus that can specify a time point at which an updated neural network post-processing filter is applied.

[0027] Furthermore, according to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus that can omit signaling of filter property presence information.

[0028] Furthermore, according to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus that efficiently signals the presence or absence of filter properties and the presence or absence of a base filter.

[0029] In addition, according to the present disclosure, an image encoding / decoding method and apparatus that can repeat an NNPFC SEI message including a basic neural network post-processing filter can be provided.

[0030] Furthermore, according to the present disclosure, an image encoding / decoding method and apparatus may be provided, wherein a repeated NNPFC SEI message has the same content as a first NNPFC SEI message having the same filter identification information value within a current CLVS in decoding order.

[0031] Furthermore, according to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus for efficiently signaling base filter presence information.

[0032] Furthermore, according to the present disclosure, a non-transitory computer-readable recording medium storing a bit stream generated by the image encoding method according to the present disclosure can be provided.

[0033] Furthermore, according to the present disclosure, a non-transitory computer-readable recording medium storing a bit stream received and decoded by the image decoding apparatus according to the present disclosure and used to reconstruct an image can be provided.

[0034] Furthermore, according to the present disclosure, a method of transmitting a bitstream generated by the image encoding method according to the present disclosure may be provided.

[0035] Those skilled in the art will understand that the effects that can be achieved by the present disclosure are not limited to the effects specifically described above, and other advantages of the present disclosure will be more clearly understood from the detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 FIG. 1 is a diagram schematically illustrating a video encoding system to which embodiments of the present disclosure can be applied.

[0037] Figure 2 FIG. 1 is a diagram schematically illustrating an image encoding apparatus to which an embodiment of the present disclosure is applicable.

[0038] Figure 3 FIG. 1 is a diagram schematically illustrating an image decoding apparatus to which an embodiment of the present disclosure is applicable.

[0039] Figure 4 is a diagram illustrating an interleaving method for deriving a luma channel.

[0040] Figures 5 to 8 1 and 2 are diagrams for explaining various examples of persistence and cancellation of NNPFA.

[0041] Figure 9 This is a flowchart for explaining an image encoding method to which an embodiment of the present disclosure can be applied.

[0042] Figure 10 This is a flowchart for explaining an image decoding method to which an embodiment of the present disclosure can be applied.

[0043] Figure 11 is a diagram illustrating a content streaming system to which an embodiment of the present disclosure can be applied. DETAILED DESCRIPTION

[0044] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement them. However, the present disclosure can be implemented in various forms and is not limited to the embodiments described herein.

[0045] When describing the present disclosure, if it is determined that the detailed description of related known functions or configurations makes the scope of the present disclosure unnecessarily vague, its detailed description will be omitted. In the accompanying drawings, parts not related to the description of the present disclosure are omitted, and like reference numerals are attached to like parts.

[0046] In the present disclosure, when a component is “connected,” “coupled,” or “linked” to another component, it may include not only a direct connection relationship but also an indirect connection relationship with intermediate components. In addition, when a component “includes” or “has” other components, unless otherwise specified, it means that other components may be further included, rather than excluding other components.

[0047] In the present disclosure, the terms first, second, etc. may be used only to distinguish one component from other components and, unless otherwise specified, do not limit the order or importance of the components. Therefore, within the scope of the present disclosure, the first component in one embodiment may be referred to as the second component in another embodiment, and similarly, the second component in one embodiment may be referred to as the first component in another embodiment.

[0048] In this disclosure, components that are distinguished from each other are intended to clearly describe each feature and do not mean that the components must be separated. In other words, multiple components can be integrated and implemented in a single hardware or software unit, or a single component can be distributed and implemented in multiple hardware or software units. Therefore, even if not specifically stated, embodiments in which components are integrated or distributed are also included in the scope of this disclosure.

[0049] In the present disclosure, the components described in the various embodiments do not necessarily mean essential components, and some components may be optional components. Therefore, embodiments consisting of a subset of the components described in the embodiments are also included in the scope of the present disclosure. In addition, embodiments that include other components in addition to the components described in the various embodiments are also included in the scope of the present disclosure.

[0050] The present disclosure relates to encoding and decoding of images, and unless newly defined in the present disclosure, terms used in the present disclosure may have general meanings that are commonly used in the technical field to which the present disclosure belongs.

[0051] In this disclosure, a "picture" generally refers to the basis for representing an image in a specific time period, and a slice / tile is the coding basis that constitutes a portion of a picture. A picture can be composed of one or more slices / tiles. In addition, a slice / tile can include one or more coding tree units (CTUs).

[0052] In the present disclosure, "pixel" or "picture element (pel)" may refer to the smallest unit constituting a picture (or image). In addition, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luma component or only a pixel / pixel value of a chroma component.

[0053] In the present disclosure, a "unit" may refer to a basic unit of image processing. The unit may include at least one of a specific area of a picture and information related to the area. A unit may include a luminance block and two chrominance (e.g., Cb, Cr) blocks. In some cases, a unit may be used interchangeably with terms such as "sample array," "block," or "area." In general, an M×N block may include M columns and N rows of samples (or sample arrays) or a set (or array) of transform coefficients.

[0054] In the present disclosure, the term "current block" may refer to one of a "current coding block," a "current coding unit," an "encoding target block," a "decoding target block," or a "processing target block." When prediction is performed, the term "current block" may refer to a "current prediction block" or a "prediction target block." When transform (inverse transform) / quantization (dequantization) is performed, the term "current block" may refer to a "current transform block" or a "transform target block." When filtering is performed, the term "current block" may refer to a "filtering target block."

[0055] In addition, in the present disclosure, unless explicitly stated as a chroma block, the "current block" may refer to a block including both a luma component block and a chroma component block or a "luma block of the current block." The luma component block of the current block may be represented by an explicit description including the luma component block such as "luma block" or "current luma block." Furthermore, the "chroma component block of the current block" may be expressed by an explicit description including the chroma component block such as "chroma block" or "current chroma block."

[0056] In the present disclosure, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expressions "A / B" and "A, B" can mean "A and / or B". In addition, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C".

[0057] In the present disclosure, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" may include 1) only "A", 2) only "B", and / or 3) both "A and B". In other words, in the present disclosure, the term "or" should be interpreted as indicating "additionally or alternatively".

[0058] Video Coding System Overview

[0059] Figure 1 is a diagram illustrating a video encoding system to which embodiments of the present disclosure can be applied.

[0060] The video encoding system according to an embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may transfer encoded video and / or image information or data to the decoding device 20 in the form of a file or stream via a digital storage medium or a network.

[0061] The encoding device 10 according to the embodiment may include a video source generator 11, an encoding unit (encoder) 12, and a transmitter 13. The decoding device 20 according to the embodiment may include a receiver 21, a decoding unit (decoder) 22, and a renderer 23. The encoding unit 12 may be referred to as a video / image encoding device, and the decoding unit 22 may be referred to as a video / image decoding device. The transmitter 13 may be included in the encoding unit 12. The receiver 21 may be included in the decoding unit 22. The renderer 23 may include a display, and the display may be configured as a separate device or an external component.

[0062] The video source generator 11 can obtain video / images by capturing, synthesizing, or generating video / images. The video source generator 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured videos / images, etc. The video / image generation device may include, for example, a computer, a tablet computer, and a smartphone, and may (electronically) generate the video / image. For example, a virtual video / image may be generated by a computer, etc. In this case, the video / image capture process may be replaced by a process that generates relevant data.

[0063] The encoding unit 12 may encode the input video / image. The encoding unit 12 may perform a series of processes such as prediction, transformation, and quantization for compression and coding efficiency. The encoding unit 12 may output the encoded data (encoded video / image information) in the form of a bitstream.

[0064] The transmitter 13 can obtain the encoded video / image information or data output as a bitstream and forward it to the receiver 21 of the decoding device 20 or another external device in the form of a file or streaming via a digital storage medium or network. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter 13 can include components for generating media files using a predetermined file format and can also include components for transmitting via a broadcast / communication network. The transmitter 13 can be configured as a transmitting device separate from the encoding device 12. In this case, the transmitting device can include at least one processor that obtains the encoded video / image information or data output as a bitstream and a transmitting unit for transmitting the encoded video / image information or data in the form of a file or streaming. The receiver 21 can extract / receive the bitstream from the storage medium or network and transmit the bitstream to the decoding unit 22.

[0065] The decoding unit 22 may decode a video / image by performing a series of processes such as dequantization, inverse transformation, and prediction corresponding to the operations of the encoding unit 12 .

[0066] The renderer 23 may render the decoded video / image. The rendered video / image may be displayed on a display.

[0067] Overview of Image Coding Devices

[0068] Figure 2 FIG. 1 is a diagram schematically illustrating an image encoding apparatus to which an embodiment of the present disclosure is applicable.

[0069] like Figure 2 As shown, the image encoding apparatus 100 may include an image splitter 110, a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter-frame prediction unit 180, an intra-frame prediction unit 185, and an entropy encoder 190. The inter-frame prediction unit 180 and the intra-frame prediction unit 185 may be collectively referred to as a "prediction unit." The transformer 120, the quantizer 130, the dequantizer 140, and the inverse transformer 150 may be included in a residual processor. The residual processor may further include a subtractor 115.

[0070] In some embodiments, all or at least some of the components configuring the image encoding apparatus 100 may be configured by one hardware component (eg, an encoder or a processor). In addition, the memory 170 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium.

[0071] The image splitter 110 can split the input image (or picture or frame) input to the image encoding device 100 into one or more processing units. For example, the processing unit can be called a coding unit (CU). The coding unit can be obtained by recursively splitting the coding tree unit (CTU) or the largest coding unit (LCU) according to the quadtree binary tree ternary tree (QT / BT / TT) structure. For example, a coding unit can be split into multiple coding units of a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. For the splitting of the coding unit, the quadtree structure can be applied first, and the binary tree structure and / or the ternary structure can be applied later. The encoding process according to the present disclosure can be performed based on the final coding unit that is no longer split. The maximum coding unit can be used as the final coding unit, or the deeper coding unit obtained by splitting the maximum coding unit can be used as the final coding unit. Here, the encoding process may include prediction, transformation, and reconstruction processes, which will be described later. As another example, the processing unit of the encoding process can be a prediction unit (PU) or a transform unit (TU). The prediction unit and the transform unit may be divided or partitioned from the final coding unit. A prediction unit may be a unit of sample prediction, and a transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from the transform coefficient.

[0072] The prediction unit (inter prediction unit 180 or intra prediction unit 185) may perform prediction on the block to be processed (current block) and generate a prediction block including prediction samples for the current block. The prediction unit may determine whether to apply intra prediction or inter prediction based on the current block or CU. The prediction unit may generate various information related to the prediction of the current block and transmit the generated information to the entropy encoder 190. The information about the prediction may be encoded in the entropy encoder 190 and output in the form of a bitstream.

[0073] The intra prediction unit (intra predictor) 185 can predict the current block by referring to samples in the current picture. The reference samples can be located in the neighborhood of the current block, or can be located separately according to the intra prediction mode and / or intra prediction technology. The intra prediction mode may include multiple non-directional modes and multiple directional modes. The non-directional mode may include, for example, a DC mode and a planar mode. The directional mode may include, for example, 33 directional prediction modes or 65 directional prediction modes according to the level of detail of the prediction direction. However, this is only an example, and more or fewer directional prediction modes may be used depending on the settings. The intra prediction unit 185 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.

[0074] The inter-frame prediction unit (inter-frame predictor) 180 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporally neighboring block can be the same or different. Temporally neighboring blocks can be referred to as collocated reference blocks, collocated CUs (colCUs), etc. A reference picture including temporally neighboring blocks can be referred to as collocated pictures (colPics). For example, the inter-frame prediction unit 180 can construct a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter-frame prediction unit 180 can use the motion information of the neighboring block as the motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted. In the case of motion vector prediction (MVP) mode, the motion vector of the neighboring block may be used as a motion vector predictor, and the motion vector of the current block may be signaled by encoding a motion vector difference and an indicator of the motion vector predictor. The motion vector difference may represent the difference between the motion vector of the current block and the motion vector predictor.

[0075] The prediction unit can generate a prediction signal based on various prediction methods and prediction techniques described below. For example, the prediction unit can apply not only intra-frame prediction or inter-frame prediction, but also both intra-frame prediction and inter-frame prediction simultaneously to predict the current block. The prediction method that simultaneously applies both intra-frame prediction and inter-frame prediction to predict the current block is referred to as combined inter-frame and intra-frame prediction (CIIP). In addition, the prediction unit can perform intra-frame block copying (IBC) for the prediction of the current block. Intra-frame block copying can be used for content image / video encoding such as gaming, such as screen content coding (SCC). IBC is a method of predicting the current picture using a previously reconstructed reference block in the current picture at a predetermined distance from the current block. When IBC is applied, the position of the reference block in the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC essentially performs prediction within the current picture, but can be performed similarly to inter-frame prediction because the reference block is derived within the current picture. In other words, IBC can use at least one of the inter-frame prediction techniques described in this disclosure.

[0076] The prediction signal generated by the prediction unit can be used to generate a reconstructed signal or a residual signal. The subtractor 115 can generate a residual signal (residual block or residual sample array) by subtracting the prediction signal (prediction block or prediction sample array) output from the prediction unit from the input image signal (original block or original sample array). The generated residual signal can be sent to the transformer 120.

[0077] The transformer 120 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented by a graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process may be applied to square pixel blocks of the same size, or may be applied to blocks of variable size other than square.

[0078] The quantizer 130 may quantize the transform coefficients and transmit them to the entropy encoder 190. The entropy encoder 190 may encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 130 may rearrange the quantized transform coefficients of the block type into a one-dimensional vector form based on the coefficient scanning order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.

[0079] The entropy encoder 190 can perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 190 can encode information required for video / image reconstruction (e.g., values of syntax elements, etc.) in addition to quantized transform coefficients, together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of a network abstraction layer (NAL). The video / image information may also include information about various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may also include general constraint information. The signaled information, transmitted information, and / or syntax elements described in the present disclosure may be encoded through the above-mentioned encoding process and included in the bitstream.

[0080] The bitstream may be transmitted over a network or may be stored in a digital storage medium. The network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that transmits a signal output from the entropy encoder 190 and / or a storage unit (not shown) that stores the signal may be included as an internal / external element of the image encoding apparatus 100. Alternatively, the transmitter may be provided as a component of the entropy encoder 190.

[0081] The quantized transform coefficients output from the quantizer 130 may be used to generate a residual signal. For example, the residual signal (residual block or residual sample) may be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients through the dequantizer 140 and the inverse transformer 150.

[0082] The adder 155 adds the reconstructed residual signal to the prediction signal output from the inter-frame prediction unit 180 or the intra-frame prediction unit 185 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). If there is no residual for the block to be processed (such as when skip mode is applied), the prediction block can be used as the reconstructed block. The adder 155 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture, and can be used for inter-frame prediction of the next picture through filtering as described below.

[0083] The filter 160 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 160 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 160 can generate various information related to filtering and send the generated information to the entropy encoder 190, as described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoder 190 and output in the form of a bitstream.

[0084] The modified reconstructed picture transmitted to the memory 170 may be used as a reference picture in the inter prediction unit 180. When inter prediction is applied by the image encoding apparatus 100, prediction mismatch between the image encoding apparatus 100 and the image decoding apparatus may be avoided, and encoding efficiency may be improved.

[0085] The DPB of the memory 170 may store the modified reconstructed picture for use as a reference picture in the inter-frame prediction unit 180. The memory 170 may store the motion information of the block from which the motion information in the current picture is derived (or encoded) and / or the motion information of the block in the reconstructed picture. The stored motion information may be sent to the inter-frame prediction unit 180 and used as the motion information of the spatially adjacent block or the motion information of the temporally adjacent block. The memory 170 may store the reconstructed samples of the reconstructed block in the current picture and may transmit the reconstructed samples to the intra-frame prediction unit 185.

[0086] Overview of Image Decoding Equipment

[0087] Figure 3 FIG. 1 is a diagram schematically illustrating an image decoding apparatus to which an embodiment of the present disclosure is applicable.

[0088] like Figure 3 As shown, the image decoding apparatus 200 may include an entropy decoder 210, a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame predictor 260, and an intra-frame prediction unit 265. The inter-frame predictor (inter-frame prediction unit) 260 and the intra-frame predictor (intra-frame prediction unit) 265 may be collectively referred to as a "prediction unit (predictor)". The dequantizer 220 and the inverse transformer 230 may be included in a residual processor.

[0089] According to an embodiment, all or at least some of the components configuring the image decoding apparatus 200 may be configured by hardware components (eg, a decoder or a processor). In addition, the memory 170 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium.

[0090] The image decoding apparatus 200 having received a bit stream including video / image information may decode the image by performing the same operation as that performed by Figure 2 The image is reconstructed by processing corresponding to the processing performed by the image encoding device 100. For example, the image decoding device 200 can perform decoding using the processing unit applied in the image encoding device. Therefore, the processing unit of decoding can be, for example, a coding unit. The coding unit can be obtained by dividing the coding tree unit or the maximum coding unit. The reconstructed image signal decoded and output by the image decoding device 200 can be reproduced by a reproduction device (not shown).

[0091] The image decoding device 200 may receive the image in the form of a bit stream from Figure 2The received signal can be decoded by the entropy decoder 210. For example, the entropy decoder 210 can parse the bitstream to derive information required for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may also include information about various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may also include general constraint information. The image decoding device may also decode the picture based on the information about the parameter set and / or the general constraint information. The signaled / received information and / or syntax elements described in the present disclosure may be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 210 may decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, as well as the output values of the syntax elements required for image reconstruction and the quantized values of the transform coefficients of the residual. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, use the decoding target syntax element information, the decoding information of the neighboring blocks and the decoding target block, or the information of the symbol / bin decoded in the previous stage to determine the context model, and perform arithmetic decoding on the bin by predicting the probability of occurrence of the bin according to the determined context model, and generate a symbol corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method can update the context model by using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. The information related to prediction among the information decoded by the entropy decoder 210 can be provided to the prediction unit (inter-frame predictor 260 and intra-frame prediction unit 265), and the residual value (i.e., quantized transform coefficient and related parameter information) on which entropy decoding is performed in the entropy decoder 210 can be input to the dequantizer 220. In addition, the information about filtering among the information decoded by the entropy decoder 210 can be provided to the filter 240. Meanwhile, a receiver (not shown) for receiving a signal output from the image encoding apparatus may also be configured as an internal / external element of the image decoding apparatus 200 , or the receiver may be a component of the entropy decoder 210 .

[0092] In addition, the image decoding device according to the present disclosure may be referred to as a video / image / picture decoding device. The image decoding device can be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoder 210. The sample decoder may include a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame prediction unit 260, or at least one of an intra-frame prediction unit 265.

[0093] The dequantizer 220 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 220 may rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the image encoding device. The dequantizer 220 may dequantize the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.

[0094] The inverse transformer 230 may inversely transform the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0095] The prediction unit may perform prediction on the current block and generate a prediction block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block based on the prediction information output from the entropy decoder 210, and may determine a specific intra / inter prediction mode (prediction technique).

[0096] As described in the prediction unit of the image encoding device 100 , the prediction unit can generate a prediction signal based on various prediction methods (techniques) to be described later.

[0097] The intra-frame predictor 265 may predict the current block by referring to samples in the current picture. The description of the intra-frame prediction unit 185 is also applicable to the intra-frame prediction unit 265.

[0098] The inter-frame predictor 260 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include information on the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter-frame predictor 260 may configure a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the prediction information may include information indicating the inter-frame prediction mode for the current block.

[0099] The adder 235 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter-frame predictor 260 and / or the intra-frame prediction unit 265). If there is no residual for the block to be processed, for example, when the skip mode is applied, the prediction block can be used as the reconstructed block. The description of the adder 155 also applies to the adder 235. The adder 235 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture, and can be used for inter-frame prediction of the next picture through filtering as described below.

[0100] The filter 240 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 240 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 250, specifically, in the DPB of the memory 250. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0101] The (modified) reconstructed picture stored in the DPB of the memory 250 can be used as a reference picture in the inter-frame predictor 260. The memory 250 can store the motion information of the block from which the motion information in the current picture is derived (or decoded) and / or the motion information of the block in the reconstructed picture. The stored motion information can be sent to the inter-frame predictor 260 to be used as the motion information of the spatial neighboring block or the motion information of the temporal neighboring block. The memory 250 can store the reconstructed samples of the reconstructed block in the current picture and transmit the reconstructed samples to the intra-frame prediction unit 265.

[0102] In the present disclosure, the embodiments described in the filter 160, the inter-frame prediction unit 180, and the intra-frame prediction unit 185 of the image encoding device 100 may be equally or correspondingly applied to the filter 240, the inter-frame predictor 260, and the intra-frame predictor 265 of the image decoding device 200.

[0103] Neural Network Post-Filter Characteristic (NNPFC)

[0104] The combination of Table 1 and Table 2 represents the NNPFC grammatical structure.

[0105] [Table 1]

[0106]

[0107] [Table 2]

[0108]

[0109] The NNPFC syntax structures of Table 1 and Table 2 may be signaled in the form of a Supplemental Enhancement Information (SEI) message. The SEI message signaling the NNPFC syntax structures of Table 1 and Table 2 may be referred to as an NNPFC SEI message.

[0110] The Neural Network Post-Filter Characteristics (NNPFC) SEI message can specify a neural network that can be used as a post-processing filter. The Neural Network Post-Filter Activation (NNPFA) SEI message can be used to indicate the use of a specified neural network post-processing filter (NNPF) for a particular picture. Here, "post-processing filter" and "post-filter" can have the same meaning.

[0111] Using this SEI message requires defining the following variables:

[0112] - The width and height of the decoded output picture can be cropped in units of luma samples, which can be expressed as CroppedWidth and CroppedHeight respectively.

[0113] - CroppedYPic[idx] (luminance sample array of the cropped decoded output picture) and CroppedCbPic[idx] and CroppedCrPic[idx] (which are chroma sample arrays), when present, can be used as input to the post-processing filter, and idx can have a range of 0 to numInputPics-1 (including 0 and numInputPics-1).

[0114] -BitDepth Y The bit depth of the luma sample array of the cropped decoded output picture can be represented.

[0115] -BitDepth C The bit depth of the chroma sample array of the cropped decoded output picture (if any) may be represented.

[0116] -ChromaFormatIdc may represent a chroma format identifier.

[0117] - When the value of nnpfc_auxiliary_inp_idc is equal to 1, the filtering strength control value StrengthControlVal shall be a real number in the range of 0 to 1 (inclusive).

[0118] The variables SubWidthC and SubHeightC can be derived from ChromaFormatIdc. Two or more NNPFC SEI messages may exist for the same picture. If two or more NNPFC SEI messages with different nnpfc_id values exist or are enabled for the same picture, the two or more NNPFC SEI messages may have the same or different nnpfc_purpose and nnpfc_mode_idx values.

[0119] nnpfc_id can contain an identification number that can be used to identify the post-processing filter. The nnpfc_id value should be between 0 and 2. 32 -2 (including 0 and 2 32 -2). You can keep the range of 256 to 511 (including 256 and 511) and 2 31 to 2 32 -2 (including 2 31 and 2 32 -2) are for future use. Decoders should ignore nnpfc_id values in the range of 256 to 511 (inclusive) or in the range of 2 31 to 2 32 -2 (including 2 31 and 2 32 -2) NNPFC SEI message within the scope of.

[0120] If the NNPFC SEI message is the first NNPFC SEI message with a particular nnpfc_id value within the current coding layer video sequence (CLVS) in decoding order, the following applies:

[0121] -SEI messages may specify base post-processing filters.

[0122] - A SEI message may be associated with the current decoded picture and all subsequent decoded pictures of the current layer in output order until the end of the current CLVS.

[0123] The NNPFC SEI message may be a repetition of a previous NNPFC SEI message within the current CLVS in decoding order, and subsequent semantics may apply as if this SEI message were the only NNPFC SEI message with the same content within the current CLVS.

[0124] If the NNPFC SEI message is not the first NNPFC SEI message with a particular nnpfc_id value within the current CLVS in decoding order, the following may apply:

[0125] - The SEI message may relate to all subsequently decoded pictures of the current CLVS or the current layer in output order, until the current CLVS or until the end of the current CLVS, or may relate to the next NNPFC SEI message with a specific nnpfc_id value within the current CLVS in output order.

[0126] When the NNPFC SEI message is the first NNPFC SEI message with a specific nnpfc_id value within the previous CLVS in decoding order, the value of nnpfc_mode_idc equal to 1 can specify that the basic post-processing filter associated with the nnpfc_id value is a neural network, where the neural network can be a neural network identified by the format identified by the tag URI nnpfc_tag_uri using the URI represented by nnpfc_uri.

[0127] When the NNPFC SEI message is not the first NNPFC SEI message with a particular nnpfc_id value within the current CLVS in decoding order, the value of nnpfc_mode_idc equal to 1 may specify that updates to the base post-processing filters with the same nnpfc_id value are defined by the URI denoted by nnpfc_uri using the format identified by the tag URI nnpfc_tag_uri.

[0128] The value of nnpfc_mode_idc may be restricted to the range of 0 to 1 in the bitstream. Values of nnpfc_mode_idc in the range of 2 to 255 (inclusive) may be reserved for future use and may not be present in the bitstream. The decoder shall ignore NNPFC SEI messages with nnpfc_mode_idc in the range of 2 to 255 (inclusive). Values of nnpfc_mode_idc greater than 255 are not present in the bitstream and may not be reserved for future use.

[0129] When the SEI message is the first NNPFC SEI message with a specific nnpfc_id value within the current CLVS in decoding order, the post-processing filter PostProcessingFilter() may be identically assigned to the base post-processing filter.

[0130] When the SEI message is not the first NNPFC SEI message with a particular nnpfc_id value within the current CLVS in decoding order, the post-processing filter PostProcessingFilter() may be obtained by applying the updates defined by the SEI message to the base post-processing filter.

[0131] Updates are not cumulative; instead, each update may be applied to the base post-processing filter, which is the post-processing filter specified by the first NNPFC SEI message with a particular nnpfc_id value within the current CLVS in decoding order.

[0132] nnpfc_reserved_zero_bit_a may be restricted due to bitstream constraints to have a value equal to 0. The decoder shall ignore NNPFC SEI messages with a non-zero value of nnpfc_reserved_zero_bit_a.

[0133] nnpfc_tag_uri can contain a tag URI with the syntax and semantics specified in IETF RFC 4151 that identifies a neural network used as a base post-processing filter, or an update relative to a base post-processing filter using the nnpfc_id value specified by nnpfc_uri. The use of nnpfc_tag_uri allows the format of the neural network data specified by nnrpf_uri to be uniquely identified without the need for a central registry. nnpfc_tag_uri equal to "tag:iso.org,2023:15938-17" can specify that the neural network data identified by nnpfc_uri conforms to ISO / IEC 15938-17.

[0134] nnpfc_uri may contain a URI having the syntax and semantics specified in IETF Internet Standard 66 that identifies a neural network used as a base post-processing filter, or an update relative to a base post-processing filter using the same nnpfc_id value.

[0135] The value of nnpfc_formatting_and_purpose_flag equal to 1 may specify the presence of syntax elements related to filter purpose, input formatting, output formatting, and complexity. The value of nnpfc_formatting_and_purpose_flag equal to 0 may specify the absence of syntax elements related to filter purpose, input formatting, output formatting, and complexity.

[0136] When the SEI message is the first NNPFC SEI message with a specific nnpfc_id value in the current CLVS in decoding order, the value of nnpfc_formatting_and_purpose_flag may be equal to 1. When the SEI message is not the first NNPFC SEI message with a specific nnpfc_id value in the current CLVS in decoding order, the value of nnpfc_formatting_and_purpose_flag may be equal to 0.

[0137] nnpfc_purpose may specify the purpose of the post-processing filter as specified in Table 3.

[0138] Due to bitstream constraints, the value of nnpfc_purpose shall be in the range of 0 to 5 (inclusive). Values of nnpfc_purpose in the range of 6 to 1023 (inclusive) are not present in the bitstream and may be reserved for future use. The decoder shall ignore NNPFC SEI messages with nnpfc_purpose in the range of 6 to 1203 (inclusive). Values of nnpfc_purpose greater than 1023 are not present in the bitstream and are not reserved for future use.

[0139] [Table 3]

[0140]

[0141] When the reserved value of nnpfc_purpose is used in the future, the syntax of this SEI message may be extended to syntax elements that exist only when nnpfc_purpose is equal to the corresponding value. When the value of SubWidthC is 1 and the value of SubHeightC is 1, nnpfc_purpose shall not have a value of 2 or 4.

[0142] A value of nnpfc_out_sub_c_flag equal to 1 specifies that outSubWidthC is 1 and outSubHeightC is 1. A value of nnpfc_out_sub_c_flag equal to 0 specifies that outSubWidthC is 2 and outSubHeightC is 1. If nnpfc_out_sub_c_flag is not present, outSubWidthC is inferred to be equal to SubWidthC, and outSubHeightC is inferred to be equal to SubHeightC. When the value of ChromaFormatIdc is 2 and nnpfc_out_sub_c_flag is present, the value of nnpfc_out_sub_c_flag shall be 1.

[0143] nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples may represent the width and height, respectively, of the luma sample array of the picture resulting from applying the post-processing filter identified by nnpfc_id to the cropped decoded output picture. When nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples are not present, they may be inferred to be equal to CroppedWidth and CroppedHeight, respectively. The value of nnpfc_pic_width_in_luma_samples shall be in the range of CroppedWidth to CroppedWidth*16-1. The value of nnpfc_pic_height_in_luma_samples shall be in the range of CroppedHeight to CroppedHeight*16-1.

[0144] nnpfc_num_input_pics_minus2+2 may represent the number of decoded output pictures used as input to the post-processing filter.

[0145] nnpfc_interpolated_pics[i] may represent the number of interpolated pictures generated by the post-processing filter between an i-th picture and an (i+1)-th picture serving as an input of the post-processing filter.

[0146] As shown in Table 4, a variable numInputPics representing the number of pictures used as input to the post-processing filter and a variable numOutputPics representing the total number of pictures generated as a result of the post-processing filter may be derived.

[0147] [Table 4]

[0148]

[0149] The value of nnpfc_component_last_flag equal to 1 can specify the last dimension of the input tensor inputTensor for the post-processing filter and the output tensor outputTensor as the result of the post-processing filter for the current channel. The value of nnpfc_component_last_flag equal to 0 can specify the third dimension of the input tensor inputTensor for the post-processing filter and the output tensor outputTensor as the result of the post-processing filter for the current channel. The first dimension of the input tensor and the output tensor can be used in the batch index used in some neural network frameworks. The formula in the semantics of this SEI message uses the batch size corresponding to the batch index equal to 0, but the batch size used as input for neural network inference can be determined by the implementation of post-processing.

[0150] For example, when the value of nnpfc_inp_order_idc is equal to 3 and the value of nnpfc_auxiliary_inp_idc is equal to 1, the input tensor may have 7 channels, including 4 luma matrices, 2 chroma matrices, and 1 auxiliary input matrix. In this case, the DeriveInputTensors() process may derive each of the 7 channels of the input tensor one by one, and when a particular channel among these channels is processed, the channel may be referred to as the current channel during the process.

[0151] nnpfc_inp_format_idc can specify a method for converting the sample value of the cropped decoded output picture into the input value of the post-processing filter. If nnpfc_inp_format_idc is 0, the input value of the post-processing filter is a real number, and the InpY() and InpC() functions can be specified as in Equation 1.

[0152] [Equation 1]

[0153] InpY(x)=x÷((1< <BitDepthY)-1)

[0154] InpC(x)=x÷((1< <BitDepthC)-1)

[0155] When the value of nnpfc_inp_format_idc is 1, the input value of the post-processing filter is an unsigned integer, and the InpY() and InpC() functions can be derived as shown in Table 5.

[0156] [Table 5]

[0157]

[0158] variable inpTensorBitDepth Y It can be derived from the syntax element nnpfc_inp_tensor_bitdepth_minus8, as described below. Values of nnpfc_inp_format_idc greater than 1 may be reserved for future use and may not be present in the bitstream. A decoder shall ignore NNPFC SEI messages containing reserved values of nnpfc_inp_format_idc.

[0159] nnpfc_inp_tensor_luma_bitdepth_minus8+8 can specify the bit depth of the luminance sample values in the input integer tensor. Y The value of can be derived as shown in Equation 2.

[0160] [Equation 2]

[0161] inpTensorBitDepth=nnpfc_inp_tensor_bitdepth_minus8+8

[0162] The value of nnpfc_inp_tensor_bitlength_minus8 can be clamped to the range of 0 to 24.

[0163] nnpfc_inp_order_idc may specify a method of ordering the sample array of the cropped decoded output picture into one of the input pictures for the post-processing filter.

[0164] In the bitstream, the value of nnpfc_inp_order_idc shall be in the range of 0 to 3 (inclusive). Values of nnpfc_inp_order_idc from 4 to 255 (inclusive) shall not be present in the bitstream. The decoder shall ignore NNPFC SEI messages with nnpfc_inp_order_idc in the range of 4 to 255 (inclusive). Values of nnpfc_inp_order_idc greater than 255 shall not be present in the bitstream and are not reserved for future use.

[0165] When ChromaFormatIdc is not equal to 1, the value of nnpfc_inp_order_idc shall not be 3.

[0166] Table 6 shows an informative description of nnpfc_inp_order_idc values.

[0167] [Table 6]

[0168]

[0169]

[0170] A patch is a rectangular array of samples of a component (eg, luma or chroma components) from a picture.

[0171] nnpfc_auxiliary_inp_idc greater than 0 specifies that auxiliary input data is present in the input tensor of the neural network postfilter. nnpfc_auxiliary_inp_idc equal to 0 specifies that auxiliary input data is not present in the input tensor. nnpfc_auxiliary_inp_idc equal to 1 specifies that auxiliary input data is derived using the methods shown in Tables 7 to 9.

[0172] In the bitstream, the value of nnpfc_auxiliary_inp_idc shall be in the range of 0 to 1 (inclusive). Values of nnpfc_inp_order_idc from 2 to 255 (inclusive) shall not be present in the bitstream. The decoder shall ignore NNPFC SEI messages with nnpfc_inp_order_idc in the range of 2 to 255 (inclusive). Values of nnpfc_inp_order_idc greater than 255 shall not be present in the bitstream and are not reserved for future use.

[0173] The procedure DeriveInputTensors() for deriving the input tensor inputTensor for a given vertical sample coordinate cTop and horizontal sample coordinate cLeft of the top left sample position of a sample patch contained in the specified input tensor can be represented as a combination of Tables 7 to 9.

[0174] [Table 7]

[0175]

[0176] [Table 8]

[0177]

[0178] [Table 9]

[0179]

[0180] nnpfc_separate_colour_description_present_flag equal to 1 may specify that the unique combination of color primaries, transform characteristics, and matrix coefficients for the picture due to the post-processing filters is specified in the SEI message syntax structure. nnfpc_separate_colour_description_present_flag equal to 0 may specify that the combination of color primaries, transform characteristics, and matrix coefficients for the picture due to the post-processing filters is the same as indicated in the VUI parameters of the CLVS.

[0181] The nnpfc_colour_primaries may have the same semantics as defined for the vui_colour_primaries syntax element, except that:

[0182] -nnpfc_colour_primaries may specify the primaries of the picture resulting from applying the neural network postfilter specified in the SEI message, instead of the primaries used in CLVS.

[0183] - If nnpfc_colour_primaries is not present in the NNPFC SEI message, the value of nnpfc_colour_primaries may be inferred to be the same as the value of vui_colour_primaries.

[0184] The nnpfc_transfer_characteristics may have the same semantics as defined for the vui_transfer_characteristics syntax element, except that:

[0185] -nnpfc_transfer_characteristics may specify the transform characteristics of the picture resulting from applying the neural network postfilter specified in the SEI message, instead of the transform characteristics used in CLVS.

[0186] - If nnpfc_transfer_characteristics is not present in the NNPFC SEI message, the value of nnpfc_transfer_characteristics may be inferred to be the same as the value of vui_transfer_characteristics.

[0187] The nnpfc_matrix_coeffs may have the same semantics as specified for the vui_matrix_coeffs syntax element, except that:

[0188] - The nnpfc_matrix_coeffs may specify the matrix coefficients of the pictures generated by the neural network post-filter specified in the applied SEI message, rather than the matrix coefficients for CLVS.

[0189] - If the nnpfc_matrix_coeffs does not exist in the NNPFC SEI message, the value of nnpfc_matrix_coeffs may be inferred to be the same as the value of vui_matrix_coeffs.

[0190] - The allowed values of nnpfc_matrix_coeffs may not be constrained by the chroma format of the decoded video picture, as indicated by the ChromaFormatIdc value of the semantics of the VUI parameters.

[0191] - If the value of nnpfc_matrix_coeffs is equal to 0, the value of nnpfc_out_order_idc cannot be equal to 1 or 3.

[0192] nnpfc_out_format_id being equal to 0 may specify that the sample values of the post-processing filter output are real numbers, which are linearly mapped from values ranging from 0 to 1 to the unsigned integer value range from 0 to (1<<bitDepth)-1 for the bit depth bitDepth required for subsequent post-processing or display. nnpfc_out_format_flag being equal to 1 may specify that the sample values of the post-processing filter output are unsigned integers within the range from 0 to (1<<(nnpfc_out_tensor_bitlength_minus8+8))-1. Values of nnpfc_out_format_idc greater than 1 do not exist in the bitstream. The decoder shall ignore the NNPFC SEI message containing the reserved values of nnpfc_out_format_idc. ‘+8’ may specify the bit depth of the sample values in the output integer tensor. The value of nnpfc_out_tensor_bitlength_minus8 shall be in the range from 0 to 24.

[0193] nnpfc_out_order_idc may specify the output order of samples output from the post-processing filter. The value of nnpfc_out_order_idc shall be in the range of 0 to 3 in the bitstream. Values of nnpfc_out_order_idc of 4 to 255 shall not be present in the bitstream. The decoder shall ignore NNPFC SEI messages with nnpfc_out_order_idc in the range of 4 to 255. Values of nnpfc_out_order_idc greater than 255 shall not be present in the bitstream and are not reserved for future use. When the value of nnpfc_purpose is 2 or 4, the value of nnpfc_out_order_idc shall not be equal to 3.

[0194] Table 10 describes the values of nnpfc_out_order_idc.

[0195] [Table 10]

[0196]

[0197] For a given vertical sample coordinate cTop and horizontal sample coordinate cLeft indicating the top left sample position of the sample patch contained in the input tensor, the procedure StoreOutputTensors() for deriving sample values in the filtered output sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor outputTensor can be expressed as a combination of Table 11 and Table 12.

[0198] [Table 11]

[0199]

[0200] [Table 12]

[0201]

[0202] nnpfc_constant_patch_size_flag equal to 1 may specify that the post-processing filter receives as input exactly the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. nnpfc_constant_patch_size_flag equal to 0 may specify that the post-processing filter receives as input any patch size that is a positive integer multiple of the patch size indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. npfc_patch_width_minus1+1 may specify the number of horizontal samples of the patch size required for input to the post-processing filter when the value of nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_width_minus1 should be in the range of 0 to Min(32766, croppedWidth-1).

[0203] nnpfc_patch_height_minus1+1 may specify the number of vertical samples of the patch size required for the input to the post-processing filter when the value of nnpfc_constant_patch_size_flag is equal to 1. The value of nnpfc_patch_height_minus1 should be in the range of 0 to Min(32766, croppedHeight-1).

[0204] The variables inpPatchWidth and inpPatchHeight can be set to patch_size_width and patch_size_height respectively.

[0205] When the value of nnpfc_constant_patch_size_flag is equal to 0, the following applies:

[0206] -The values of inpPatchWidth and inpPatchHeight can be provided by an external device or set by the post-processor itself.

[0207] -inpPatchWidth should be a positive integer multiple of nnpfc_patch_width_minus1+1 and should be less than or equal to CroppedWidth. -inpPatchHeight should be a positive integer multiple of nnpfc_patch_height_minus1+1 and should be less than or equal to CroppedHeight.

[0208] Otherwise (if the value of nnpfc_constant_patch_size_flag is equal to 1), the value of inpPatchWidth may be set equal to nnpfc_patch_width_minus1+1, and the value of inpPatchHeight may be set equal to nnpfc_patch_height_minus1+1.

[0209] nnpfc_overlap can specify the number of overlapping horizontal and vertical samples of adjacent input tensors for the post-processing filter. The value of nnpfc_overlap should be in the range of 0 to 16383.

[0210] The variables outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, outPatchCHeight, and overlappSize can be derived as shown in Table 13.

[0211] [Table 13]

[0212]

[0213] The bitstream conformance requirement is that outPatchWidth*CroppedWidth shall be equal to nnpfc_pic_width_in_luma_samples*inpPatchWidth, and outPatchHeight*CroppedHeight shall be equal to nnpfc_pic_height_in_luma_samples*inpPatchHeight. As described in Table 14, nnpfc_padding_type may specify the padding treatment when referencing sample positions outside the boundaries of the cropped decoded output picture. The value of nnpfc_padding_type shall be in the range of 0 to 15.

[0214] [Table 14]

[0215] nnpfc_padding_type describe 0 Zero padding 1 Copy Fill 2 Mirror Fill 3 Surround Fill 4 Fixed padding 5…15 reserve

[0216] nnpfc_luma_padding_val can specify the luma value to be used for padding when the value of nnpfc_padding_type is 4.

[0217] nnpfc_cb_padding_val may specify the Cb value to be used for padding when the value of nnpfc_padding_type is 4.

[0218] nnpfc_cr_padding_val specifies the Cr value to be used for padding when the value of nnpfc_padding_type is 4.

[0219] The InpSampleVal(y, x, picHeight, picWidth, CroppedPic) function (whose inputs are the vertical sample position y, the horizontal sample position x, the picture height picHeight, the picture width picWidth and the sample array CroppedPic) can return the derived SampleVal value as shown in Table 15.

[0220] For the input to the InpSampleVal() function, the vertical position can be listed before the horizontal position to be compatible with the input tensor rules of some inference engines.

[0221] [Table 15]

[0222]

[0223] The processing in Table 16 may be used to filter the cropped decoded output picture in a patch-wise manner using the post-processing filter PostProcessingFilter() to generate a filtered picture, which may include a Y sample array FilteredYPic, a Cb sample array FilteredCbPic, and a Cr sample array FilteredCrPic, as indicated by nnpfc_out_order_idc.

[0224] [Table 16]

[0225]

[0226] nnpfc_complexity_info_present_flag equal to 1 may specify the presence of one or more syntax elements indicating the complexity of the post-processing filter associated with nnpfc_id.nnpfc_complexity_info_present_flag equal to 0 may specify the absence of syntax elements indicating the complexity of the post-processing filter associated with nnpfc_id.

[0227] nnpfc_parameter_type_idc equal to 0 specifies that the neural network uses only integer parameters. nnpfc_parameter_type_flag equal to 1 specifies that the neural network can use floating-point or integer parameters. nnpfc_parameter_type_idc equal to 2 specifies that the neural network uses only binary parameters. nnpfc_parameter_type_idc equal to 3 is reserved for future use and is not present in the bitstream. The decoder should ignore the NNPFC SEI message with nnpfc_parameter_type_idc equal to 3.

[0228] nnpfc_log2_parameter_bit_length_minus3 equal to 0, 1, 2, and 3 may specify that the neural network does not use parameters with bit lengths greater than 8, 16, 32, and 64, respectively. If nnpfc_parameter_type_idc is present and nnpfc_log2_parameter_bit_length_minus3 is not present, the neural network may not use parameters with bit lengths greater than 1.

[0229] nnpfc_num_parameters_idc may specify the maximum number of neural network parameters for the post-processing filter in powers of 2048. nnpfc_num_parameters_idc equal to 0 may specify that the maximum number of neural network parameters is unknown. The value of nnpfc_num_parameters_idc shall be in the range of 0 to 52. Values of nnpfc_num_parameters_idc greater than 52 shall not be present in the bitstream. The decoder shall ignore NNPFC SEI messages with nnpfc_num_parameters_idc greater than 52.

[0230] When the value of nnpfc_num_parameters_idc is greater than 0, the variable maxNumParameters may be derived as in Equation 3.

[0231] [Equation 3]

[0232] maxNumParameters=(2048< <nnpfc_num_parameters_ide)-1

[0233] The number of neural network parameters in a post-processing filter can be restricted to be less than or equal to maxNumParameters.

[0234] nnpfc_num_kmac_operations_idc greater than 0 specifies that the maximum number of multiply-accumulate operations per sample in the post-processing filter is less than or equal to nnpfc_num_kmac_operations_idc * 1000. nnpfc_num_kmac_operations_idc equal to 0 specifies that the maximum number of multiply-accumulate operations in the network is unknown. The value of nnpfc_num_kmac_operations_idc should be between 0 and 2. 32 -1 (including 0 and 2 32 -1).

[0235] nnpfc_total_kilobyte_size greater than 0 specifies the total size (in kilobytes) required to store the uncompressed parameters of the neural network. The total size in bits can be a number greater than or equal to the sum of the bits used to store each parameter. nnpfc_total_kilobyte_size can be obtained by dividing the total size (in bits) by 8000 and rounding it up. nnpfc_total_kilobyte_size equal to 0 specifies that the total size required to store the parameters of the neural network is unknown. The value of nnpfc_total_kilobyte_size should be between 0 and 2. 32 -1 (including 0 and 2 32 -1).

[0236] nnpfc_reserved_zero_bit_b shall be equal to 0 in the bitstream. A decoder shall ignore NNPFC SEI messages where nnpfc_reserved_zero_bit_b is not 0.

[0237] nnpfc_payload_byte[i] may contain the i-th byte of the bitstream. The byte sequence nnpfc_payload_byte[i] for all current values of i shall be a complete bitstream conforming to ISO / IEC 15938-17.

[0238] Neural Network Post-Filter Activation (NNFPA)

[0239] The grammatical structure of NNFPA is shown in Table 17.

[0240] [Table 17]

[0241]

[0242] The NNPFA syntax structure of Table 17 may be signaled in the form of an SEI message. The SEI message signaling the NNPFA syntax structure of Table 17 may be referred to as an NNPFA SEI message.

[0243] The NNPFASEI message enables or disables the possible use of the target neural network post-processing filter identified by nnpfa_target_id for post-processing filtering of a set of pictures.

[0244] If the post-processing filters are used for different purposes or filter different color components, there may be multiple NNPFASEI messages for the same picture.

[0245] nnpfa_target_id may specify a target neural network post-processing filter that is associated with the current picture and specified by one or more NNPFC SEI messages with nnpfc_id equal to nnfpa_target_id.

[0246] The value of nnpfa_target_id should be between 0 and 2 32 -2 (including 0 and 2 32 -2). You can keep the range of 256 to 511 (including 256 and 511) and 2 31 to 2 32 -2 (including 2 31 and 2 32 -2) are for future use. Decoders should ignore nnpfa_target_id values in the range of 256 to 511 (inclusive) or in the range of 2 31 to 2 32 -2 (including 2 31 and 2 32 -2) within the scope of NNPFASEI message.

[0247] Unless one or both of the following conditions are true, an NNPFASEI message with a specific value of nnpfa_target_id shall not be present in the current picture unit (PU): Here, a PU may be a set of NAL units containing the VCL NAL units of a coded picture and their associated non-VCL VAL units.

[0248] - Within the current CLVS, there is an NNPFC SEI message with nnpfc_id equal to a specific value of nnpfa_target_id present in a PU (picture unit) preceding the current PU (picture unit) in decoding order.

[0249] - There is an NNPFC SEI message with nnpfc_id equal to a specific value of nnpfa_target_id of the current PU (picture unit).

[0250] If a PU (picture unit) includes both an NNPFC SEI message with a specific value of nnpfc_id and an NNPFA SEI message with nnpfa_target_id equal to the specific value of nnpfc_id, the NNPFC SEI message shall precede the NNPFA SEI message in decoding order.

[0251] nnpfa_cancel_flag equal to 1 may specify that the persistence of the target neural network post-processing filter set by any previous NNPFASEI message with the same nnpfa_target_id as the current SEI message is canceled. That is, the target neural network post-processing filter is no longer used unless it is activated by another NNPFASEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 0. nnpfa_cancel_flag equal to 0 may specify that nnpfa_persistence_flag follows.

[0252] nnpfa_persistence_flag can specify the persistence of the target neural network post-processing filter of the current layer. nnpfa_persistence_flag equal to 0 can specify that the target neural network post-processing filter can be used for post-processing filtering only for the current picture. nnpfa_persistence_flag equal to 1 can specify that the target neural network post-processing filter can be used for post-processing filtering for the current picture and all subsequent pictures in the current layer in output order until one or more of the following conditions are true:

[0253] - A new CLVS starts for the current layer.

[0254] - End of bitstream.

[0255] - Pictures in the current layer associated with an NNPFA SEI message having the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1 are output after the current picture in output order.

[0256] The target neural network post-processing filter is not applied to subsequent pictures within the current layer associated with an NNPFA SEI message with the same nnpfa_target_id and nnpfa_cancel_flag equal to 1 as the current SEI message.

[0257] Post-Filter Tips

[0258] The syntax structure of the post-filter prompt is shown in Table 18.

[0259] [Table 18]

[0260]

[0261] The post-filter hint syntax structure of Table 18 may be signaled in the form of a SEI message. The SEI message that signals the post-filter hint syntax structure of Table 18 may be referred to as a post-filter hint SEI message.

[0262] The post-filter hint SEI message may provide post-filter coefficients or correlation information for design of a post-filter to potentially use in post-processing a set of decoded and output pictures to obtain improved display quality.

[0263] filter_hint_cancel_flag equal to 1 may specify that the SEI message cancels the persistence of the previous post-filter hint SEI message in the output order applied to the current layer. filter_hint_cancel_flag equal to 0 may specify that the post-filter hint message follows.

[0264] filter_hint_persistence_flag may specify the persistence of the post-filter hint SEI message for the current layer. filter_hint_persistence_flag equal to 0 may specify that the post-filter hint applies only to the currently decoded picture. filter_hint_persistence_flag equal to 1 may specify that the post-filter hint SEI message applies to the currently decoded picture and persists for all subsequent pictures in the current layer in output order until one or more of the following conditions are true:

[0265] - A new CLVS starts for the current layer.

[0266] - End of bitstream.

[0267] - The pictures in the current layer of the AU associated with the post-filter hint SEI message are output after the current picture in output order.

[0268] filter_hint_size_y can specify the vertical size of the filter coefficient or correlation array. The value of filter_hint_size_y should be in the range of 1 to 15.

[0269] filter_hint_size_x can specify the horizontal size of the filter coefficient or correlation array. The value of filter_hint_size_x should be in the range of 1 to 15.

[0270] filter_hint_type may specify the type of filter hint sent, as shown in Table 19. The value of filter_hint_type shall be in the range 0 to 2. filter_hint_type equal to 3 shall not be present in the bitstream. The decoder shall ignore the post-filter hint SEI message with filter_hint_type equal to 3.

[0271] [Table 19]

[0272] value describe 0 2D-FIR filter coefficients 1 1D-FIR filter coefficients 2 Cross-correlation matrix

[0273] filter_hint_chroma_coeff_present_flag equal to 1 can specify that the filter coefficients for chroma are present. filter_hint_chroma_coeff_present_flag equal to 0 can specify that the filter coefficients for chroma are not present. filter_hint_value[cIdx][cy][cx] can specify the filter coefficients or the elements of the cross-correlation matrix between the original signal and the decoded signal with 16-bit precision. The value of filter_hint_value[cIdx][cy][cx] should be between -2 31 +1 to 2 31 -1 (including -2 31 +1 and 2 31 -1). cIdx can specify the relevant color component, cy represents the counter in the vertical direction, and cx represents the counter in the horizontal direction. Depending on the value of filter_hint_type, the following applies:

[0274] - If filter_hint_type is equal to 0, the coefficients of a two-dimensional finite impulse response (FIR) filter of size filter_hint_size_y*filter_hint_size_x may be sent.

[0275] Otherwise, if filter_hint_type is equal to 1, the filter coefficients of two 1-dimensional FIR filters can be sent. In this case, filter_hint_size_y should be equal to 2. The index cy equal to 0 specifies the filter coefficients of the horizontal filter and cy equal to 1 specifies the filter coefficients of the vertical filter. In the filtering process, the horizontal filter is applied first, and the result is filtered by the vertical filter.

[0276] - Otherwise (filter_hint_type equals 2), the sent hint can specify the cross-correlation matrix between the original signal s and the decoded signal s'.

[0277] The normalized cross-correlation matrix of the correlated color component identified by cIdx of size filter_hint_size_y * filter_hint_size_x can be defined as in Equation 4.

[0278] [Equation 4]

[0279]

[0280] In Equation 4, s represents the sample array of the color component cIdx of the original picture, s' represents the corresponding array of the decoded picture, h represents the vertical height of the correlated color component, w represents the horizontal width of the correlated color component, bitDepth represents the bit depth of the color component, OffsetY equals (filter_hint_size_y >> 1), OffsetX equals (filter_hint_size_x >> 1), the range of cy is 0 <= cy < filter_hint_size_y, and the range of cx is 0 <= cx < filter_hint_size_x.

[0281] The decoder can derive a Wiener post-filter from the cross-correlation matrix of the original signal and the decoded signal and the auto-correlation matrix of the decoded signal.

[0282] Related technical issues

[0283] The problems of the related art according to the present disclosure are as follows.

[0284] The NNPFC SEI message can provide a neural network post-filter, and the NNPFA SEI message can provide the activation of the post-filter specified in the NNPFC SEI message to the picture set. The NNPFA SEI message can specify that the target post-filter is only applied to the current picture or applied to the current picture and subsequent pictures in the output order until one of the following events occurs.

[0285] - A new CLVS of the current layer starts.

[0286] - The bitstream ends.

[0287] - Pictures in the current layer associated with an NNPFA SEI message having the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1 are output after the current picture in output order.

[0288] Signaling designs (including and persistent NNPFC and NNPFA) may be used in applications such as Figure 5 The situation in which this causes problems. Figure 5 In the example of

[15] , a new NNPFC SEI message may be present in an AU containing a PU (picture unit) with POC3. This may create a situation where it is unclear whether the neural network post-filter is applied to the PU (picture unit) containing POC3. If applied, NNPFC is applied after the previous NNPFA enabled base NNPFC SEI, but an update is provided.

[0289] One possible solution to this problem is to specify that the neural network post-filter applied to the picture is the neural network post-filter in the NNPFC that immediately precedes the NNPFA in output order. Figure 6 As shown in FIG, since the presence of NNPFC SEI does not cancel the persistence of NNPFA SEI, NNPFC with ID 1 can still be applied. That is, the original (basic) filter can be applied to the PU containing POC3.

[0290] Other problems of the related art according to the present disclosure are as follows.

[0291] The Neural Network Post Filter Characteristics (NNPFC) SEI message includes a flag called nnpfc_formatting_and_purpose_flag. The nnpfc_formatting_and_purpose_flag indicates the presence of a syntax element related to a description that may include formatting, purpose, and / or complexity information of the filter. The presence of this flag is specified as follows:

[0292] When this SEI message is the first NNPFC SEI message with a particular nnpfc_id value in the current CLVS in decoding order, nnpfc_formatting_and_purpose_flag shall be equal to 1. When this SEI message is not the first NNPFC SEI message with a particular nnpfc_id value in the current CLVS in decoding order, nnpfc_formatting_and_purpose_flag shall be equal to 0.

[0293] However, if the purpose of the above constraint is simply to specify the purpose of the filter and the presence of signaling format information, then the above constraint may not be necessary, as it can be simply derived based on whether the SEI message with a specific nnpfc_id is the first SEI message. On the other hand, if nnpfc_formatting_and_purpose_flag also has the purpose of indicating whether the SEI message contains a basic neural network post filter (NNPF), then it should be designed to reflect such a purpose.

[0294] Other problems of the related art according to the present disclosure are as follows.

[0295] The aforementioned constraint on the nnpfc_formatting_and_purpose_flag has at least one problem: it prohibits repeating the NNPFC SEI message containing the underlying neural network post-processing filters. When repeating an SEI message, the constraint content is generally the same. However, the above constraint prohibits signaling of formatting, purpose, and complexity in repeated SEI messages.

[0296] Even if the above issues are fixed, regarding the repeated NNPFC SE I There are more issues to be resolved, as follows:

[0297] Question 1. Is it allowed to repeat an NNPFC SEI message regardless of whether the NNPFC SEI message contains the underlying neural network post-processing filter or an update of the underlying neural network post-processing filter? If not, which NNPFC SEI message is allowed to be repeated?

[0298] Question 2. If an NNPFC SEI message containing a base neural network post-processing filter can be repeated, can the NNPFC SEI message be repeated after the filter is updated by the presence of other NNPFC SEI messages (i.e., NNPFC SEI messages with the same nnpfc_id)?

[0299] Question 3. If the answer to Question 2 above is yes, when a repeated NNPFC SEI message containing a base neural network post-processing filter follows another NNPFC SEI message with the same nnpfc_id containing an update to the neural network post-processing filter, does the repeated SEI message overrule the update (i.e., it cancels the update)?

[0300] The semantics of the NNPFC SEI message should clarify all the issues described above.

[0301] Implementation Method

[0302] Hereinafter, NNPFC may be signaled in the form of an SEI message as the NNPFC syntax structure of Tables 1 and 2. In this case, NNPFC may be an NNPFC SEI message. NNPFA may be signaled in the form of an SEI message as the NNPFA syntax structure of Table 17. In this case, NNPFA may be an NNPFA SEI message. Postfilter hint may be signaled in the form of an SEI message as the postfilter hint syntax structure of Table 18. In this case, postfilter hint may be a postfilter hint SEI message.

[0303] Implementations according to the present disclosure may include various aspects to improve some or all of the above-mentioned problems. Each aspect may be applied alone or in combination of two or more.

[0304] Aspect 1: When the NNFPA SEI message is related to the current picture, the target neural network post-filter applied to the current picture / related to the current picture can be specified as the filter described / carried in the first NNPFC SEI message with the same id (e.g., nnpfc_id), which is immediately before the NNPFA SEI message in the output order.

[0305] Aspect 2: Instead of explicitly signaling a flag (ie, nnpfc_formatting_and_purpose_flag), the presence of syntax elements for purpose and format information in the NNPFC SEI message may be determined based on whether the NNPFC SEI message is the first SEI message with a specific id in the CLVS.

[0306] Aspect 3: The flag nnpfc_formatting_and_purpose_flag can be changed to nnpfc_base_filter_flag, and the semantics of the flag can be updated so that it specifies that the SEI message carries (includes) the base neural network post-processing filter. In this case, the flag can also specify whether the SEI message contains syntax about the purpose and formatting of the filter.

[0307] Aspect 4: A new flag may be specified to indicate whether the NNPFC SEI message contains a base neural network post filter (NNPF). This flag may be called nnpfc_base_flag.

[0308] Aspect 5: For the first NNPFC SEI message with a specific nnpfc_id in CLVS, nnpfc_base_flag may be restricted to be equal to 1.

[0309] Aspect 6: When an NNPFC SEI message with a specific nnpfc_id is not the first NNPFC SEI message in the CLVS and the value of nnpfc_base_flag is equal to 1, it means that the NNPFC SEI message is a repetition of the first NNPFC SEI message with the specified nnpfc_id. In this case, the repetition NNPFC SEI message should have the same content as the first NNPFC SEI message.

[0310] Aspect 7: The presence of nnpfc_base_flag can be conditional on the value of nnpfc_formatting_and_purpose_flag. When not present, the value of nnpfc_base_flag can be inferred to be equal to 0.

[0311] In the present disclosure, nnpfc_base_flag having a value of '1' may mean that the NNPFC SEI message includes a base NNPF. In the present disclosure, nnpfc_base_flag having a value of '0' may mean that the NNPFC SEI message does not include a base NNPF.

[0312] Refer to Table 1, Table 2, and Table 17, which describe the NNPFC syntax structure and semantics, and the NNPFA syntax structure and semantics.

[0313] According to the present disclosure, at least some of the above-mentioned problems can be solved by improving at least some of the above-mentioned NNPFC syntax structure and semantics, and NNPFA syntax structure and semantics by considering at least some of the above-mentioned aspects 1 to 7.

[0314] Implementation Method 1

[0315] Embodiment 1 of the present disclosure is related to the above-mentioned aspect 1. According to embodiment 1 of the present disclosure, when an NNPFA SEI message is related to a current picture, a target neural network post-processing filter for the current picture may be designated as a filter provided in a first NNPFC SEI message that immediately precedes the NNPFA SEI message in output order and has nnpfc_id equal to nnpfa_target_id.

[0316] Figure 7 The activations of the neural network post-processing filters are shown, which shows which filter is applied to which frame. Figure 7As shown in FIG, for a PU with POC2, the neural network post-processing filters used are still the filters in SEI A, even though the update filters exist before the PU (i.e., in SEI B). This is because the filters in SEI B have not yet been activated and the last activation was through SEI AA associated with SEI A. The filters in SEI B are activated through NNPFA SEI BA and applied to PUs with POC3.

[0317] Figure 8 The application of multiple NNPFC SEI messages with the same nnpfc_id is shown. Figure 8 In the example, the location of activating each NNPFC SEI message is different. Figure 8 As shown in , the update filter can be signaled at any point and the updated filter is activated only when a new NNPFA SEI message with nnpfa_cancel_flag equal to 0 is output.

[0318] According to this embodiment, the time point at which the updated filter is applied can be clearly specified. Alternatively, the filter to be applied at each time point can be clearly specified.

[0319] Implementation Method 2

[0320] Embodiment 2 of the present disclosure is related to the above-mentioned aspect 2. According to embodiment 2 of the present disclosure, signaling of nnpfc_formatting_and_purpose_flag in the NNPFC SEI message described with reference to Table 1 may be omitted.

[0321] In addition, in Table 1, conditional signaling based on nnpfc_formatting_and_purpose_flag can be performed based on a variable called, for example, nnpfcFormattingAndPurposePresentFlag. For example, if (nnpfc_formatting_and_purpose_flag) in Table 1 can be changed to if (nnpfcFormattingAndPurposePresentFlag). According to Embodiment 2 of the present disclosure, the relevant parts of Table 1 can be changed as shown in Table 20 below.

[0322] [Table 20]

[0323]

[0324] nnpfcFormattingAndPurposePresentFlag can be a variable rather than a syntax element and can be derived as follows:

[0325] If this SEI message is the first NNPFC SEI message with a specific nnpfc_id within the current CLVS in decoding order, then nnpfcFormattingAndPurposePresentFlag may be derived as 1. Otherwise, nnpfcFormattingAndPurposePresentFlag may be derived as 0.

[0326] According to the present embodiment, whether the syntax for the purpose and formatting information of the filter is present in the NNPFC SEI message may be determined without explicitly signaling nnpfc_formatting_and_purpose_flag, thereby saving bits required for explicit signaling.

[0327] Implementation 3

[0328] Embodiment 3 of the present disclosure is related to aspect 3. According to embodiment 3 of the present disclosure, nnpfc_base_filter_flag may be signaled instead of nnpfc_formatting_and_purpose_flag in the NNPFC SEI message described with reference to Table 1. According to embodiment 3 of the present disclosure, relevant parts of Table 1 may be changed as shown in Table 21 below.

[0329] [Table 21]

[0330]

[0331] Additionally, nnpfc_base_filter_flag equal to 1 may specify that this SEI message is the first NNPFC SEI message with a particular nnpfc_id within the current CLVS in decoding order. Additionally, there are syntax elements related to filter purpose, input formatting, output formatting, and complexity.

[0332] nnpfc_base_filter_flag equal to 0 may specify that this SEI message is not the first NNPFC SEI message with a specific nnpfc_id within the current CLVS in decoding order. Additionally, there are no syntax elements related to filter purpose, input formatting, output formatting, and complexity.

[0333] According to an embodiment, by signaling nnpfc_base_filter_flag, it may be specified whether the SEI message contains information about the filter purpose, formatting and complexity and whether the SEI message contains a base NNPF.

[0334] Implementation 4

[0335] Embodiment 4 of the present disclosure is related to at least one of aspects 4 to 6. According to embodiment 4 of the present disclosure, the NNPFC SEI message described with reference to Table 1 may include nnpfc_base_flag. According to embodiment 4 of the present disclosure, the relevant part of Table 1 may be changed as shown in Table 22 below.

[0336] [Table 22]

[0337]

[0338] According to this embodiment, nnpfc_base_flag equal to 1 can specify that the SEI message contains the base neural network post filter. nnpfc_base_flag equal to 0 can specify that the SEI message contains updates relative to the base neural network post filter. The value of nnpfc_base_flag shall be as follows:

[0339] - When the NNPFC SEI message is the first NNPFC SEI message with a specific nnpfc_id within the current CLVS in decoding order, the value of nnpfc_base_flag may be restricted to be equal to 1.

[0340] - When NNPFC SEI message nnpfcB is not the first NNPFC SEI message in decoding order within the current CLVS with a specific nnpfc_id and nnpfc_base_flag equal to 1, nnpfcB is a repetition of the first NNPFC SEI message nnpfcA in decoding order with the same nnpfc_id. In this case, the content of nnpfcB may be constrained to be the same as that of nnpfcA.

[0341] When the NNPFC SEI message is not the first NNPFC SEI message with a specific nnpfc_id within the current CLVS in decoding order and is not a repetition of the first NNPFC SEI message with a specific nnpfc_id, the following applies:

[0342] - This SEI message may define updates relative to the previous base post-processing filter with the same nnpfc_id in decoding order.

[0343] - This SEI message pertains to the current decoded picture and all subsequent decoded pictures of the current layer in output order, until the end of the current CLVS or the end of the decoded picture that follows the current decoded picture in output order within the current CLVS. Additionally, the SEI message may be associated with a subsequent NNPFC SEI message with a specific nnpfc_id within the current CLVS in decoding order (whichever is earlier). When the value of nnpfc_formatting_and_purpose_flag is equal to 1 and nnpfc_base_flag is equal to 0, the SEI message may contain a full update relative to the NNPFC SEI message containing the base neural network postfilter.

[0344] In addition, among the aspects described with reference to Table 1, the aspect in which the NNPFC SEI message is a repetition of the previous NNPFC SEI message in the current CLVS in decoding order, and the subsequent semantics apply as if this SEI message is the only NNPFC SEI message with the same content within the current CLVS may be omitted and not applied to implementation mode 4.

[0345] According to Embodiment 4 of the present disclosure, an NNPFC SEI message containing a basic neural network post-processing filter can be repeated. Furthermore, the repeated NNPFC SEI message can have the same content as the first NNPFC SEI message with the same nnpfc_id value in the current CLVS in decoding order. In addition to the aforementioned effects, it is expected to resolve additional issues of the aforementioned related art.

[0346] Implementation 5

[0347] Embodiment 5 of the present disclosure is related to at least one of aspects 4 to 7 above. According to embodiment 5 of the present disclosure, the NNPFC SEI message described with reference to Table 1 may include nnpfc_base_flag. In addition, nnpfc_base_flag may be conditionally signaled based on nnpfc_formatting_and_purpose_flag. According to embodiment 5 of the present disclosure, the relevant parts of Table 1 may be changed as shown in Table 23 below.

[0348] [Table 23]

[0349]

[0350] Embodiment 5 of the present disclosure differs in that nnpfc_base_flag may be conditionally signaled based on nnpfc_formatting_and_purpose_flag. Therefore, the description of Embodiment 4 is equally applicable to Embodiment 5. However, nnpfc_base_flag may not be signaled based on nnpfc_formatting_and_purpose_flag being 0. If nnpfc_base_flag does not exist, nnpfc_base_flag may be inferred to be a value of "0."

[0351] According to Embodiment 5 of the present disclosure, the effects according to Embodiment 4 can be expected. In addition, when nnpfc_formatting_and_purpose_flag is 0, bits required for signaling nnpfc_base_flag can be saved by omitting signaling of nnpfc_base_flag.

[0352] Hereinafter, an image encoding method and an image decoding method according to various embodiments of the present invention will be described. Figure 9 The image encoding method may be performed by the image encoding apparatus 100, and Figure 10 The image decoding method can be performed by the image decoding device 200.

[0353] Reference Figure 9 , the image encoding device may generate NNPF (neural network post filter) related information about a neural network post-processing filter to be applied to the current picture (S901). The image encoding device may encode the generated NNPF related information to generate an NNPF (neural network post filter) related SEI message (S902). The image encoding device may send the generated NNPF related SEI message to, for example, an image decoding device (S903). The image encoding device performs steps S901 and S902, and step S903 may be part of a sending method performed by a separate sending device.

[0354] Reference Figure 10 , the image decoding device may receive an NNPF-related SEI message related to a neural network post-processing filter to be applied to the current picture (S1001). The image decoding device may decode the received NNPF-related SEI message to reconstruct the NNPF-related information (S1002). The image decoding device may apply the NNPF to the current picture based on the reconstructed NNPF-related information (S1003). The above steps S1002 and S1003 may be performed under the condition that the NNPF is applied to the current picture.

[0355] In reference Figure 9and Figure 10 In the description, the NNPF-related SEI message may include an NNPFC SEI message, an NNPFASEI message and / or an SEI message about a post-filter hint according to the present disclosure. In addition, the NNPF-related information may refer to information signaled by a syntax element included in the NNPF-related SEI message.

[0356] As described above, the NNPFC SEI message may include syntax elements such as nnpfc_id, nnpfc_formatting_and_purpose_flag, etc. nnpfc_id is used to identify the post-processing filter and can be expressed as "filter identification information". nnpfc_formatting_and_purpose_flag specifies whether there are syntax elements related to filter properties such as filter purpose, input formatting, output formatting and / or complexity, and can be expressed as "filter property presence information". In addition, nnpfc_base_flag is information that specifies whether the SEI message includes a base neural network post-processing filter, and can be expressed as "base filter presence information".

[0357] The embodiments and aspects of the present disclosure regarding the above NNPFC SEI message can be applied to reference Figure 9 and Figure 10 The image encoding method and / or image decoding method according to the present disclosure is described.

[0358] Specifically, in the image encoding method and / or image decoding method according to the present disclosure, the NNPF-related SEI message includes a neural network post-filter characteristics (NNPFC) SEI message, which includes filter identification information and filter property existence information, and the NNPFC SEI message may also include basic filter existence information specifying whether the NNPFC SEI message includes a basic neural network post-processing filter.

[0359] In addition, in the image encoding method and / or image decoding method according to the present disclosure, based on the basic filter existence information having a first value, the NNPFC SEI message may include the basic neural network post-processing filter, and based on the basic filter existence information having a second value, the NNPFC SEI message may include an update relative to the basic neural network post-processing filter.

[0360] In addition, in the image encoding method and / or image decoding method according to the present disclosure, based on the fact that the NNPFC SEI message is the first NNPFC SEI message having specific filter identification information within the current CLVS (coding layer video sequence) in decoding order, the basic filter existence information can be limited to having a first value.

[0361] In addition, in the image encoding method and / or image decoding method according to the present disclosure, based on the fact that the NNPFC SEI message is not the first NNPFC SEI message having specific filter identification information in the current CLVS in decoding order and the basic filter existence information has a first value, the NNPFC SEI message may be a repetition of the first NNPFC SEI message having the same filter identification information in decoding order.

[0362] In addition, in the image encoding method and / or the image decoding method according to the present disclosure, the content of the NNPFC SEI message may be restricted to be the same as the content of the first NNPFC SEI message.

[0363] In addition, in the image encoding method and / or image decoding method according to the present disclosure, based on the fact that the NNPFC SEI message is not the first NNPFC SEI message having specific filter identification information within the current CLVS in decoding order and is not a repetition of the first NNPFC SEI message, the NNPFC SEI message may be an update relative to the previous base post-processing filter having the same filter identification information in decoding order.

[0364] In addition, in the image encoding method and / or image decoding method according to the present disclosure, based on the filter property existence information having a first value and the basic filter existence information having a second value, the NNPFC SEI message may include a complete update relative to the NNPFC SEI message including the basic neural network post-processing filter.

[0365] In addition, in the image encoding method and / or the image decoding method according to the present disclosure, basic filter presence information may be included in the NNPFC SEI message based on the filter property presence information.

[0366] For example, based on the filter characteristic presence information having a first value, the basic filter presence information may be included in the NNPFC SEI message.

[0367] Alternatively, for example, based on the filter characteristic presence information having the second value, the basic filter presence information may not be included in the NNPFC SEI message.

[0368] Alternatively, for example, based on the filter property presence information having the second value, the basic filter presence information may be derived as the second value.

[0369] In addition, a bitstream generated by the image encoding method according to the present disclosure may be stored in a non-transitory computer-readable recording medium.

[0370] Furthermore, a bit stream generated by the image encoding method according to the present disclosure may be transmitted to, for example, an image decoding device.

[0371] According to embodiments of the present disclosure, an NNPFC SEI message including a base neural network post-processing filter can be repeated. Furthermore, the repeated NNPFC SEI message can have the same content as the first NNPFC SEI message with the same filter identification information in the current CLVS in decoding order. Furthermore, when the filter attribute presence information is the second value, signaling of the base filter presence information can be omitted, thereby saving bits. The aforementioned effects are expected to resolve the additional problems of the aforementioned related art.

[0372] In the present disclosure, the filter property presence information and the basic filter presence information may have a first value or a second value. The first value may refer to 1 or "true". The second value may refer to 0 or "false".

[0373] Figure 11 is a diagram illustrating a content streaming system to which an embodiment of the present disclosure can be applied.

[0374] like Figure 11 As shown, the content streaming system to which the embodiments of the present disclosure are applied may mainly include an encoding server, a streaming server, a network server, a media storage, a user device, and a multimedia input device.

[0375] The encoding server compresses content input from a multimedia input device such as a smartphone, a camera, a camcorder, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, a camcorder, etc. directly generates a bitstream, the encoding server can be omitted.

[0376] A bitstream may be generated by the image encoding method or the image encoding apparatus to which the embodiments of the present disclosure are applied, and a streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0377] The streaming server transmits multimedia data to user devices via a network server based on user requests, and the network server serves as an intermediary for notifying users of services. When a user requests a desired service from the network server, the network server can pass it on to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server is used to control commands and responses between devices in the content streaming system.

[0378] The streaming server can receive content from a media storage and / or encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined time.

[0379] Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation, tablet PCs, touch-screen tablet PCs, ultrabooks, wearable devices (e.g., smart watches, smart glasses, head-mounted displays), digital TVs, desktop computers, digital signage, etc.

[0380] Each server in the content streaming system may operate as a distributed server, in which case data received from each server may be distributed.

[0381] The scope of the present disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) for enabling operations according to the methods of various embodiments to be performed on a device or computer, and non-transitory computer-readable media on which such software or commands are stored and which can be executed on a device or computer.

[0382] Industrial Applicability

[0383] The embodiments of the present disclosure may be used to encode or decode an image.

Claims

1. An image decoding method performed by an image decoding device, the image decoding method comprising: Receive a supplemental enhancement information (SEI) message related to a neural network post filter (NNPF) to be applied to the current picture; reconstructing NNPF related information based on the NNPF related SEI message; as well as Applying NNPF to the current picture based on the NNPF related information, The NNPF related SEI message includes a neural network post filter characteristic NNPFC SEI message, the NNPFC SEI message includes filter identification information and filter attribute existence information, and The NNPFC SEI message further includes basic filter existence information specifying whether the NNPFC SEI message includes a basic neural network post-processing filter.

2. The image decoding method according to claim 1, in, Based on the base filter presence information having a first value, the NNPFC SEI message includes the base neural network post-processing filter, and Wherein, based on the base filter existence information having a second value, the NNPFC SEI message includes an update relative to the base neural network post-processing filter.

3. The image decoding method according to claim 1, wherein: Based on the fact that the NNPFC SEI message is the first NNPFC SEI message having specific filter identification information in the current coding layer video sequence CLVS in decoding order, the basic filter presence information is restricted to have a first value.

4. The image decoding method according to claim 1, wherein: Based on the NNPFC SEI message not being the first NNPFC SEI message with specific filter identification information in the current CLVS in decoding order and the basic filter presence information having a first value, the NNPFC SEI message is a repetition of the first NNPFC SEI message with the same filter identification information in decoding order.

5. The image decoding method according to claim 4, wherein: The content of the NNPFC SEI message is restricted to be the same as the content of the first NNPFC SEI message. The image decoding method according to claim 1 , wherein: Based on the NNPFC SEI message not being the first NNPFC SEI message with specific filter identification information within the current CLVS in decoding order and not being a repetition of the first NNPFC SEI message, the NNPFC SEI message is an update relative to a previous base post-processing filter with the same filter identification information in decoding order.

7. The image decoding method according to claim 1, wherein: Based on the filter property presence information having a first value and the base filter presence information having a second value, the NNPFC SEI message includes a complete update relative to the NNPFC SEI message that includes the base neural network post-processing filter.

8. The image decoding method according to claim 1, wherein: Based on the filter property presence information, the basic filter presence information is included in the NNPFC SEI message.

9. The image decoding method according to claim 8, wherein: Based on the filter property presence information having a first value, the base filter presence information is included in the NNPFC SEI message.

10. The image decoding method according to claim 8, wherein: Based on the filter property presence information having a second value, the basic filter presence information is not included in the NNPFC SEI message.

11. The image decoding method according to claim 10, wherein: Based on the filter property presence information having a second value, the basic filter presence information is derived as a second value.

12. An image encoding method performed by an image encoding device, the image encoding method comprising: Generate information related to the neural network post filter NNPF to be applied to the current picture; as well as Generate NNPF related supplementary enhancement information SEI message based on the NNPF related information, The NNPF related SEI message includes a neural network post filter characteristic NNPFC SEI message, the NNPFC SEI message includes filter identification information and filter attribute existence information, and The NNPFC SEI message further includes basic filter existence information specifying whether the NNPFC SEI message includes a basic neural network post-processing filter.

13. A computer-readable recording medium storing a bit stream generated by an image encoding method, the image encoding method comprising: Generate information related to the neural network post filter NNPF to be applied to the current picture; as well as Generate NNPF related supplementary enhancement information SEI message based on the NNPF related information, The NNPF related SEI message includes a neural network post filter characteristic NNPFC SEI message, the NNPFC SEI message includes filter identification information and filter attribute existence information, and The NNPFC SEI message further includes basic filter existence information specifying whether the NNPFC SEI message includes a basic neural network post-processing filter.

14. A method for transmitting a bitstream generated by an image encoding method, the image encoding method comprising: Generate information related to the neural network post filter NNPF to be applied to the current picture; as well as Generate NNPF related supplementary enhancement information SEI message based on the NNPF related information, The NNPF related SEI message includes a neural network post filter characteristic NNPFC SEI message, the NNPFC SEI message includes filter identification information and filter attribute existence information, and The NNPFC SEI message further includes basic filter existence information specifying whether the NNPFC SEI message includes a basic neural network post-processing filter.