Image encoding / decoding methods, bitstream transmission methods, and non-forwardable computer readability.

VN126304APending Publication Date: 2026-06-15LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
VN · VN
Patent Type
Applications
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2024-10-02
Publication Date
2026-06-15

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality images, such as HD and UHD, leads to a significant increase in image data, resulting in higher transmission and storage costs. Current image compression technologies are not efficient enough to effectively manage these high-data volumes.

Method used

An improved image coding/decoding method that clarifies the meaning of Picture 0 and Source Picture timing information (SPTI SEI message) to reduce errors in decoders and enhance coding quality and efficiency. This method also includes a non-timely computer-readable recording medium for storing bitstreams generated by the image encoding method and a method for transmitting these bitstreams.

Benefits of technology

The proposed method achieves improved encoding/decoding efficiency, reduces errors in decoders, and enhances coding quality by clearly defining the timing information related to Picture 0 and SPTI SEI messages, thereby reducing transmission and storage costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure VN1202603401_0
    Figure VN1202603401_0
Patent Text Reader

Abstract

The invention relates to a method for encoding / decoding images, a method for transmitting bit streams, and a means of reading them by a non-forwarding computer. The method for decoding images under the invention comprises the following steps: acquiring supplemental enhancement information (SEI) messages and source picture timing information (SPTI); and reconstructing the image based on the SEI SPTI messages, wherein the SEI SPTI messages may include source picture timing information for the image combined with the SEI.
Need to check novelty before this filing date? Find Prior Art

Description

Video encoding / decoding method, method for transmitting bitstream, and recording medium storing bitstream The present disclosure relates to a video encoding / decoding method, a method for transmitting a bitstream, and a recording medium storing the bitstream, and relates to a video encoding / decoding method related to source picture timing, a method for transmitting a bitstream, and a recording medium storing the bitstream. Recently, demand for high-resolution, high-quality images, such as HD (High Definition) and UHD (Ultra High Definition) images, has been increasing across various fields. As image data becomes higher resolution and higher quality, the amount of information transmitted, or bits, increases relative to conventional image data. This increase in information or bits transmitted leads to increased transmission and storage costs. Accordingly, a highly efficient image compression technology is required to effectively transmit, store, and play high-resolution, high-quality image information. The present disclosure aims to provide a video encoding / decoding method and device with improved encoding / decoding efficiency. Additionally, the present disclosure aims to provide a method for coding and processing SPTI SEI messages. Additionally, the present disclosure aims to more clearly specify the source picture timing associated with the SPTI SEI message. Additionally, the present disclosure aims to clarify the meaning of picture 0 for deriving source picture timing. Additionally, the present disclosure aims to reduce decoder errors and improve coding quality and efficiency by clarifying the meaning of information related to picture 0 and source picture timing. In addition, the present disclosure aims to improve coding efficiency by clarifying constraints on picture 0. In addition, the present disclosure aims to provide a non-transitory computer-readable recording medium for storing a bitstream generated by an image encoding method according to the present disclosure. In addition, the present disclosure aims to provide a non-transitory computer-readable recording medium that stores a bitstream received and decoded by an image decoding device according to the present disclosure and used for restoring an image. In addition, the present disclosure aims to provide a method for transmitting a bitstream generated by an image encoding method according to the present disclosure. The technical problems to be achieved in the present disclosure are not limited to the technical problems mentioned above, and other technical problems not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present disclosure belongs from the description below. In an image decoding method according to one embodiment of the present disclosure, the method comprises the steps of obtaining a source picture timing information (SPTI) SEI (Supplemental Enhancement Information) message and restoring an image based on the SPTI SEI message, wherein the SPTI SEI message may include information on source picture timing for a picture associated with the SEI. In a video encoding method according to one embodiment of the present disclosure, the method comprises the steps of determining a source picture timing information (SPTI) SEI (Supplemental Enhancement Information) message and encoding a bitstream including the SPTI SEI message, wherein the SPTI SEI message may include information on source picture timing for a picture associated with the SEI. In addition, according to the present disclosure, a non-transitory computer-readable recording medium for storing a bitstream generated by an image encoding method according to the present disclosure can be provided. In addition, according to the present disclosure, a non-transitory computer-readable recording medium can be provided that stores a bitstream received and decoded by an image decoding device according to the present disclosure and used for restoring an image. Additionally, according to the present disclosure, a method for transmitting a bitstream generated by an image encoding method can be provided. The features briefly summarized above regarding the present disclosure are merely exemplary aspects of the detailed description of the present disclosure that follows and do not limit the scope of the present disclosure. According to the present disclosure, a video encoding / decoding method and device with improved encoding / decoding efficiency can be provided. In addition, according to the present disclosure, the semantics of information within an SPTI SEI message can be modified to enable clearer meaning delivery and reduce decoder errors. Additionally, according to the present disclosure, coding quality and efficiency can be improved by clearly deriving source picture timing associated with SPTI SEI messages. Additionally, according to the present disclosure, coding quality and efficiency can be improved by more clearly specifying the meaning of picture 0 for source picture timing. In addition, according to the present disclosure, a non-transitory computer-readable recording medium for storing a bitstream generated by an image encoding method according to the present disclosure can be provided. In addition, according to the present disclosure, a non-transitory computer-readable recording medium can be provided that stores a bitstream received and decoded by an image decoding device according to the present disclosure and used for restoring an image. Additionally, according to the present disclosure, a method for transmitting a bitstream generated by an image encoding method can be provided. The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the description below. FIG. 1 is a diagram schematically illustrating a video coding system to which an embodiment according to the present disclosure can be applied. FIG. 2 is a schematic diagram of an image encoding device to which an embodiment according to the present disclosure can be applied. FIG. 3 is a schematic diagram illustrating an image decoding device to which an embodiment according to the present disclosure can be applied. Figure 4 is a drawing for explaining an interleaved method for deriving a luma channel. FIG. 5 is a flowchart for explaining an image decoding method to which an embodiment according to the present disclosure can be applied. FIG. 6 is a flowchart for explaining an image encoding method to which an embodiment according to the present disclosure can be applied. FIG. 7 is a diagram exemplifying a content streaming system to which an embodiment according to the present disclosure can be applied. Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In describing embodiments of the present disclosure, detailed descriptions of known configurations or functions will be omitted if they are deemed to obscure the gist of the present disclosure. Furthermore, portions unrelated to the description of the present disclosure in the drawings have been omitted, and similar portions have been designated with similar reference numerals. In the present disclosure, when a component is said to be "connected," "coupled," or "connected" to another component, this may include not only a direct connection, but also an indirect connection in which another component exists in between. Furthermore, when a component is said to "include" or "have" another component, unless otherwise specifically stated, this does not exclude the other component, but rather implies that the other component may be included. In this disclosure, terms such as first, second, etc. are used solely to distinguish one component from another, and do not limit the order or importance of components unless specifically stated otherwise. Accordingly, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment. In this disclosure, distinct components are used to clearly illustrate their respective characteristics, and do not necessarily imply that the components are separated. That is, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed into multiple hardware or software units. Therefore, even if not specifically mentioned, such integrated or distributed embodiments are also included within the scope of this disclosure. In the present disclosure, the components described in various embodiments are not necessarily essential components, and some may be optional components. Therefore, embodiments comprising a subset of the components described in one embodiment are also within the scope of the present disclosure. Furthermore, embodiments including other components in addition to the components described in various embodiments are also within the scope of the present disclosure. The present disclosure relates to encoding and decoding of images, and terms used in the present disclosure may have their usual meanings commonly used in the technical field to which the present disclosure belongs, unless newly defined in the present disclosure. In the present disclosure, a "picture" generally refers to a unit representing one image of a specific time period, and a slice / tile is a coding unit that constitutes a part of a picture, and a single picture may be composed of one or more slices / tiles. In addition, a slice / tile may include one or more coding tree units (CTUs). In the present disclosure, "pixel" or "pel" may refer to the smallest unit that constitutes a picture (or image). Additionally, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luma component or only a pixel / pixel value of a chroma component. In the present disclosure, a "unit" may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. In some cases, the term "unit" may be used interchangeably with terms such as "sample array," "block," or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows. In the present disclosure, the "current block" may mean one of the following: a "current coding block," a "current coding unit," a "block to be encoded," a "block to be decoded," or a "block to be processed." When prediction is performed, the "current block" may mean a "current prediction block" or a "block to be predicted." When transformation (inverse transformation) / quantization (inverse quantization) is performed, the "current block" may mean a "current transformation block" or a "block to be transformed." When filtering is performed, the "current block" may mean a "block to be filtered." In the present disclosure, a "current block" may mean a block that includes both a luma component block and a chroma component block, or a "luma block of the current block," unless explicitly described as a chroma block. The luma component block of the current block may be explicitly expressed by including an explicit description of the luma component block, such as "luma block" or "current luma block." Additionally, the chroma component block of the current block may be explicitly expressed by including an explicit description of the chroma component block, such as "chroma block" or "current chroma block." In this disclosure, " / " and "," can be interpreted as "and / or". For example, "A / B" and "A, B" can be interpreted as "A and / or B". Additionally, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C." In this disclosure, "or" may be interpreted as "and / or." For example, "A or B" may mean 1) "A" only, 2) "B" only, or 3) "A and B." Alternatively, "or" in this disclosure may mean "additionally or alternatively." Overview of Video Coding Systems FIG. 1 is a diagram schematically illustrating a video coding system to which an embodiment according to the present disclosure can be applied. A video coding system according to one embodiment may include an encoding device (10) and a decoding device (20). The encoding device (10) may transmit encoded video and / or image information or data to the decoding device (20) in the form of a file or streaming through a digital storage medium or a network. An encoding device (10) according to one embodiment may include a video source generation unit (11), an encoding unit (12), and a transmission unit (13). A decoding device (20) according to one embodiment may include a reception unit (21), a decoding unit (22), and a rendering unit (23). The encoding unit (12) may be referred to as a video / image encoding unit, and the decoding unit (22) may be referred to as a video / image decoding unit. The transmission unit (13) may be included in the encoding unit (12). The reception unit (21) may be included in the decoding unit (22). The rendering unit (23) may include a display unit, and the display unit may be configured as a separate device or an external component. The video source generation unit (11) can obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source generation unit (11) can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, a virtual video / image can be generated through a computer, etc., in which case the video / image capture process can be replaced with a process of generating related data. The encoding unit (12) can encode input video / images. The encoding unit (12) can perform a series of procedures such as prediction, transformation, and quantization to improve compression and encoding efficiency. The encoding unit (12) can output encoded data (encoded video / image information) in the form of a bitstream. The transmission unit (13) can obtain encoded video / image information or data output in the form of a bitstream, and transmit it to the reception unit (21) of the decoding device (20) or another external object through a digital storage medium or a network in the form of a file or streaming. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit (13) may include an element for generating a media file through a predetermined file format, and may include an element for transmission through a broadcasting / communication network. The transmission unit (13) may be provided as a separate transmission device from the encoding device (12), and in this case, the transmission device may include at least one processor for obtaining encoded video / image information or data output in the form of a bitstream, and a transmission unit for transmitting it in the form of a file or streaming. The reception unit (21) can extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit (22). The decoding unit (22) can decode video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding unit (12). The rendering unit (23) can render the decrypted video / image. The rendered video / image can be displayed through the display unit. Overview of the video encoding device FIG. 2 is a schematic diagram illustrating an image encoding device to which an embodiment according to the present disclosure can be applied. As illustrated in FIG. 2, the image encoding device (100) may include an image segmentation unit (110), a subtraction unit (115), a transformation unit (120), a quantization unit (130), an inverse quantization unit (140), an inverse transformation unit (150), an addition unit (155), a filtering unit (160), a memory (170), an inter prediction unit (180), an intra prediction unit (185), and an entropy encoding unit (190). The inter prediction unit (180) and the intra prediction unit (185) may be collectively referred to as a “prediction unit.” The transformation unit (120), the quantization unit (130), the inverse quantization unit (140), and the inverse transformation unit (150) may be included in a residual processing unit. The residual processing unit may further include a subtraction unit (115). All or at least some of the plurality of components constituting the video encoding device (100) may be implemented as a single hardware component (e.g., an encoder or a processor) according to an embodiment. In addition, the memory (170) may include a decoded picture buffer (DPB) and may be implemented by a digital storage medium. The image segmentation unit (110) can segment an input image (or picture, frame) input to the image encoding device (100) into one or more processing units. For example, the processing unit may be called a coding unit (CU). The coding unit may be obtained by recursively segmenting a coding tree unit (CTU) or a largest coding unit (LCU) according to a QT / BT / TT (Quad-tree / binary-tree / ternary-tree) structure. For example, one coding unit may be segmented into a plurality of coding units of deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. For segmenting the coding unit, the quad-tree structure may be applied first, and the binary-tree structure and / or the ternary-tree structure may be applied later. The coding procedure according to the present disclosure may be performed based on the final coding unit that is no longer segmented. The maximum coding unit can be used directly as the final coding unit, and the coding unit of the lower depth obtained by dividing the maximum coding unit can be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and / or restoration described below. As another example, the processing unit of the coding procedure may be a prediction unit (PU) or a transformation unit (TU). The prediction unit and the transformation unit may each be divided or partitioned from the final coding unit. The prediction unit may be a unit of sample prediction, and the transformation unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from a transform coefficient. The prediction unit (inter-prediction unit (180) or intra-prediction unit (185)) can perform prediction on a block to be processed (current block) and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block or CU unit. The prediction unit can generate various information regarding the prediction of the current block and transmit the information to the entropy encoding unit (190). The information regarding the prediction can be encoded by the entropy encoding unit (190) and output in the form of a bitstream. The intra prediction unit (185) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block or may be located away from it, depending on the intra prediction mode and / or intra prediction technique. The intra prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of detail in the prediction direction. However, this is merely an example, and a greater or lesser number of directional prediction modes may be used depending on the settings. The intra prediction unit (185) may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks. The inter prediction unit (180) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different from each other. The temporal neighboring block may be called a collocated reference block, a collocated CU (colCU), etc. A reference picture including the above temporal neighboring blocks may be called a collocated picture (colPic). For example, the inter prediction unit (180) may construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter prediction unit (180) may use the motion information of neighboring blocks as the motion information of the current block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of the current block can be signaled by using the motion vector of the surrounding blocks as the motion vector predictor and encoding the motion vector difference and an indicator for the motion vector predictor. The motion vector difference can mean the difference between the motion vector of the current block and the motion vector predictor. The prediction unit can generate a prediction signal based on various prediction methods and / or prediction techniques described below. For example, the prediction unit can apply intra prediction or inter prediction to predict the current block, and can also apply intra prediction and inter prediction simultaneously. A prediction method that simultaneously applies intra prediction and inter prediction to predict the current block may be called combined inter and intra prediction (CIIP). In addition, the prediction unit may perform intra block copy (IBC) to predict the current block. Intra block copy can be used for video / image coding of content such as games, such as screen content coding (SCC). IBC is a method of predicting the current block using a previously restored reference block within the current picture located at a predetermined distance from the current block. When IBC is applied, the location of the reference block within the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives reference blocks within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in the present disclosure. The prediction signal generated through the prediction unit can be used to generate a restoration signal or a residual signal. The subtraction unit (115) can generate a residual signal (residual block, residual sample array) by subtracting the prediction signal (predicted block, predicted sample array) output from the prediction unit from the input image signal (original block, original sample array). The generated residual signal can be transmitted to the conversion unit (120). The transform unit (120) can apply a transform technique to the residual signal to generate transform coefficients. For example, the transform technique can include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is expressed as a graph. CNT refers to a transform obtained based on generating a prediction signal using all previously reconstructed pixels. The transform process can be applied to a pixel block having a square equal size, or can be applied to a block of a non-square variable size. The quantization unit (130) can quantize the transform coefficients and transmit them to the entropy encoding unit (190). The entropy encoding unit (190) can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients can be called residual information. The quantization unit (130) can rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and can also generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit (190) can perform various encoding methods, such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoding unit (190) can also encode, together or separately, information necessary for video / image restoration (e.g., values ​​of syntax elements) in addition to quantized transform coefficients. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in the form of a network abstraction layer (NAL) unit. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The signaling information, transmitted information and / or syntax elements mentioned in the present disclosure may be encoded through the encoding procedure described above and included in the bitstream. The above bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmitting the signal output from the entropy encoding unit (190) and / or a storage unit (not shown) for storing the signal may be provided as an internal / external element of the video encoding device (100), or the transmission unit may be provided as a component of the entropy encoding unit (190). The quantized transform coefficients output from the quantization unit (130) can be used to generate a residual signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (140) and inverse transformation unit (150), a residual signal (residual block or residual samples) can be restored. The addition unit (155) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit (180) or the intra prediction unit (185). When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit (155) can be called a reconstructor or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after filtering as described below. The filtering unit (160) can improve subjective / objective picture quality by applying filtering to the restoration signal. For example, the filtering unit (160) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and store the modified restoration picture in the memory (170), specifically, in the DPB of the memory (170). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (160) can generate various information regarding filtering and transmit the information to the entropy encoding unit (190), as described later in the description of each filtering method. The information regarding filtering may be encoded by the entropy encoding unit (190) and output in the form of a bitstream. The modified restored picture transmitted to the memory (170) can be used as a reference picture in the inter prediction unit (180). Through this, when inter prediction is applied, the image encoding device (100) can avoid prediction mismatch between the image encoding device (100) and the image decoding device, and can also improve encoding efficiency. The DPB in the memory (170) can store a modified reconstructed picture to be used as a reference picture in the inter prediction unit (180). The memory (170) can store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of blocks in a picture that has already been reconstructed. The stored motion information can be transferred to the inter prediction unit (180) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (170) can store reconstructed samples of reconstructed blocks in the current picture and transfer them to the intra prediction unit (185). Video Decoding Device Overview FIG. 3 is a schematic diagram illustrating an image decoding device to which an embodiment according to the present disclosure can be applied. As illustrated in FIG. 3, the image decoding device (200) may be configured to include an entropy decoding unit (210), an inverse quantization unit (220), an inverse transformation unit (230), an addition unit (235), a filtering unit (240), a memory (250), an inter prediction unit (260), and an intra prediction unit (265). The inter prediction unit (260) and the intra prediction unit (265) may be collectively referred to as a “prediction unit.” The inverse quantization unit (220) and the inverse transformation unit (230) may be included in a residual processing unit. All or at least some of the plurality of components constituting the video decoding device (200) may be implemented as a single hardware component (e.g., a decoder or processor) depending on the embodiment. In addition, the memory (170) may include a DPB and may be implemented by a digital storage medium. The video decoding device (200) that receives a bitstream including video / image information can restore the image by performing a process corresponding to the process performed in the video encoding device (100) of FIG. 2. For example, the video decoding device (200) can perform decoding using a processing unit applied in the video encoding device. Therefore, the processing unit for decoding may be, for example, a coding unit. The coding unit may be a coding tree unit or may be obtained by dividing a maximum coding unit. In addition, the restored image signal decoded and output by the video decoding device (200) can be reproduced through a reproduction device (not shown). The video decoding device (200) can receive a signal output from the video encoding device of FIG. 2 in the form of a bitstream. The received signal can be decoded through the entropy decoding unit (210). For example, the entropy decoding unit (210) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The video decoding device may additionally use information on the parameter set and / or the general constraint information to decode the image. The signaling information, received information, and / or syntax elements mentioned in the present disclosure can be obtained from the bitstream by being decoded through the decoding procedure. For example, the entropy decoding unit (210) can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values ​​of syntax elements required for image restoration and the quantized values ​​of transform coefficients for residuals. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element in the bitstream, determines a context model using information of the syntax element to be decoded and the decoding information of the surrounding block and the decoding target block or the information of the symbol / bin decoded in the previous step, and predicts the occurrence probability of the bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Information regarding prediction among the information decoded by the entropy decoding unit (210) is provided to the prediction unit (inter prediction unit (260) and intra prediction unit (265)), and the residual value on which entropy decoding is performed by the entropy decoding unit (210), i.e., quantized transform coefficients and related parameter information, can be input to the inverse quantization unit (220). In addition, information regarding filtering among the information decoded by the entropy decoding unit (210) can be provided to the filtering unit (240). Meanwhile, a receiving unit (not shown) that receives a signal output from an image encoding device may be additionally provided as an internal / external element of the image decoding device (200), or the receiving unit may be provided as a component of an entropy decoding unit (210). Meanwhile, the video decoding device according to the present disclosure may be referred to as a video / video / picture decoding device. The video decoding device may include an information decoder (video / video / picture information decoder) and / or a sample decoder (video / video / picture sample decoder). The information decoder may include an entropy decoding unit (210), and the sample decoder may include at least one of an inverse quantization unit (220), an inverse transformation unit (230), an addition unit (235), a filtering unit (240), a memory (250), an inter prediction unit (260), and an intra prediction unit (265). The inverse quantization unit (220) can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit (220) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the image encoding device. The inverse quantization unit (220) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients. In the inverse transform unit (230), the transform coefficients can be inversely transformed to obtain a residual signal (residual block, residual sample array). The prediction unit can perform a prediction on the current block and generate a predicted block containing prediction samples for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block based on the prediction information output from the entropy decoding unit (210), and can determine a specific intra / inter-prediction mode (prediction technique). The fact that the prediction unit can generate a prediction signal based on various prediction methods (techniques) described below is the same as that mentioned in the description of the prediction unit of the image encoding device (100). The intra prediction unit (265) can predict the current block by referring to samples within the current picture. The description of the intra prediction unit (185) can be equally applied to the intra prediction unit (265). The inter prediction unit (260) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (260) can construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes (techniques), and the information about the prediction can include information indicating the mode (technique) of inter prediction for the current block. The addition unit (235) can generate a restoration signal (restored picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (predicted block, prediction sample array) output from the prediction unit (including the inter prediction unit (260) and / or the intra prediction unit (265)). When there is no residual for the block to be processed, such as when the skip mode is applied, the predicted block can be used as the restoration block. The description of the addition unit (155) can be equally applied to the addition unit (235). The addition unit (235) can be called a restoration unit or a restoration block generation unit. The generated restoration signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after going through filtering as described below. The filtering unit (240) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (240) can apply various filtering methods to the restored picture to generate a modified restored picture, and store the modified restored picture in the memory (250), specifically, in the DPB of the memory (250). The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The (modified) reconstructed picture stored in the DPB of the memory (250) can be used as a reference picture in the inter prediction unit (260). The memory (250) can store motion information of a block from which motion information is derived (or decoded) within the current picture and / or motion information of blocks within a picture that has already been reconstructed. The stored motion information can be transferred to the inter prediction unit (260) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (250) can store reconstructed samples of reconstructed blocks within the current picture and transfer them to the intra prediction unit (265). In this specification, the embodiments described in the filtering unit (160), the inter prediction unit (180), and the intra prediction unit (185) of the image encoding device (100) can be applied to the filtering unit (240), the inter prediction unit (260), and the intra prediction unit (265) of the image decoding device (200) in the same or corresponding manner, respectively. General post-processing filtering process using NNPFs The input of the post-processing filtering process is a bitstream BitstreamToFilter, and the output can be a list of NNPF output pictures ListNnpfOutputPics. First, BitstreamToFilter is decoded, and the list ListNnpfOutputPics can be set to a list of cropped reconstructed pictures in output order resulting from decoding BitstreamToFilter. Next, the post-processing filtering process for one picture is repeatedly called for each cropped reconstructed picture in CroppedDecodedPictures in output order, and one or more NNPFs can be activated for the picture. The order of the pictures in ListNnpfOutputPics can be the output order. There must be exactly one picture associated with a specific output time instance in ListNnpfOutputPics. If multiple NNPFs are enabled for any given picture in CroppedDecodedPictures and only one NNPF can be selected to apply (any of the NNPFs can be selected), the above constraints can be applied regardless of which NNPF is selected to apply to the given picture. A post-processing filtering process is applied to each cropped reconstructed picture referred to as the current picture (existing in CroppedDecodedPictures), and one or more NNPFs may be activated. When applying an NNPF to the current picture, pictures filtered and / or interpolated by the NNPF may be generated by applying an NNPF process specific to the semantics of the NNPFC SEI message described below. When applying an NNPF to the current picture, the order of the pictures generated by the NNPF may be the output order by applying the NNPF stored in the output tensor of the NNPF. If the applied NNPF is the last NNPF applied to the current picture, the pictures generated by the NNPF and output by the NNPF process may be included in ListNnpfOutputPics in the same order as the order in which the pictures are stored in the output tensor of the NNPF. Neural-network post-filter characteristics (NNPFC) The combinations in Tables 1 to 3 represent the NNPFC syntax structure. The NNPFC syntax structures of Tables 1 to 3 may be signaled in the form of supplemental enhancement information (SEI) messages. An SEI message signaling the NNPFC syntax structures of Tables 1 to 3 may be referred to as an NNPFC SEI message. The NNPFC SEI message can specify a neural network that can be used as a post-processing filter. The use of specific post-processing filters for specific pictures can be indicated using the neural-network post-filter activation (NNPFA) SEI message. Here, "post-processing filter" and "post-filter" can have the same meaning. Using these SEI messages may require defining the following variables: - The width and height of the input picture can be cropped in luma sample units, and this width and height can be expressed as CroppedWidth and CroppedHeight, respectively. - CroppedYPic[idx], which is a luma sample array of input pictures, and CroppedCbPic[idx] and CroppedCrPic[idx], which are chroma sample arrays, can be used as inputs to NNPF if they exist, and the index idx can have a range from 0 to numInputPics-1. - BitDepth Y can represent the bit depth for the luma sample array of input pictures. - BitDepth C can represent the bit depth of the chroma sample arrays of the input pictures, if they exist. - ChromaFormatIdc can represent a chroma format identifier. - When the value of nnpfc_auxiliary_inp_idc is 1, the filtering strength control value array StrengthControlVal[idx] of the input pictures must contain real numbers in the range of 0 to 1, and the index idx can have a range of 0 to numInputPics-1. An input picture with index 0 may correspond to a picture whose NNPF defined by the NNPFC SEI message is activated by the NNPFA SEI message. An input picture whose index i is in the range of 1 to numInputPics - 1 may have precedence over an input picture with index i-1 in the output order. nnpfc_purpose may indicate the purpose of the NNPF as shown in Table 4. In Table 4, a non-zero ( nnpfc_purpose & bitMask ) may indicate that the NNPF in Table 4 has a purpose associated with the bitMask value. If nnpfc_purpose is greater than 0 and ( nnpfc_purpose & bitMask ) is equal to 0, the purpose associated with the bitMask value may not be applicable to the NNPF. If nnpfc_purpose is equal to 0, the NNPF may be used at the discretion of the application. The value of nnpfc_purpose may be constrained to be in the range of 0 to 63 in the bitstream. The values ​​in the range of 64 to 65 535 for nnpfc_purpose may be reserved for future use. Decoders should ignore NNPFC SEI messages with nnpfc_purpose in the range of 64 to 65 535. The variable chromaUpsamplingFlag indicating whether the purpose of the NNPF indicated by nnpfc_purpose is chroma upsampling, the variable resolutionResamplingFlag indicating whether the purpose of the NNPF indicated by nnpfc_purpose is resolution resampling, the variable pictureRateUpsamplingFlag indicating whether the purpose of the NNPF indicated by nnpfc_purpose is picture rate upsampling, the variable bitDepthUpsamplingFlag indicating whether the purpose of the NNPF indicated by nnpfc_purpose is bit depth upsampling, and the variable colorizationFlag indicating whether the purpose of the NNPF indicated by nnpfc_purpose is colorization can be derived as shown in Table 5 below. When ChromaFormatIdc is equal to 3, chromaUpsamplingFlag may be constrained to be equal to 0. When ChromaFormatIdc or chromaUpsamplingFlag is not equal to 0, colorizationFlag may be constrained to be equal to 0. When pictureRateUpsamplingFlag is equal to 1 and the input picture with index 0 is associated with a frame packing array SEI message with fp_arrangement_type equal to 5, all input pictures may be associated with frame packing array SEI messages with fp_arrangement_type equal to 5 and fp_current_frame_is_frame0_flag equal to 0. nnpfc_id may contain an identification number that can be used to identify the NNPF. The nnpfc_id value is between 0 and 2. 32 - Must be in the range of 2. 256 to 511 and 2 31 Inland 232 - nnpfc_id values ​​in the range 256 to 511 may be reserved for future use. Decoders may use values ​​in the range 256 to 511 or 2 31 Inland 2 32 - NNPFC SEI messages with nnpfc_id in the range 2 must be ignored. If the NNPFC SEI message is the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS, the following may apply: - The above SEI message may represent a base NNPF. - The above SEI message may be associated with the current decoded picture and all subsequent decoded pictures of the current layer until the current CLVS ends, in output order. An NNPFC SEI message may be a repetition of a previous NNPFC SEI message within the current CLVS in decoding order, and subsequent semantics may apply as if this SEI message were the only NNPFC SEI message with the same content within the current CLVS. A value of 1 for nnpfc_base_flag may indicate that the SEI message refers to the base NNPF. A value of 0 for nnpfc_base_flag may indicate that the SEI message refers to an update related to the base NNPF. The following restrictions may apply to the value of nnpfc_base_flag: - If the NNPFC SEI message is the first NNPFC SEI message with a particular nnpfc_id value within the current CLVS in decoding order, the value of nnpfc_base_flag may be equal to 1. - If NNPFC SEI message nnpfcB is not the first NNPFC SEI message with a particular nnpfc_id value within the current CLVS in decoding order and the value of nnpfc_base_flag is equal to 1, then the NNPFC SEI message may correspond to a repetition of the first NNPFC SEI message nnpfcA with the same nnpfc_id in decoding order. That is, the payload content of nnpfcB may be identical to the payload content of nnpfcA. If the value of nnpfc_base_flag is 0, the following restrictions may apply: - This SEI message can define updates relative to a primary NNPF preceding it in decoding order using the same nnpfc_id value. Updates are not cumulative, but each update can be applied to the primary NNPF, which is the NNPF specified in the first NNPFC SEI message in decoding order with a particular nnpfc_id value within the current CLVS. The NNPF defined by this SEI message can be obtained by applying updates defined by SEI messages relative to the primary NNPF with the same nnpfc_id value. - This SEI message may be associated with the currently decoded picture and all subsequent decoded pictures up to the end of the current CLVS in output order of the current layer, excluding decoded pictures that follow the current decoded picture in output order within the current CLVS. This SEI message may be associated with subsequent NNPFC SEI messages in decoding order that have nnpfc_base_flag equal to 0 and have an earlier value of a specific nnpfc_id within the current CLVS. A value of nnpfc_mode_idc of 0 may indicate that the SEI message contains a bitstream representing a base NNPF (if the value of nnpfc_base_flag is 1) or an update to a base NNPF with the same nnpfc_id value (if the value of nnpfc_base_flag is 0). If the value of nnpfc_base_flag is 1, an nnpfc_mode_idc equal to 1 may indicate that the base NNPF associated with the value of nnpfc_id is associated with a neural network identified by a URI, where the URI may be indicated by nnpfc_uri in the format indicated by the tag URI nnpfc_tag_uri. When the value of nnpfc_base_flag is 0, nnpfc_mode_idc equal to 1 may indicate that updates associated with the base NNPF of the same nnpfc_id value are defined by a URI, where the URI may be indicated by nnpfc_uri in the format indicated by the tag URI nnpfc_tag_uri. The value of nnpfc_mode_idc may be constrained to range from 0 to 1 in the bitstream. Values ​​of nnpfc_mode_idc in the range 2 to 255 may be reserved for future use and may not be present in the bitstream. Decoders must ignore NNPFC SEI messages with nnpfc_mode_idc in the range 2 to 255. Values ​​of nnpfc_mode_idc greater than 255 may not be present in the bitstream and may not be reserved for future use. nnpfc_reserved_zero_bit_a may be restricted to have a value equal to 0 due to bitstream restrictions. Decoders may be restricted to ignore NNPFC SEI messages where the value of nnpfc_reserved_zero_bit_a is not 0. The nnpfc_tag_uri may contain a tag URI with syntax and semantics specified in IETF RFC 4151 that identifies a neural network used as the primary NNPF or an update to the primary NNPF using the nnpfc_id value specified by the nnpfc_uri. The nnpfc_tag_uri allows for uniquely identifying the format of the neural network data specified by the nnrpf_uri without the need for a central registry. An nnpfc_tag_uri equal to "tag:iso.org,2023:15938-17" may indicate that the neural network data identified by the nnpfc_uri is ISO / IEC 15938-17 compliant. nnpfc_uri may contain a URI with syntax and semantics specified in IETF Internet Standard 66 that identifies the neural network used as the default NNPF or an update related to the default NNPF using the same nnpfc_id value. A value of 1 for nnpfc_property_present_flag may indicate the presence of syntax elements related to filter purpose, input formatting, output formatting, and complexity. A value of 0 for nnpfc_property_present_flag may indicate the absence of syntax elements related to filter purpose, input formatting, output formatting, and complexity. If the value of nnpfc_base_flag is 1, nnpfc_property_present_flag may be constrained to have a value of 1. If the value of nnpfc_property_present_flag is 0, the values ​​of all syntax elements that can be present only when the value of nnpfc_property_present_flag is 1 may be inferred to be equal to the values ​​of their corresponding syntax elements in the NNPFC SEI message that contains the base NNPF to which the SEI message provides updates. If the NNPFC SEI message nnpfcCurr is not the first NNPFC SEI message with a particular nnpfc_id value within the current CLVS in decoding order, is not a repeat of the first NNPFC SEI message with a particular nnpfc_id value (i.e., nnpfc_base_flag has a value of 0), and nnpfc_property_present_flag has a value of 1, the following restrictions may apply: - The value of nnpfc_purpose in an NNPFC SEI message must be identical to the value of nnpfc_purpose in the first NNPFC SEI message with a specific nnpfc_id value within the current CLVS in decoding order. - The values ​​of the syntax elements nnpfc_base_flag and preceding nnpfc_complexity_info_present_flag in the NNPFC SEI message must be identical to the values ​​of the corresponding syntax elements in the first NNPFC SEI message with a specific nnpfc_id value in the current CLVS in decoding order. - In decoding order, the nnpfc_complexity_info_present_flag in the first NNPFC SEI message with a specific nnpfc_id value within the current CLVS must be equal to 0 or all 1s, and the following may apply: (1) nnpfc_parameter_parameter_type_idc in nnpfcCurr must be the same as nnpfc_parameter_parameter_type_idc in nnpfcBase. (2) If nnpfc_log2_parameter_bit_length_minus3 exists in nnpfcCurr, nnpfc_log2_parameter_bit_length_minus3 in nnpfcCurr must be less than or equal to nnpfc_log2_parameter_bit_length_minus3 in nnpfcBase. (3) If nnpfc_num_parameters_idc in nnpfcBase is equal to 0, nnpfc_num_parameters_idc in nnpfcCurr must be equal to 0. (4) Otherwise (nnpfc_num_parameters_idc in nnpfcBase is greater than 0), nnpfc_num_parameters_idc in nnpfcCurr must be greater than 0 or less than or equal to nnpfc_num_parameters_idc in nnpfcBase. (5) If nnpfc_num_kmac_operations_idc in nnpfcBase is equal to 0, nnpfc_num_kmac_operations_idc in nnpfcCurr must be equal to 0. (6) Otherwise (nnpfc_num_kmac_operations_idc in nnpfcBase is greater than 0), nnpfc_num_kmac_operations_idc in nnpfcCurr must be greater than 0 and less than or equal to nnpfc_num_kmac_operations_idc in nnpfcBase. (7) If nnpfc_total_kilobyte_size in nnpfcBase is equal to 0, nnpfc_total_kilobyte_size in nnpfcCurr must be equal to 0. (8) Otherwise (nnpfc_total_kilobyte_size in nnpfcBase is greater than 0), nnpfc_total_kilobyte_size in nnpfcCurr must be greater than 0 or less than or equal to nnpfc_total_kilobyte_size in nnpfcBase. nnpfc_num_input_pics_minus1 + 1 may represent the number of decoded output pictures used as input to NNPF. The value of nnpfc_num_input_pics_minus1 may be constrained to be in the range of 0 to 63. If the value of pictureRateUpsamplingFlag is 1, the value of nnpfc_num_input_pics_minus1 may be constrained to be greater than 0. The variable numInputPics, which represents the number of pictures used as input to NNPF, can be derived as in mathematical expression 1. The value 1 of nnpfc_input_pic_output_flag[ i ] can indicate that NNPF generates the corresponding output picture for the i-th input picture. The value 0 of nnpfc_input_pic_output_flag[ i ] can indicate that NNPF does not generate the corresponding output picture for the i-th input picture. If the value of nnpfc_num_input_pics_minus1 is 0, nnpfc_input_pic_output_flag

[0000] can be inferred to be 1. If the value of pictureRateUpsamplingFlag is equal to 0 and the value of nnpfc_num_input_pics_minus1 is greater than 0, nnpfc_input_pic_output_flag[ i ] may be constrained to be equal to 1 for at least one value of i in the range from 0 to nnpfc_num_input_pics_minus1. nnpfc_input_pic_output_flag[ i ] may be referred to as nnpfc_input_pic_filtering_flag[ i ]. A value of 1 for nnpfc_absent_input_pic_zero_flag may indicate that NNPF expects that input pictures that are not present in the bitstream are represented by sample arrays with sample values ​​of 0. A value of 0 for nnpfc_absent_input_pic_zero_flag may indicate that NNPF expects that input pictures that are not present in the bitstream are represented by the input pictures that are closest in output order within the bitstream. nnpfc_out_sub_c_flag can indicate the values ​​of variables outSubWidthC and outSubHeightC when the value of chromaUpsamplingFlag is 0. A value of nnpfc_out_sub_c_flag of 1 can indicate that the value of outSubWidthC is 1 and the value of outSubHeightC is 1. A value of nnpfc_out_sub_c_flag of 0 can indicate that the value of outSubWidthC is 2 and the value of outSubHeightC is 1. If the value of ChromaFormatIdc is 2 and nnpfc_out_sub_c_flag is present, the value of nnpfc_out_sub_c_flag must be equal to 1. nnpfc_out_colour_format_idc can indicate the color format of the NNPF output and the values ​​of the variables outSubWidthC and outSubHeightC corresponding to it when the value of colorizationFlag is 1. A value of 1 in nnpfc_out_colour_format_idc can indicate that the color format of the NNPF output is a 4:2:0 format and that outSubWidthC and outSubHeightC are both equal to 2. A value of 2 in nnpfc_out_colour_format_idc can indicate that the color format of the NNPF output is a 4:2:2 format, that outSubWidthC is 2, and that outSubHeightC is 1. A value of 3 in nnpfc_out_colour_format_idc can indicate that the color format of the NNPF output is a 4:4:4 format and that outSubWidthC and outSubHeightC are both 1. The value of nnpfc_out_colour_format_idc may be constrained to not be equal to 0. If both chromaUpsamplingFlag and colourizationFlag are equal to 0, then outSubWidthC and outSubHeightC may be inferred to be equal to SubWidthC and SubHeightC, respectively. nnpfc_pic_width_num_minus1+1 and nnpfc_pic_width_denom_minus1+1 can represent the numerator and denominator, respectively, of the resampling ratio of the NNPF output picture width to CroppedWidth. The value of ( nnpfc_pic_width_num_minus1 + 1 ) divided by ( nnpfc_pic_width_denom_minus1 + 1 ) must be in the range of 1 / 16 to 16. If nnpfc_pic_width_num_minus1 and nnpfc_pic_width_denom_minus1 do not exist, both nnpfc_pic_width_num_minus1 and nnpfc_pic_width_denom_minus1 can be inferred to be 0. The variable nnpfcOutputPicWidth represents the width of the luma sample arrays of the picture(s) corresponding to the result of applying the NNPF identified by nnpfc_id to the input picture(s), and can be derived as in mathematical expression 2. The remainder of nnpfcOutputPicWidth divided by outSubWidthC must be 0. nnpfc_pic_height_num_minus1+1 and nnpfc_pic_height_denom_minus1+1 can represent the numerator and denominator of the resampling ratio of the NNPF output picture height to CroppedHeight, respectively. The value of ( nnpfc_pic_height_num_minus1 + 1 ) divided by ( nnpfc_pic_height_denom_minus1 + 1 ) must be in the range of 1 / 16 to 16. If nnpfc_pic_height_num_minus1 and nnpfc_pic_height_denom_minus1 do not exist, both nnpfc_pic_height_num_minus1 and nnpfc_pic_height_denom_minus1 can be inferred to be 0. The variable nnpfcOutputPicHeight represents the height of the luma sample arrays of the picture(s) corresponding to the result of applying the NNPF identified by nnpfc_id to the input picture(s), and can be derived as in mathematical expression 3. The remainder of nnpfcOutputPicHeight divided by outSubHeightC must be equal to 0. If nnpfc_pic_width_num_minus1, nnpfc_pic_width_denom_minus1, nnpfc_pic_height_num_minus1, and nnpfc_pic_height_denom_minus1 exist, at least one of the following restrictions may be true: - The value of nnpfcOutputPicWidth is not equal to CroppedWidth - The value of nnpfcOutputPicHeight is not equal to CroppedHeight nnpfc_interpolated_pics[ i ] may represent the number of interpolated pictures generated by NNPF between the i-th picture used as input to NNPF and the (i + 1)-th picture. The value of nnpfc_interpolated_pics[ i ] may be constrained to be within the range of 0 to 63. The value of nnpfc_interpolated_pics[ i ] may be constrained to be greater than 0 for at least one i within the range of 0 to nnpfc_num_input_pics_minus1 - 1. The variable NumInpPicsInOutputTensor, which represents the number of pictures that have corresponding input pictures and exist in the output tensor of the NNPF, the variable InpIdx[ idx ], which represents the input picture index of the idx-th picture that has corresponding input pictures and exists in the output tensor of the NNPF, and the variable numOutputPics, which represents the total number of pictures that exist in the output tensor of the NNPF, can be derived as shown in Table 6. A value of 1 for nnpfc_component_last_flag can indicate that the last dimension of the input tensor inputTensor for NNPF and the output tensor outputTensor resulting from NNPF are used for the current channel. A value of 0 for nnpfc_component_last_flag can indicate that the third dimension of the input tensor inputTensor for NNPF and the output tensor outputTensor resulting from NNPF are used for the current channel. The first dimensions of the input tensor and the output tensor can be used as batch indices used in some neural network frameworks. The formula in the semantics of this SEI message uses a batch size corresponding to a batch index equal to 0, but determining the batch size used as input for neural network inference can be determined by the implementation of postprocessing. For example, when the value of nnpfc_inp_order_idc is equal to 3 and the value of nnpfc_auxiliary_inp_idc is equal to 1, the input tensor can have 7 channels, including 4 luma matrices, 2 chroma matrices, and 1 auxiliary input matrix. In this case, the DeriveInputTensors() process can derive one for each of the 7 channels of the input tensor, and when a specific channel among these channels is processed, that channel can be referred to as the current channel during the process. nnpfc_inp_format_idc can indicate how to convert the sample values ​​of the input picture into the input values ​​of NNPF. When the value of nnpfc_inp_format_idc is 1, the input values ​​of NNPF can be real numbers, and the functions InpY( ) and InpC( ) can be expressed as in mathematical expression 4. If the value of nnpfc_inp_format_idc is 1, the input values ​​of NNPF are unsigned integer numbers, and the functions InpY() and InpC() can be derived as shown in Table 7. Variable inpTensorBitDepth Y can be derived from the syntax element nnpfc_inp_tensor_luma_bitdepth_minus8 described below. inpTensorBitDepth C can be derived from the syntax element nnpfc_inp_tensor_chroma_bitdepth_minus8 described below. Values ​​of nnpfc_inp_format_idc greater than 1 may be reserved for future use and must not be present in the bitstream. Decoders must ignore NNPFC SEI messages containing reserved values ​​of nnpfc_inp_format_idc. A value of nnpfc_auxiliary_inp_idc greater than 0 may indicate the presence of auxiliary input data in the input tensor of NNPF. A value of 0 for nnpfc_auxiliary_inp_idc may indicate that the auxiliary input data is not present in the input tensor. A value of 1 for nnpfc_auxiliary_inp_idc may indicate that the auxiliary input data is derived using the methods shown in Tables 10 to 12. The value of nnpfc_auxiliary_inp_idc must be in the range of 0 to 1 in the bitstream. Values ​​of 2 to 255 for nnpfc_auxiliary_inp_idc may be reserved for future use and are not present in the bitstream. Decoders should ignore NNPFC SEI messages with a value of nnpfc_auxiliary_inp_idc in the range of 2 to 255. Values ​​of nnpfc_auxiliary_inp_idc greater than 255 are not present in the bitstream and are not reserved for future use. nnpfc_inp_order_idc may indicate how to order the sample arrays of the input picture to form the input tensor for NNPF. The value of nnpfc_inp_order_idc must be in the range of 0 to 3 in the bitstream. Values ​​of 4 to 255 for nnpfc_inp_order_idc may be reserved for future use and are not present in the bitstream. Decoders should ignore NNPFC SEI messages with nnpfc_inp_order_idc in the range of 4 to 255. Values ​​of nnpfc_inp_order_idc greater than 255 are not present in the bitstream and are not reserved for future use. The value of nnpfc_inp_order_idc must not be 3 if the value of ChromaFormatIdc is not 1. If the value of ChromaFormatIdc is 0, the value of nnpfc_inp_order_idc must be 3. If the value of chromaUpsamplingFlag is 1, the value of nnpfc_inp_order_idc must not be 0. Table 8 describes the nnpfc_inp_order_idc values. nnpfc_inp_tensor_luma_bitdepth_minus8+8 can represent the bit depth of luma sample values ​​in the input integer tensor. inpTensorBitDepth Y can be derived as in mathematical equation 5. The value of nnpfc_inp_tensor_luma_bitlength_minus8 may be constrained to be in the range 0 to 24. nnpfc_inp_tensor_chroma_bitdepth_minus8+8 can represent the bit depth of chroma sample values ​​in the input integer tensor. The value of inpTensorBitDepthC can be derived as in Equation 6. The value of nnpfc_inp_tensor_chroma_bitdepth_minus8 may be constrained to be in the range 0 to 24. When the value of nnpfc_auxiliary_inp_idc is equal to 1, the variable strengthControlScaledVal can be derived as shown in Table 9. A patch may be a rectangular array of samples from a component of a picture (e.g., a luma or chroma component). The process DeriveInputTensors() for deriving an input tensor inputTensor for given vertical sample coordinates cTop and horizontal sample coordinates cLeft, which represent the upper left sample location of a sample patch contained in the input tensor, can be represented as a combination of Tables 10 to 12. A value of 0 for nnpfc_out_format_idc may indicate that the sample values ​​output by the NNPF are real numbers whose value range from 0 to 1 is linearly mapped to an unsigned integer value range from 0 to (1 << bitDepth) - 1, for the bit depth bitDepth required for subsequent post-processing or display. A value of 1 for nnpfc_out_format_idc may indicate that the luma sample values ​​output by the NNPF are in the range from 0 to (1 << outTensorBitDepth Y) - 1 can be an unsigned integer in the range 0 to ( 1 << ( 1 << outTensorBitDepth C ) - 1 may indicate an unsigned integer in the range 1 to 1. Values ​​of nnpfc_out_format_idc greater than 1 may be reserved for future use and are not present in the bitstream. Decoders must ignore NNPFC SEI messages containing reserved values ​​of nnpfc_out_format_idc. nnpfc_out_order_idc can indicate the output order of samples output from the NNPF. The value of nnpfc_out_order_idc must be in the range of 0 to 3 in the bitstream. Values ​​of nnpfc_out_order_idc from 4 to 255 may be reserved for future use and are not present in the bitstream. Decoders must ignore NNPFC SEI messages with nnpfc_out_order_idc in the range of 4 to 255. Values ​​of nnpfc_out_order_idc greater than 255 are not present in the bitstream and are not reserved for future use. If the value of chromaUpsamplingFlag is 1, the value of nnpfc_out_order_idc must not be equal to 0 or 3. If the value of colorizationFlag is 1, the value of nnpfc_out_order_idc must not be equal to 0. Table 13 describes the values ​​of nnpfc_out_order_idc. nnpfc_out_tensor_luma_bitdepth_minus8+8 can represent the bit depth of the luma sample values ​​in the output integer tensor. The value of nnpfc_out_tensor_luma_bitdepth_minus8 must be in the range 0 to 24. outTensorBitDepth Y The value of can be derived as in mathematical formula 7. nnpfc_out_tensor_chroma_bitdepth_minus8+8 can represent the bit depth of the chroma sample values ​​in the output integer tensor. The value of nnpfc_out_tensor_chroma_bitdepth_minus8 must be in the range of 0 to 24. outTensorBitDepth C The value of can be derived as in mathematical expression 8. If the value of bitDepthUpsamplingFlag is 1, the value of nnpfc_out_format_idc must be equal to 1, and at least one of the following restrictions may be true: - nnpfc_out_tensor_luma_bitdepth_minus8+8 exists, and outTensorBitDepth Y This BitDepth Y greater than - nnpfc_out_tensor_chroma_bitdepth_minus8+8 exists, and outTensorBitDepth C This BitDepth C greater than nnpfc_inp_tensor_luma_bitdepth_minus8, nnpfc_inp_tensor_chroma_bitdepth_minus8, nnpfc_out_tensor_luma_bitdepth_minus8, and nnpfc_out_tensor_chroma_bitdepth_minus8 exist and outTensorBitDepth Y This inpTensorBitDepth Y If greater than outTensorBitDepth C is inpTensorBitDepth C It must not be less than nnpfc_inp_tensor_luma_bitdepth_minus8, nnpfc_inp_tensor_chroma_bitdepth_minus8, nnpfc_out_tensor_luma_bitdepth_minus8, and nnpfc_out_tensor_chroma_bitdepth_minus8 exist, and outTensorBitDepth C This inpTensorBitDepth C If greater than outTensorBitDepth Y is inpTensorBitDepth Y It must not be less than. The process StoreOutputTensors( ) that derives sample values ​​in the filtered sample arrays FilteredYPic, FilteredCbPic, and FilteredCrPic from the output tensor for the given vertical sample coordinate cTop and horizontal sample coordinate cLeft, which represent the upper-left sample locations of the sample patches contained in the input tensor, can be represented as a combination of Tables 14 and 15. A value of 1 for nnpfc_separate_colour_description_present_flag may indicate that a distinct combination of colour primaries, transform properties, matrix coefficients, and scaling and offset values ​​for a picture due to NNPF is specified in the SEI message syntax structure. A value of 0 for nnfpc_separate_colour_description_present_flag may indicate that the combination of colour primaries, transform properties, matrix coefficients, and scaling and offset values ​​for a picture due to NNPF is the same as indicated in the VUI parameters for CLVS. nnpfc_colour_primaries may have the same semantics as defined for the vui_colour_primaries syntax element, except that: - nnpfc_colour_primaries can represent the primary colors of a picture that result from applying NNPF specified in the SEI message, rather than the primary colors used in CLVS. - If nnpfc_colour_primaries does not exist in the NNPFC SEI message, the value of nnpfc_colour_primaries can be inferred to be the same as the value of vui_colour_primaries. nnpfc_transfer_characteristics can have the same semantics as those defined for the vui_transfer_characteristics syntax element, except that: - nnpfc_transfer_characteristics can indicate the transfer characteristics of the picture that result from applying the NNPF specified in the SEI message, rather than the transfer characteristics used in CLVS. - If nnpfc_transfer_characteristics does not exist in the NNPFC SEI message, the value of nnpfc_transfer_characteristics can be inferred to be the same as the value of vui_transfer_characteristics. nnpfc_matrix_coeffs can describe the equations used to derive luma and chroma signals from green, blue, and red or Y, Z, and X primaries. The semantics of nnpfc_matrix_coeffs can be applied to pictures resulting from applying the NNPF specified in the SEI message, and BitDepth as shown for MatrixCoefficients. Y and BitDepth C Each one is outTensorBitDepth Y and outTensorBitDepth C may be identical to vui_matrix_coeffs. If nnpfc_matrix_coeffs does not exist in the NNPFC SEI message, the value of nnpfc_matrix_coeffs may be inferred to be identical to vui_matrix_coeffs. nnpfc_matrix_coeffs must not be equal to 0 unless the following conditions are met: - nnpfc_out_tensor_chroma_bitdepth_minus8 is equivalent to nnpfc_out_tensor_luma_bitdepth_minus8 - nnpfc_out_order_idc is equal to 2, outSubHeightC is equal to 1, and outSubWidthC is equal to 1 nnpfc_matrix_coeffs must not be equal to 8 unless one of the following conditions is met: - nnpfc_out_tensor_chroma_bitdepth_minus8 is equivalent to nnpfc_out_tensor_luma_bitdepth_minus8 - nnpfc_out_tensor_chroma_bitdepth_minus8 is equal to nnpfc_out_tensor_luma_bitdepth_minus8+1, nnpfc_out_order_idc is equal to 2, outSubHeightC is equal to 2, and outSubWidthC is equal to 1. nnpfc_full_range_flag may indicate scaling and offset values ​​applied with respect to matrix coefficients, as specified by nnpfc_matrix_coeffs. The semantics of nnpfc_full_range_flag may be the same as those specified for VideoFullRangeFlag. If nnpfc_full_range_flag is not present, the value of nnpfc_full_range_flag may be inferred to be equal to 0. A value of 1 for nnpfc_chroma_loc_info_present_flag may indicate the presence of the nnpfc_chroma_sample_loc_type_frame syntax element in the NNPFC SEI message. A value of 0 for nnpfc_chroma_loc_info_present_flag may indicate the absence of the nnpfc_chroma_sample_loc_type_frame syntax element in the NNPFC SEI message. If the value of colorizationFlag is 0 or nnpfc_out_colour_format_idc is not 1, the value of nnpfc_chroma_loc_info_present_flag may be constrained to be equal to 0. If nnpfc_chroma_sample_loc_type_frame is not equal to 6 and nnpfc_out_colour_format_idc is equal to 1, nnpfc_chroma_sample_loc_type_frame may indicate the locations of chroma samples of the output pictures. If nnpfc_chroma_sample_loc_type_frame is equal to 6 and nnpfc_out_colour_format_idc is equal to 1, it may indicate that the locations of the chroma samples are unknown, unspecified, or otherwise specified. The value of nnpfc_chroma_sample_loc_type_frame must be in the range 0 to 6, inclusive. nnpfc_overlap can indicate the number of overlapping horizontal and vertical samples of adjacent input tensors of NNPF. The value of nnpfc_overlap must be in the range of 0 to 16,383. A value of 1 for nnpfc_constant_patch_size_flag can indicate that NNPF accepts as input exactly the patch sizes indicated by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. A value of 0 for nnpfc_constant_patch_size_flag can indicate that NNPF accepts as input any arbitrary patch size with width inpPatchWidth and height inpPatchHeight. Here, the width of the extended patch (i.e., the patch plus the overlapping area) equal to inpPatchWidth+2*nnpfc_overlap is a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1+1+2*nnpfc_overlap, and the height of the extended patch equal to inpPatchHeight+2*nnpfc_overlap is a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1+1+2*nnpfc_overlap. npfc_patch_width_minus1+1 can represent the number of horizontal samples of the patch size required for the input of NNPF when the value of nnpfc_constant_patch_size_flag is 1. The value of nnpfc_patch_width_minus1 must be in the range of 0 to Min(32 766, CroppedWidth - 1). npfc_patch_height_minus1+1 can represent the number of vertical samples of the patch size required for the input of NNPF when the value of nnpfc_constant_patch_size_flag is 1. The value of nnpfc_patch_height_minus1 must be in the range of 0 to Min(32 766, CroppedHeight - 1). nnpfc_extended_patch_width_cd_delta_minus1+1+2*nnpfc_overlap can represent the common divisor of the allowed values ​​of the width of the extended patch required to be input to NNPF when the value of nnpfc_constant_patch_size_flag is 0. The value of nnpfc_extended_patch_width_cd_delta_minus1 must be in the range of 0 to Min(32 766, CroppedWidth - 1). nnpfc_extended_patch_height_cd_delta_minus1+1+2*nnpfc_overlap can represent the common divisor of the allowed values ​​of extended patch heights required to input to NNPF when the value of nnpfc_constant_patch_size_flag is 0. The value of nnpfc_extended_patch_height_cd_delta_minus1 must be in the range of 0 to Min(32 766, CroppedHeight - 1). The variables inpPatchWidth and inpPatchHeight can be set to the patch size width and patch size height respectively. If the value of nnpfc_constant_patch_size_flag is 0, the following can be applied: - The values ​​of inpPatchWidth and inpPatchHeight can be provided by external means or set by the post-processor. - The value of inpPatchWidth+2*nnpfc_overlap must be a positive integer multiple of nnpfc_extended_patch_width_cd_delta_minus1+1+2*nnpfc_overlap, and inpPatchWidth must be less than or equal to CroppedWidth. The value of inpPatchHeight+2*nnpfc_overlap must be a positive integer multiple of nnpfc_extended_patch_height_cd_delta_minus1+1+2*nnpfc_overlap, and inpPatchHeight must be less than or equal to CroppedHeight. Otherwise (if the value of nnpfc_constant_patch_size_flag is 1), the value of inpPatchWidth may be set equal to nnpfc_patch_width_minus1+1, and the value of inpPatchHeight may be set equal to nnpfc_patch_height_minus1+1. The variables outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, and outPatchCHeight can be derived as shown in Table 16. outPatchWidth * CroppedWidth must equal nnpfcOutputPicWidth * inpPatchWidth, and outPatchHeight * CroppedHeight must equal nnpfcOutputPicHeight * inpPatchHeight. nnpfc_padding_type can indicate the padding process when referencing sample positions outside the boundary of the input picture, as described in Table 17. The value of nnpfc_padding_type must be in the range 0 to 4, inclusive. The values ​​5 to 15 for nnpfc_padding_type may be reserved for future use and may not be present in the bitstream. A decoder must ignore NNPFC SEI messages with nnpfc_padding_type in the range 5 to 15. The values ​​of nnpfc_padding_type exceeding 15 shall not be present in the bitstream and may not be reserved for future use. nnpfc_luma_padding_val can indicate the luma value to use for padding when the value of nnpfc_padding_type is 4. The value of nnpfc_luma_padding_val must be in the range of 0 to (1 << BitDepthY) - 1. nnpfc_cb_padding_val can indicate the Cb value to be used for padding when the value of nnpfc_padding_type is 4. The value of nnpfc_cb_padding_val must be in the range of 0 to (1 << BitDepthC) - 1. nnpfc_cr_padding_val can indicate the Cr value to use for padding when the value of nnpfc_padding_type is 4. The value of nnpfc_cr_padding_val must be in the range of 0 to (1 << BitDepthC) - 1. The function InpSampleVal( y, x, picHeight, picWidth, croppedPic, cIdx ) with inputs of vertical sample position y, horizontal sample position x, picture height picHeight, picture width picWidth, sample array CroppedPic, and common index cIdx (0 for luma, 1 for Cb, and 2 for Cr) can return the value of SampleVal derived as shown in Table 18. NNPF PostProcessingFilter( ) may be a target NNPF derived from the semantics of the NNPFA SEI message. The process in Table 19 can be used to generate filtered and / or interpolated pictures by filtering them patch-wise using NNPF PostProcessingFilter(), where the filtered and / or interpolated pictures can include a Y sample array FilteredYPic, a Cb sample array FilteredCbPic, and a Cr sample array FilteredCrPic, as indicated by nnpfc_out_order_idc. An NNPF-generated picture with index i may contain sample arrays FilteredYPic[ i ], FilteredCbPic[ i ], and FilteredCrPic[ i ], if any. An NNPF-generated picture may not contain overlapping regions. The NNPF process consists of outputting NNPF-generated pictures in increasing index order according to the process defined in Table 19, where all NNPF-generated pictures interpolated by the NNPF are output and all NNPF-generated pictures corresponding to pictures input to the NNPF can be output as specified in the semantics of the NNPFA SEI message. A value of 1 for nnpfc_complexity_info_present_flag may indicate the presence of one or more syntax elements indicating the complexity of the NNPF associated with the nnpfc_id. A value of 0 for nnpfc_complexity_info_present_flag may indicate the absence of syntax elements indicating the complexity of the NNPF associated with the nnpfc_id. A value of 0 for nnpfc_parameter_type_idc may indicate that the neural network uses only integer parameters. A value of 1 for nnpfc_parameter_type_flag may indicate that the neural network can use either floating-point or integer parameters. A value of 2 for nnpfc_parameter_type_idc may indicate that the neural network uses only binary parameters. A value of 3 for nnpfc_parameter_type_idc may be reserved for future use and is not present in the bitstream. Decoders should ignore NNPFC SEI messages with a value of 3 for nnpfc_parameter_type_idc. The values ​​0, 1, 2, and 3 of nnpfc_log2_parameter_bit_length_minus3 can indicate that the network does not use parameters with bit lengths greater than 8, 16, 32, and 64, respectively. If nnpfc_parameter_type_idc exists and nnpfc_log2_parameter_bit_length_minus3 does not exist, the network may not use parameters with bit lengths greater than 1. nnpfc_num_parameters_idc can represent the maximum number of neural network parameters for NNPF in units of powers of 2048. A value of 0 for nnpfc_num_parameters_idc can indicate that the maximum number of neural network parameters is unknown. The value of nnpfc_num_parameters_idc must be in the range of 0 to 52. A value of nnpfc_num_parameters_idc greater than 52 must not be present in the bitstream. Decoders must ignore NNPFC SEI messages with nnpfc_num_parameters_idc greater than 52. If the value of nnpfc_num_parameters_idc is greater than 0, the variable maxNumParameters can be derived as in Equation 9. The number of neural network parameters in NNPF can be limited to be less than or equal to maxNumParameters. A value of nnpfc_num_kmac_operations_idc greater than 0 may indicate that the maximum number of multiply-accumulate operations per sample of the NNPF is less than or equal to nnpfc_num_kmac_operations_idc * 1 000. A value of nnpfc_num_kmac_operations_idc of 0 may indicate that the maximum number of multiply-accumulate operations of the network is unknown. The value of nnpfc_num_kmac_operations_idc may be between 0 and 2. 32 - Must be within the range of 2. A value of nnpfc_total_kilobyte_size greater than 0 may indicate the total size (in kilobytes) required to store the uncompressed parameters of the neural network. The total size in bits may be a number greater than or equal to the sum of the bits used to store each parameter. nnpfc_total_kilobyte_size may be the total size (in bits) divided by 8,000, rounded up. A value of 0 for nnpfc_total_kilobyte_size may indicate that the total size required to store the parameters for the neural network is unknown. The value of nnpfc_total_kilobyte_size may be between 0 and 2. 32 - Must be within the range of 2. A value of 0 for nnpfc_metadata_extension_num_bits may indicate that nnpfc_reserved_metadata_extension does not exist. A value of nnpfc_metadata_extension_num_bits greater than 0 may indicate the length (in bits) of nnpfc_reserved_metadata_extension. nnpfc_metadata_extension_num_bits must be equal to 0. Values ​​in the range 1 to 2048 for nnpfc_metadata_extension_num_bits are reserved for future use and are not present in the bitstream. Decoders may accept any value of nnpfc_metadata_extension_num_bits in the range 0 to 2048. Values ​​of nnpfc_metadata_extension_num_bits greater than 2 048 are not present in the bitstream and are reserved for future use. nnpfc_reserved_metadata_extension is not present in the bitstream. However, decoders must ignore the presence and value of nnpfc_reserved_metadata_extension. If nnpfc_reserved_metadata_extension is present, its length can be equal to nnpfc_metadata_extension_num_bits. nnpfc_reserved_zero_bit_b must be equal to 0 in the bitstream. Decoders must ignore NNPFC SEI messages with nnpfc_reserved_zero_bit_b not equal to 0. nnpfc_payload_byte[ i ] may contain the ith byte of the bitstream. For all existing values ​​of i, the byte sequence nnpfc_payload_byte[ i ] must be a complete bitstream conforming to ISO / IEC 15938-17. Neural-network post-filter activation (NNFPA) The syntax structure for NNFPA is shown in Table 20. The NNPFA syntax structure of Table 20 can be signaled in the form of an SEI message. An SEI message signaling the NNPFA syntax structure of Table 20 may be referred to as an NNPFA SEI message. The NNPFA SEI message can enable or disable the possible use of a target neural network post-processing filter (NNPF), identified by nnpfa_target_id, for post-processing filtering of a picture set. For a given picture with NNPF enabled, the target NNPF can be derived as follows. - If nnpfa_target_base_flag is 1, the target NNPF can be the base NNPF with the same nnpfc_id as nnpfa_target_id. - Otherwise (if nnpfa_target_base_flag is 0), the target NNPF may be the NNPF specified by the last NNPFC SEI message with nnpfc_id equal to nnpfa_target_id, where the last NNPFC SEI message may not be a repeat of the NNPFC SEI message containing the base NNPF and preceding the first VCL NAL unit of the current picture in decoding order. Multiple NNPFA SEI messages may exist for the same picture if NNPF is used for different purposes or filters different color components. nnpfa_target_id may indicate a target NNPF associated with the current picture and specified by one or more NNPFC SEI messages having an nnpfc_id equal to nnfpa_target_id. The value of nnpfa_target_id is 0 to 2. 32 - Must be within the range of 2. An NNPFA SEI message with an nnpfa_target_id of a particular value shall not exist in the current PU unless one or both of the following conditions are true: - Within the current CLVS, there exists an NNPFC SEI message with a specific value of nnpfa_target_id and the same nnpfc_id that exists within the PU that precedes the current PU in decoding order. - There is an NNPFC SEI message with an nnpfc_id equal to the nnpfa_target_id of a specific value of the current PU. If a PU contains both an NNPFC SEI message with a nnpfc_id of a particular value and an NNPFA SEI message with a nnpfa_target_id equal to a particular nnpfc_id, the NNPFC SEI message shall precede the NNPFA SEI message in decoding order. A value of 1 for nnpfa_cancel_flag may indicate that the persistence of the target NNPF established by any previous NNPFA SEI message with the same nnpfa_target_id as the current SEI message is canceled. That is, the target NNPF is no longer used unless it is activated by another NNPFA SEI message with the same nnpfa_target_id as the current SEI message and an nnpfa_cancel_flag equal to 0. A value of 0 for nnpfa_cancel_flag may indicate that nnpfa_target_base_flag, nnpfa_persistence_flag, and nnpfa_num_output_entries follow. A value of 1 for nnpfa_target_base_flag may indicate that the target NNPF is a base NNPF having an nnpfc_id equal to nnpfa_target_id. A value of 0 for nnpfa_target_base_flag may indicate that the target NNPF is an NNPF specified by the last NNPFC SEI message having an nnpfc_id equal to nnpfa_target_id, where the last NNPFC SEI message may not correspond to a repetition of an NNPFC SEI message that precedes the first VCL NAL unit of the current picture in decoding order and contains the base NNPF. nnpfa_persistence_flag can indicate the persistence of the target NNPF for the current layer. A value of 0 for nnpfa_persistence_flag can indicate that the target NNPF can only be used for post-processing filtering for the current picture. A value of 1 for nnpfa_persistence_flag can indicate that the target NNPF can be used for post-processing filtering for the current picture and all subsequent pictures in the current layer in output order until one or more of the following conditions are true: - A new CLVS for the current layer is started. - Bitstream ended - The picture in the current layer associated with the NNPFA SEI message with the same nnpfa_target_id as the current SEI message and nnpfa_cancel_flag equal to 1 is output after the current picture in the output order. The target NNPF does not apply to subsequent pictures within the current layer associated with an NNPFA SEI message that has the same nnpfa_target_id as the current SEI message and an nnpfa_cancel_flag equal to 1. nnpfcTargetPictures may be a set of pictures associated with the last NNPFC SEI message that precedes the current NNPFA SEI message in decoding order and has an nnpfc_id equal to nnpfa_target_id. nnpfaTargetPictures may be a set of pictures whose target NNPF is activated by the current NNPFA SEI message. Any picture included in nnpfaTargetPictures must also be included in nnpfcTargetPictures. nnpfa_num_output_entries can indicate the number of nnpfa_output_flag[ i ] syntax elements in the NNPFA SEI message. The value of nnpfa_output_flag[ i ] must be in the range 0 to NumInpPicsInOutputTensor. A value of 1 in nnpfa_output_flag[ i ] may indicate that the NNPF-generated picture corresponding to the input picture having the index InpIdx[ i ] is output by the NNPF process activated by the NNPFA SEI message. Here, the NNPF process may be specified by the semantics of the NNPFC SEI message. A value of 0 in nnpfa_output_flag[ i ] may indicate that the NNPF-generated picture corresponding to the input picture having the index InpIdx[ i ] is not output by the NNPF process activated by the NNPFA SEI message. If nnpfa_num_output_entries is less than NumInpPicsInOutputTensor, nnpfa_output_flag[ i ] can be inferred to be 1 for each value of i in the range nnpfa_num_output_entries to NumInpPicsInOutputTensor - 1. Neural-network post-filter group characteristics (NNPFGC) The syntax structure for NNFPGC is shown in Table 21. The NNPFGC syntax structure of Table 21 can be signaled in the form of an SEI message. An SEI message signaling the NNPFGC syntax structure of Table 21 may be referred to as an NNPFGC SEI message. The NNPFGC SEI message can indicate a neural network post-filter group. This SEI message can specify whether an NNPF group defines an NNPF cascade, or whether it defines NNPF groups of NNPF or NNPF cascades that replace each other. The use of an NNPF group of an NNPF cascade for a particular picture can be indicated using the Neural Network Post-Filter Group Activation (NNPFGA) SEI message. nnpfgc_id may contain an identification number that can be used to identify an NNPF group. The value of nnpfgc_id is 0 to 2. 32 - Must be within the range of 256 to 511 and nnpfgc_id must be within the range of 256 to 511 and 2 31 Inland 2 32 - nnpfgc_id in the range 256 to 511 may be reserved for future use. The decoder may return an nnpfgc_id in the range 256 to 511 or a 2 31 Inland 2 32 - NNPFGC SEI messages with nnpfgc_id in the range -2 may be restricted to be ignored. The value of nnpfgc_id shall not be equal to the value of nnpfgc_id of any NNPFGC SEI message existing within the same CLVS. If the value of nnpfgc_id of NNPFGC SEI message nnpfgcSeiA is equal to the value of nnpfgc_id of another NNPFGC SEI message nnpfgcSeiB existing within the same CLVS, then nnpfgcSeiA and nnpfgcSeiB shall be equal. A value of nnpfgc_grouping_type of 0 may indicate that this SEI message specifies a group of serial NNPFs. A value of nnpfgc_grouping_type of 1 may indicate that the NNPFs or NNPF groups identified by nnpfgc_member_id[ i ] are substituted for each other and the post-processor should select only one of them to apply. A value of nnpfgc_grouping_type of 2 may indicate that this SEI message is intended to be used jointly or in a substituted manner, with at most one NNPF active for any given picture. A value of nnpfgc_grouping_type of 3 may indicate that the NNPFs or NNPF groups identified by nnpfgc_member_id[ i ] are intended to be used in parallel. A value of 4 for nnpfgc_grouping_type may indicate that the NNPF or NNPF group identified by nnpfgc_member_id[ i ] is optional, i.e., may or may not be applied by the post-processor. The value of nnpfgc_grouping_type must be in the range 0 to 255. Values ​​of nnpfgc_grouping_type in the range 5 to 255 are reserved for future use and may be restricted from being present in the bitstream. Decoders may be restricted to ignore NNPFGC SEI messages with nnpfgc_grouping_type in the range 5 to 255. nnpfgc_purpose has the semantics of nnpfc_purpose, but with the exception that the semantics are specific to an NNPF group defined by the NNPFGC SEI message, rather than to an NNPF defined by the NNPFC SEI message. nnpfgc_num_members_minus2+2 can represent the number of NNPFs defined by the NNPFGC SEI message or the number of NNPF groups within an NNPF group. nnpfgc_member_id[ i ] can represent the ith member in the NNPF group defined by the NNPFGC SEI message as follows. - If there is an NNPF having an nnpfc_id equal to nnpfgc_member_id[ i ] defined in CLVS, the i-th member in the NNPF group defined by the NNPFGC SEI message may be an NNPF having an nnpfc_id equal to nnpfgc_member_id[ i ]. - Otherwise (if there is no NNPF with nnpfc_id equal to nnpfgc_member_id[ i ] defined in CLVS), the ith member in the NNPF group defined by the NNPFGC SEI message may be the NNPF with nnpfgc_id equal to nnpfgc_member_id[ i ]. If the nnpfgc_member_id[ i ] value refers to the nnpfgc_id value of an NNPFGC SEI message nnpfgcSei, the NNPFGC SEI message nnpfgcSei may be constrained to have nnpfgc_grouping_type equal to 0. If nnpfgc_grouping_type is equal to 0 or 2, there must exist an NNPF with nnpfc_id equal to nnpfgc_member_id[ i ] defined in CLVS. If nnpfgc_grouping_type is equal to 1, 3, or 4, there must exist an NNPF with nnpfc_id equal to nnpfgc_member_id[ i ], or an NNPF with nnpfgc_id equal to nnpfgc_member_id[ i ]. When nnpfgc_grouping_type is equal to 0, NNPFs having nnpfc_id equal to nnpfgc_member_id[ i ] can be executed serially in increasing order of i, as activated by NNPFGA SEI messages having nnpfga_target_id equal to nnpfgc_id.nnpfgc_complexity_info_present_flag, nnpfgc_parameter_type_idc, nnpfgc_log2_parameter_bit_length_minus3, nnpfgc_num_parameters_idc, nnpfgc_num_kmac_operations_idc, and nnpfgc_total_kilobyte_size have the semantics of nnpfc_complexity_info_present_flag, nnpfc_parameter_type_idc, nnpfc_log2_parameter_bit_length_minus3, nnpfc_num_parameters_idc, nnpfc_num_kmac_operations_idc, and nnpfc_total_kilobyte_size, respectively, but these semantics are specific to the NNPF defined in this SEI message, not the NNPF defined in the NNPFC SEI message. The difference is that if nnpfgc_grouping_type is 1, nnpfgc_complexity_info_present_flag must be equal to 0. Neural-network post-filter group activation (NNPFGA) The syntax structure for NNFPGA is shown in Table 22. The NNPFGA syntax structure of Table 22 can be signaled in the form of an SEI message. An SEI message signaling the NNPFGA syntax structure of Table 22 may be referred to as an NNPFGA SEI message. The NNPFGA SEI message can enable or disable the possible use of a target NNPFGC, identified by the nnpfga_target_id of NNPF groups, for post-processing filtering of a picture set. The nnpfgc_grouping_type for the identified NNPF group must be equal to 0 (serial) or 1 (alternative). If nnpfgc_grouping_type is equal to 1, each member of the group must have the same number of input pictures and NNPF output pictures. For a particular picture with NNPFGC enabled, the target NNPFG may be the NNPFG specified by the last NNPFGC SEI message that precedes the first VCL NAL unit of the current picture in decoding order and whose nnpfgc_id is equal to nnpfga_target_id. The NNPFs of the target NNPFG may be defined as NNPFC SEI messages having an nnpfc_id equal to the nnpfgc_member_id[i] value of the target NNPFG, which may be within the current picture unit or may precede the current picture in decoding order. Use of the NNPFGC SEI message may require the following definitions: - For input pictures that can be used as input to NNPFGC and have an index idx in the range 0 to numCandInputPics - 1, the input picture width InitCroppedWidth[ idx ] and height InitCroppedHeight[ idx ] in luma sample units. - For input pictures that can be used as input to NNPFGC and have an index idx in the range 0 to numCandInputPics - 1, a luma sample array InitCroppedYPic[ idx ], and chroma sample arrays InitCroppedCbPic[ idx ] and InitCroppedCrPic[ idx ]. - Bit depth for the luma sample array of candidate input pictures BitDepthY - Bit depth for the chroma sample array of candidate input pictures BitDepth C - Chroma format indicator indicated by ChromaFormatIdc - An array of filtering strength control values ​​StrengthControlVal[ idx ] containing real values ​​in the range 0 to 1, for input pictures with indices idx in the range 0 to numCandInputPics - 1, if nnpfc_auxiliary_inp_idc is equal to 1. A candidate input picture with index 0 can represent a picture for which NNPFGC is activated by the NNPFGA SEI message. Input pictures with indices idx in the range of 0 to numCandInputPics - 1 can precede the candidate input picture with index i-1 in the output order. candInputPicList[0] can represent a list of candidate input pictures listed in reverse order of the output order. nnpfga_target_id indicates the target NNPFG specified by the NNPFGC SEI message related to the current picture, and may have the same nnpfgc_id as nnpfga_target_id. The nnpfga_target_id value is between 0 and 2. 32- Must be within the range of 2. An NNPFGA SEI message with a particular value of nnpfga_target_id cannot be included in the current PU unless an NNPFGC SEI message with an nnpfgc_id equal to that value and an nnpfgc_grouping_type equal to 0 is present in the current PU or in a PU preceding the current PU in decoding order. If a PU contains both an NNPFGC SEI message with a particular nnpfgc_id value and an NNPFGA SEI message with an nnpfga_target_id equal to that nnpfgc_id value, the NNPFGC SEI message must precede the NNPFGA SEI message in decoding order. A value of 1 for nnpfga_cancel_flag may indicate that the persistence of the target NNPFG established by a previous NNPFGA SEI message with the same nnpfga_target_id as the current SEI message has been canceled. That is, the target NNPFG may no longer be used unless it is activated by another NNPFGA SEI message with the same nnpfga_target_id and nnpfga_cancel_flag set to 0. A value of 0 for nnpfga_cancel_flag may indicate that the target NNPFG is activated for use. nnpfga_persistence_flag can indicate the persistence of the target NNPFG for the current layer. A value of 0 for nnpfga_persistence_flag can indicate that the target NNPFG can be used for post-processing filtering only for the current picture. A value of 1 for nnpfga_persistence_flag can indicate that the target NNPFG can be used for post-processing filtering for the current picture and all subsequent pictures in output order of the current layer until one of the following conditions is true: - A new CLVS for the current layer is started. - End of bitstream - A picture of the current layer that contains an NNPFGA SEI message with the same nnpfga_target_id appears in the output order after the current picture. The target NNPFG does not apply to subsequent pictures in the current layer associated with an NNPFGA SEI message having the same nnpfga_target_id as the current SEI message. nnpfgcTargetPictures may be the set of pictures corresponding to the last NNPFGC SEI message with the same nnpfgc_id as the nnpfga_target_id that precedes the current NNPFGA SEI message in decoding order. nnpfgaTargetPictures may be the set of pictures for which the target NNPFG is activated by the current NNPFGA SEI message. All pictures included in nnpfgaTargetPictures must also be included in nnpfgcTargetPictures. The value of nnpfga_num_filters_minus2 plus 2 can represent the number of NNPFs in the NNPFG that the NNPFGA SEI message activates. The value of nnpfga_num_filters_minus2 must be equal to the value of nnpfgc_num_members_minus2 in the NNPFGC SEI message where nnpfgc_id is equal to nnpfga_target_id. A value of 1 for nnpfga_target_base_flag[i] may indicate that the ith NNPF in the target NNPFGC is the base NNPF with nnpfc_id equal to nnpfgc_member_id[i] in the NNPFGC SEI message whose nnpfgc_id is equal to nnpfga_target_id. A value of 0 for nnpfga_target_base_flag[i] may indicate that the ith NNPF in the target NNPFGC is the NNPF specified by the last NNPFC SEI message with nnpfc_id equal to nnpfgc_member_id[i] in the NNPFGC SEI message whose nnpfgc_id is equal to nnpfga_target_id. This message precedes the first VCL NAL unit of the current picture in decoding order and may not be a repeat of the NNPFC SEI message containing the base NNPF. A value of 1 in nnpfga_input_all_pics_flag[i] may indicate that the input pictures for the ith NNPF are selected from the list of candidate input pictures candInputPicList[i] without skipping. A value of 0 in nnpfga_input_all_pics_flag[i] may indicate that the input pictures for the ith NNPF are selected from candInputPicList[i] by skipping some candidate input pictures. nnpfga_num_input_pics_minus1[i] may represent the number of input pictures for the ith NNPF of the target NNPFG. If nnpfga_num_input_pics_minus1[i] exists, nnpfga_num_input_pics_minus1[i] must be equal to nnpfc_num_input_pics_minus1 for the NNPF with nnpfc_id equal to nnpfgc_member_id[i] in the NNPFGC SEI message whose nnpfgc_id is equal to nnpfga_target_id. If nnpfga_num_input_pics_minus1[i] does not exist, nnpfga_num_input_pics_minus1[i] can be inferred to be equal to nnpfc_num_input_pics_minus1 of the NNPF equal to nnpfgc_member_id[i] in the NNPFGC SEI message where nnpfgc_id is equal to nnpfga_target_id. nnpfga_input_pic_skip_count[i][j] may represent the j-th picture count to be skipped from the list of candidate input pictures candInputPicList[i] when selecting an input picture for the NNPF activated by the ith iteration loop entry. If nnpfga_input_pic_skip_count[i][j] does not exist, nnpfga_input_pic_skip_count[i][j] may be inferred to be 0 for all j values ​​when j is in the range of 0 to nnpfga_num_input_pics_minus1[i]. The variable numCandInputPics representing the number of candidate input pictures for the NNPFG may be derived as shown in Table 23. candInputPicList[m] (where m is in the range 1 to nnpfga_num_filters_minus2 + 1) is a list of pictures in reverse order of output, initially empty, and formed in descending order for n in the range 0 to m - 1. This list can be constructed by first adding pictures output by the NNPF process of the nth loop entry that are not already included in candInputPicList[m], and then lastly adding pictures in candInputPicList[0] that are not already included in candInputPicList[m]. For any value of m in the range of 1 to nnpfga_num_filters_minus2 + 1, if the candidate input picture candInputPicList[m][idx] is the NNPF output picture of the nth NNPF process with a value less than m, the width and height of the candidate input picture may be equal to nnpfcOutputPicWidth and nnpfcOutputPicHeight of the NNPF output picture, respectively. The list of input pictures for the NNPF of the mth loop entry, inputPicList[m], can be derived as shown in Table 24. candIdx shall not exceed the number of pictures in candInputPicList[m]. For any value of m in the range 1 to nnpfga_num_filters_minus2 + 1, the pictures in inputPicList[m] shall have the same width, height, bit depth, and chroma format. For the interpretation of an NNPFC SEI message whose nnpfc_id is equal to nnpfgc_member_id[i] in an NNPFGC SEI message whose nnpfgc_id is equal to nnpfga_target_id, the following variables may be specified for the i-th loop entry. - The variables BitDepthY, BitDepthC, and ChromaFormatIdc can be used as provided for interpreting the NNPFGA SEI message. - CroppedWidth and CroppedHeight are set to the width and height of the picture in inputPicList[i], respectively, and can be expressed in luma sample units. - For each input picture k in the range 0 to nnpfga_num_input_pics_minus1[i], the following can be applied: (1) CroppedYPic[k], CroppedCbPic[k], and CroppedCrPic[k] (if present) can be set to the sample array of inputPicList[i][k], respectively. (2) If nnpfc_auxiliary_inp_idc is equal to 1 for an NNPF where nnpfc_id is equal to nnpfgc_member_id[i] in an NNPFGC SEI message (where nnpfgc_id is equal to nnpfga_target_id), the following may apply. (2.1) inputPicList[i][k] must be equal to candInputPicList[0][idx] for idx values ​​in the range 0 to numCandInputPics - 1. (2.2) StrengthControlVal[k] can be set equal to InitStrengthControlVal[idx]. nnpfga_num_output_entries[i] may indicate the number of nnpfga_output_flag[i][j] syntax elements present in the NNPFGA SEI message. The value of nnpfga_num_output_entries[i] must be in the range 0 to NumInpPicsInOutputTensor for NNPF where nnpfc_id is equal to nnpfgc_member_id[i] in the NNPFGC SEI message where nnpfgc_id is equal to nnpfga_target_id. A value of 1 in nnpfga_output_flag[i][j] may indicate that the NNPF-generated picture corresponding to the input picture having the derived index InpIdx[j] for the i-th NNPF of the target NNPFG is output by the NNPF process activated by this loop entry, where the NNPF process may be specified in the semantics of the NNPFC SEI message. A value of 0 in nnpfga_output_flag[i][j] may indicate that the NNPF-generated picture corresponding to the input picture having the derived index InpIdx[j] for the i-th NNPF of the target NNPFG is not output by the NNPF process activated by this loop entry. If nnpfga_num_output_entries[i] is less than NumInpPicsInOutputTensor derived for the i-th NNPF of the target NNPFG, nnpfga_output_flag[i][j] can be inferred to be 1 for each value of i in the range of nnpfga_num_output_entries[i] to NumInpPicsInOutputTensor - 1. NnpfgaOutputPicList, a list of pictures output in output order by the NNPF process of NNPFG, is initially empty, and is formed in descending order for n in the range of 0 to nnpfga_num_filters_minus2 + 1, and can be constructed by adding each picture output by the NNPF process of the nth loop entry that is not already included in NnpfgaOutputPicList. Post-filter hint The syntax structure for post-filter hints is shown in Table 25. The post-filter hint syntax structure of Table 25 can be signaled in the form of an SEI message. An SEI message signaling the post-filter hint syntax structure of Table 25 may be referred to as a post-filter hint SEI message. The post-filter hint SEI message can provide post-filter coefficients or correlation information for the design of a post-filter, potentially allowing the decoded and output picture set to be used in post-processing to achieve improved display quality. A value of 1 for filter_hint_cancel_flag may indicate that the SEI message cancels the persistence of the previous post-filter hint SEI message in the output order applied to the current layer. A value of 0 for filter_hint_cancel_flag may indicate that post-filter hint information follows. filter_hint_persistence_flag can indicate the persistence of the post-filter hint SEI message for the current layer. A value of 0 for filter_hint_persistence_flag can indicate that the post-filter hint applies only to the currently decoded picture. A value of 1 for filter_hint_persistence_flag can indicate that the post-filter hint SEI message applies to the currently decoded picture and persists for all subsequent pictures in the current layer in output order until one or more of the following conditions are true: - A new CLVS for the current layer is started. - Bitstream ended - Pictures in the current layer of the AU associated with the post-filter hint SEI message are output after the current picture in the output order. filter_hint_size_y can indicate the vertical size of the filter coefficients or correlation array. The value of filter_hint_size_y must be in the range of 1 to 15. filter_hint_size_x can indicate the horizontal size of the filter coefficients or correlation array. The value of filter_hint_size_x must be in the range of 1 to 15. filter_hint_type can indicate the type of filter hint transmitted, as shown in Table 26. The value of filter_hint_type must be in the range of 0 to 2. A filter_hint_type value equal to 3 is reserved for future use and is not present in the bitstream. Decoders must ignore post-filter hint SEI messages with filter_hint_type equal to 3. A value of 1 for filter_hint_chroma_coeff_present_flag may indicate that filter coefficients for chroma exist. A value of 0 for filter_hint_chroma_coeff_present_flag may indicate that filter coefficients for chroma do not exist. filter_hint_value[ cIdx ][ cy ][ cx ] can represent filter coefficients or cross-correlation matrix elements between the original signal and the decoded signal with 16-bit precision. The value of filter_hint_value[ cIdx ][ cy ][ cx ] is -2 31 + 1 to 2 31 - Must be in the range of 1. cIdx represents the related color element, cy represents the vertical counter, and cx can represent the horizontal counter. Depending on the value of filter_hint_type, the following can be applied. - If the value of filter_hint_type is 0, the coefficients of a two-dimensional FIR (Finite Impulse Response) filter of the size of filter_hint_size_y * filter_hint_size_x can be transmitted. - Otherwise, if the value of filter_hint_type is 1, the filter coefficients of two one-dimensional FIR filters can be transmitted. In this case, the value of filter_hint_size_y must be 2. An index cy of 0 can represent the filter coefficients of the horizontal filter, and a cy of 1 can represent the filter coefficients of the vertical filter. In the filtering process, the horizontal filter is applied first, and the result can be filtered by the vertical filter. - Otherwise (if the value of filter_hint_type is 2), the transmitted hint may represent a cross-correlation matrix between the original signal s and the decoded signal s'. The normalized cross-correlation matrix for the relevant color components identified by cIdx of the size of filter_hint_size_y * filter_hint_size_x can be defined as in Equation 10. In Equation 10, s represents a sample array of the color component cIdx of the original picture, s' represents an array of the corresponding decoded picture, h represents the vertical height of the relevant color component, w represents the horizontal width of the relevant color component, and bitDepth represents the bit depth of the color component. In addition, OffsetY is equal to ( filter_hint_size_y >> 1 ), OffsetX is equal to ( filter_hint_size_x >> 1 ), and the range of cy is 0 <= cy < filter_hint_size_y, and the range of cx is 0 <= cx < filter_hint_size_x. The decoder can derive a Wiener post-filter from the cross-correlation matrix of the original signal and the decoded signal and the auto-cross-correlation matrix of the decoded signal. Source Picture Timing Information (SPTI) The syntax structure for SPTI is shown in Table 27. The SPTI syntax structure of Table 27 can be signaled in the form of an SEI message. An SEI message signaling the SPTI structure of Table 27 may be referred to as an SPTI SEI message. The SPTI SEI message can indicate the temporal distance between source pictures relative to the decoded output picture before encoding. For example, for content captured by a camera, the temporal distance between source pictures can mean the difference between the time the image sensor was exposed to generate the source picture relative to the currently decoded picture and the time the image sensor was exposed to generate the source picture relative to the previously decoded picture in the output order. A value of 1 for spti_cancel_flag may indicate that the SPTI SEI message cancels the persistence of a previous SPTI SEI message applied to the current layer. A value of 0 for spti_cancel_flag may indicate that source picture timing information follows. spti_persistence_flag can indicate the persistence of SPTI SEI messages for the current layer. A value of 0 for spti_persistence_flag can indicate that the SPTI SEI message applies only to the currently decoded picture. A value of 1 for spti_persistence_flag can indicate that the SPTI SEI message applies to the currently decoded picture and persists to all subsequent pictures of the current layer in output order until one of the following conditions is true: - When a new CLVS of the current layer starts - When the bitstream ends - When a picture of an AU containing a SPTI SEI message located after the current picture in the output order is output As shown in Table 28, spti_source_picture_timing_type may indicate a timing relationship between a source picture and a decoded output picture. (spti_source_picture_timing_type & bitMask) being non-zero may indicate that the timing relationship has an interpretation associated with the bitMask value. If spti_source_picture_timing_type is greater than 0 and (spti_source_picture_timing_type & bitMask) is 0, the interpretation of the bitMask value may not be applied to the SPTI SEI message. If spti_source_picture_timing_type is 0, the timing relationship may be specified by the application. The spti_source_picture_timing_type value must be in the range 0 to 127. Values ​​128 to 255 for spti_source_picture_timing_type are reserved for future use and must not be included in the bitstream. Decoders must ignore SPTI SEI messages with spti_source_picture_timing_type in the range 128 to 255. The variable temporalReversalFlag can be equal to ( spti_source_picture_timing_type & 0x10 )? 1 : 0. A value of 1 for spti_source_timing_equals_output_timing_flag can indicate that the timing of the source picture is equal to the timing of the decoded output picture. A value of 0 for spti_source_timing_equals_output_timing_flag can indicate that the timing of the source picture may not be equal to the timing of the decoded output picture. If spti_source_timing_equals_output_timing_flag is 1 and a picture timing SEI message exists for the current picture, the source picture timing can be determined from the information carried in the picture timing SEI message. spti_time_scale can represent the number of time units that pass in one second. The value of spti_time_scale cannot be 0. For example, a time coordinate system using a 27MHz clock could have a spti_time_scale of 27,000,000. spti_num_units_in_elemental_source_picture_interval can represent the number of time units of a clock operating at a frequency of spti_time_scale Hz, which corresponds to the basic source picture interval of consecutive pictures in output order in CLVS. The specified basic source picture interval, i.e., the interval represented by the variable ElementalSourcePictureInterval, is expressed in seconds, which can be equal to spti_num_units_in_elemental_source_picture_interval divided by spti_time_scale. For example, if the basic source picture interval is 0.04 seconds, spti_time_scale can be set to 27,000,000 and spti_num_units_in_elemental_source_picture_interval can be set to 1,080,000. spti_max_sublayers_minus_1+1 can represent the maximum number of temporal sublayers that can exist in CLVS. spti_sublayer_source_picture_interval_scale_factor[i], if present, may indicate a scale factor used to determine the source picture interval of consecutive pictures in output order in CLVS whose TemporalId is less than or equal to i. A value of 0 for spti_sublayer_source_picture_interval_scale_factor[i] may be used to indicate that the source picture corresponding to the currently decoded output picture is the same as the source picture corresponding to the previously decoded output picture. The source picture interval related to output pictures whose TemporalId is less than or equal to i is represented by the variable SourcePictureInterval[i] and can be derived in seconds as in mathematical expression 11. If picture n is an output picture with TemporalId less than or equal to I and is not the first picture in the bitstream in output order, the value of the variable SourcePictureTime[ n ] can be derived as follows. - If temporalReversalFlag is 0, SourcePictureTime[n] = SourcePictureTime[previousPicInOutputOrder] + SourcePictureInterval[i] - Otherwise, if temporalReversalFlag is 1, SourcePictureTime[ n ] = SourcePictureTime[ previousPicInOutputOrder ] - SourcePictureInterval[ i ] Here, previousPicInOutputOrder can indicate the last output picture that precedes picture n in the output order, with TemporalId less than or equal to i (if such a picture exists). If the SourcePictureTime[0] value is not provided by external means, the SourcePictureTime[0] value can be inferred to be 0. If spti_sublayer_synthesized_picture_flag[i] exists, a value of 1 in spti_sublayer_synthesized_picture_flag[i] may indicate that the decoded output pictures belonging to the ith temporal sublayer are synthesized and do not correspond to the unmodified original source pictures. A value of 0 in spti_sublayer_synthesized_picture_flag[i] does not provide such an indication. If spti_sublayer_synthesized_picture_flag[i] does not exist, the value of spti_sublayer_synthesized_picture_flag[i] may be inferred to be 0. Problems with prior art A picture that provides an association with a source picture by the semantics of the SPTI SEI message (e.g., picture 0, or an anchor picture, or a first source timing picture, or a standard picture, but hereinafter collectively referred to as picture 0) may need to be a picture that satisfies a specific condition (e.g., a first output picture). In this case, it is unclear in the current design which picture will be the first output picture, i.e., which picture will be referred to as picture 0. In other words, due to the absence of a precise definition of picture 0, picture 0 may be classified as more than one picture. For example, picture 0 may be classified as the first picture in the bitstream, but it may also mean the first picture in CVS / CLVS, and it may also be classified as the first picture associated with an SEI message (e.g., in terms of the SEI message). This makes picture 0 somewhat ambiguous in meaning. Additionally, the current design includes a feature indicating whether the picture associated with the SPTI SEI message is a synthesized picture or not. Since a synthesized picture is not associated with a source picture, i.e., it does not have an associated source picture, it can be restricted that picture 0 should not be associated with a synthesized picture. However, the current design does not have any restrictions regarding this, which can cause problems. Overview of the Example The present disclosure proposes various embodiments that can address the problems of conventional designs, including those described above. The following embodiments can be used independently, in combination with each other or with other embodiments, and can also be combined with the embodiments described above, which are also included within the present disclosure. 1. Picture 0 for the calculation of a specific syntax of the SPTI SEI message (for example, syntax related to source picture timing may be included, e.g., sourcepictureTime[0], etc.) may be specified as a picture within the same access unit (AU) containing the SEI message. 2. Since there is no corresponding source picture for a composite picture, it can be restricted that picture 0 should not be a composite picture. 3. The temporal identifier (TemporalId) of picture 0 can be restricted to a specific value (e.g., 0). Below, various embodiments, including the above-described embodiments, are described in more detail, and improvements to pictures in SEI messages related to neural-network post-filter (NNPF) of coded video bitstreams are presented. Although the embodiments described below are based on standard video codecs (e.g., Versatile Video Coding (VVC), etc.) and Versatile Supplemental Enhancement Information Messages for Coded Video Bitstreams (VSEI), it is self-evident that they can be applied to other video coding techniques, and thus, these are also included in the scope of the present disclosure. Meanwhile, in describing the embodiments below, the NNPF (Neural-network post-filter) SEI message or the NNPF-related SEI message may include an NNPFC (Neural-network post-filter Characteristic) SEI message and / or an NNPFA (Neural-network post-filter activation) SEI message. Meanwhile, the names of the syntaxes used in explaining the examples below are arbitrarily designated for clarity of explanation, so it is obvious that the names of the syntaxes may be changed, and even if the names of the syntaxes are changed, they will be included in the present disclosure. Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. Example The following examples detail the embodiments of Tables 1, 2, and 3 described in the above-described embodiment overview. The following describes the VSEI message syntax and semantics. As an example, the syntax and semantics of SEI messages related to Neural-network post-filter (NNPF) (e.g., NNPFA SEI messages, etc.) may be partially modified. As an example, assume that picture 0 is a picture of the same access unit as the SPTI SEI message, and for N greater than 0, picture N is the Nth picture from picture 0. As an example, if the TEMPORALID of picture N is less than or equal to any value i and is not the first picture in the bitstream, information related to the source picture timing (e.g., SourcePictureTime[ n ]) can be derived as follows: - If the value of specific information (e.g., temporalReversalFlag, etc.) is a specific value (e.g., 0), information related to source picture timing SourcePictureTime[ n ] can be derived based on SourcePictureTime[ previousPicInOutputOrder ] and information about the source picture interval. For example, SourcePictureTime[ n ] can be derived based on the addition operation of SourcePictureTime[ previousPicInOutputOrder ] and SourcePictureInterval[ i ] (SourcePictureTime[ previousPicInOutputOrder ] + SourcePictureInterval[ i ]). - If the value of specific information (e.g., temporalReversalFlag, etc.) is a specific value (e.g., 1), information related to source picture timing, SourcePictureTime[ n ], can be derived based on SourcePictureTime[ previousPicInOutputOrder ] and information about the source picture interval. For example, SourcePictureTime[ n ] can be derived based on the subtraction operation of SourcePictureTime[ previousPicInOutputOrder ] and SourcePictureInterval[ i ] (SourcePictureTime[ previousPicInOutputOrder ] - SourcePictureInterval[ i ]). Meanwhile, as an example, previousPicInOutputOrder may be the last output picture that precedes picture n in output order and has a temporal identifier (TemporalId) less than or equal to a specific threshold value (e.g., denoted as i). Meanwhile, if the value of information related to source picture timing (SourcePictureTime[0]) is not provided based on other syntax or other information, the value of SourcePictureTime[0] may be inferred to be a specific value (e.g., 0). In addition, as an example, the value of the temporal identifier of picture 0 (e.g., an anchor picture, a reference picture, the first source timing picture, etc.) may be restricted to a specific value (e.g., 0). Meanwhile, as an example, if information spti_sublayer_synthesized_picture_flag[ i ] related to a synthesized picture exists, if the value of the information is 1, it may indicate that the decoded output picture belonging to the i-th temporal sublayer is a synthesized picture that does not correspond to the unmodified original source picture. Conversely, if the value of the information is 0, the above restriction may not be indicated. In addition, if the syntax does not exist, the value of the syntax may be inferred as a specific value (e.g., 0). Meanwhile, picture 0 (e.g., reference picture, anchor picture, first source timing picture, first source timing picture, etc.) may not be a picture belonging to the synthesized temporal layer. According to the above-described embodiment described in the present disclosure, not only can the problems of the prior art described above be solved, but also the error of the decoder can be reduced and the coding quality and efficiency can be improved by diversifying the variables, syntax and / or semantics of the information, clarifying constraints on the information, or clarifying the semantics of the information. Embodiments of video decoding method and encoding method Hereinafter, image encoding methods and image decoding methods according to various embodiments of the present disclosure will be described. The image decoding method of FIG. 5 can be performed by an image decoding device (200), and the image encoding method of FIG. 6 can be performed by an image encoding device (100). In addition, the image decoding and encoding methods of FIGS. 5 and 6 can be based on the embodiments described above, as well as the examples described below. FIG. 5 is a diagram for explaining an image decoding method that can be performed by an image decoding device according to one embodiment of the present disclosure. First, a source picture timing information (SPTI) SEI (Supplemental Enhancement Information) message can be acquired (S510). As an example, the SPTI SEI message can be signaled in a bitstream, and an image can be restored (S520) based on the acquired SPTI SEI message. Meanwhile, as an example, the SPTI SEI message may include information about the source picture timing for a picture associated with the SEI. The source picture timing may be derived based on an anchor picture. As an example, the anchor picture may be represented as picture 0, the standard picture, the first source timing picture, or the first source timing picture, and the value of the temporal identifier of the anchor picture may be a specific value (e.g., 0). Also, as an example, the anchor picture may not belong to a synthesized temporal layer, which may be a constraint. Also, the anchor picture may be included in the same access unit as the access unit including the SPTI SEI message. Meanwhile, as an example, the source picture timing may be derived for a specific picture based on an order relative to the anchor picture, and the order may include, for example, an output order relative to the anchor picture. Since other explanations are the same as those described above, duplicate explanations are omitted. Meanwhile, since the image decoding method of FIG. 5 corresponds to one embodiment of the present disclosure, certain steps may be changed, the order of steps may be changed, or some steps may be added or deleted, and it will be obvious that even if there are such modifications, they are included in the present disclosure. FIG. 6 is a diagram for explaining an image encoding method that can be performed by an image encoding device according to one embodiment of the present disclosure. The image encoding method of FIG. 6 can also be performed based on the embodiment described above as well as the description below, and even though the description is based on a decoder, it can also be applied to the image encoding method as long as the description does not conflict with the operation of the encoder. As an example, a source picture timing information (SPTI) SEI (Supplemental Enhancement Information) message may be determined (S610). The step of determining the SPTI SEI message may include a step of determining information to be included in the SPTI SEI message. Thereafter, a bitstream including the SPTI SEI message may be encoded (S620). In other words, the SPTI SEI message may be encoded and the SPTI SEI message may be signaled. Meanwhile, as described above, as an example, the SPTI SEI message may include information about the source picture timing for a picture associated with the SEI. The source picture timing may be determined based on an anchor picture. As an example, the anchor picture may be represented as picture 0, the standard picture, the first source timing picture, or the first source timing picture, and the value of the temporal identifier of the anchor picture may be a specific value (e.g., 0). In addition, as an example, the anchor picture may not belong to a synthesized temporal layer, which may be a constraint. In addition, the anchor picture may be included in the same access unit as the access unit including the SPTI SEI message. Meanwhile, as an example, the source picture timing may be determined for a specific picture based on an order relative to the anchor picture, and the order may include, as an example, an output order relative to the anchor picture. Since other explanations are the same as those described above, duplicate explanations are omitted. Afterwards, the picture may be restored based on the output picture information, although it is not expressed in the drawing. Additionally, as an example, a computer-readable medium recording a bitstream generated by an image encoding method may be provided, and a method for transmitting a bitstream generated by an image encoding method may be provided. Meanwhile, since the image encoding method of FIG. 6 corresponds to one embodiment of the present disclosure, certain steps may be changed, the order of steps may be changed, or some steps may be added or deleted, and it will be obvious that even if there are such modifications, they are included in the present disclosure. According to the present invention, the meaning of information that can be included in VSEI can be clarified, reducing decoder errors and enabling more accurate representation of scenarios, thereby improving coding quality. Furthermore, the present invention enables clear handling of cases where a corresponding output picture has been generated but an input picture does not exist, thereby improving coding efficiency. FIG. 7 is a diagram illustrating an example of a content streaming system to which an embodiment according to the present disclosure can be applied. As illustrated in FIG. 7, a content streaming system to which an embodiment of the present disclosure is applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device. The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, generates a bitstream, and transmits it to the streaming server. Alternatively, if multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server may be omitted. The above bitstream can be generated by an image encoding method and / or an image encoding device to which an embodiment of the present disclosure is applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream. The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server can act as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, and the streaming server can transmit multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server may control commands / responses between each device within the content streaming system. The streaming server can receive content from a media repository and / or encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time. Examples of the user devices may include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs), digital TVs, desktop computers, digital signage, etc. Each server within the above content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner. The scope of the present disclosure includes software or machine-executable instructions (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of various embodiments to be executed on a device or a computer, and a non-transitory computer-readable medium having such software or instructions stored thereon and executable on the device or computer. Embodiments according to the present disclosure can be used to encode / decode images.

Claims

1. In the video decryption method, A step of obtaining a SPTI (source picture timing information) SEI (Supplemental Enhancement Information) message; and A step of restoring an image based on the above SPTI SEI message; including: A method for decoding a video, wherein the SPTI SEI message includes information about source picture timing for a picture associated with the SEI.

2. In paragraph 1, A video decoding method, wherein the above source picture timing is derived based on an anchor picture.

3. In paragraph 2, A video decoding method wherein the value of the temporal identifier of the above anchor picture is 0.

4. In paragraph 2, A method for decoding an image, wherein the above anchor picture is restricted from belonging to a synthesized temporal layer.

5. In paragraph 2, A video decoding method, wherein the anchor picture is included in the same access unit as the access unit including the SPTI SEI message.

6. In paragraph 2, A video decoding method wherein the above source picture timing is derived for a specific picture based on an order based on the above anchor picture.

7. In the video encoding method, A step of determining a SPTI (source picture timing information) SEI (Supplemental Enhancement Information) message; and A step of encoding a bitstream including the SPTI SEI message; comprising: A video encoding method, wherein the SPTI SEI message includes information about source picture timing for a picture associated with the SEI.

8. In paragraph 7, A video encoding method wherein the above source picture timing is derived based on an anchor picture.

9. In paragraph 8, A video encoding method wherein the value of the temporal identifier of the above anchor picture is 0.

10. A non-transitory computer-readable medium containing a bitstream generated by an image encoding method, wherein the image encoding method comprises: A step of determining a SPTI (source picture timing information) SEI (Supplemental Enhancement Information) message; and A step of encoding a bitstream including the SPTI SEI message; comprising: The above SPTI SEI message contains information about the source picture timing for the picture associated with the SEI.

11. A method for transmitting a bitstream generated by a video encoding method, A step of transmitting the above non-stream; including, The above image encoding method is, A step of determining a SPTI (source picture timing information) SEI (Supplemental Enhancement Information) message; and A step of encoding a bitstream including the SPTI SEI message; comprising: A method wherein the SPTI SEI message includes information about source picture timing for a picture associated with the SEI.