A method for encoding / decoding video, a method for transmitting a bitstream, and a recording medium for storing a bitstream.

The video encoding/decoding method clarifies neural-network post-filter information in SEI messages to enhance encoding/decoding efficiency, reduce decoder errors, and optimize video storage and transmission.

JP2026136387APending Publication Date: 2026-08-25LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026095333
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-04-06
Filing Date
2026-06-08
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

The increasing demand for high-resolution, high-quality video has led to a surge in data volume, resulting in higher transmission and storage costs, necessitating highly efficient video compression technologies.

Method used

A video encoding/decoding method that includes processing SEI messages to clarify the meaning of neural-network post-filter information, adjusting the signaling order of this information, and using non-temporary computer-readable recording media to store and transmit bitstreams generated by this method.

Benefits of technology

This approach improves encoding/decoding efficiency, reduces decoder errors, and enhances coding quality by clearly identifying and signaling neural-network post-filter information, thereby optimizing video storage and transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026136387000001_ABST
    Figure 2026136387000001_ABST
Patent Text Reader

Abstract

The present invention provides a video encoding / decoding method, a bitstream transmission method, and a computer-readable recording medium for storing bitstreams. [Solution] The video decoding method includes the steps of: obtaining post-filter-based corresponding output picture information for an input picture from an NNPFC (neural-network post-filter characteristics) SEI (supplemental enhancement information) message; and obtaining a corresponding output picture for the input picture based on the corresponding output picture information. The corresponding output picture information includes output picture generation information indicating whether or not a post-filter-based corresponding output picture has been generated for the input picture.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to a video encoding / decoding method, a method for transmitting a bitstream, and a bitstream Regarding recording media that store streams, video encoding related to neural network post-processing filters / Decoding method, bitstream transmission method, and method for storing bitstream Regarding recording media. [Background technology]

[0002] In recent years, HD (High Definition) video and UHD (Ultra High) have become popular. There is a growing demand for high-resolution, high-quality video, such as (h Definition) video. It is increasing in this field. The higher the resolution and quality of the video data, the more it compares to existing video data. The amount of information or bits transmitted will increase. An increase in data volume leads to increased transmission and storage costs.

[0003] Therefore, in order to effectively transmit, store, and play back high-resolution, high-quality video information Highly efficient video compression technology is desired. [Overview of the project] [Problems that the invention aims to solve]

[0004] This disclosure provides a video encoding / decoding method and apparatus with improved encoding / decoding efficiency. The purpose is to achieve this.

[0005] Furthermore, this disclosure also includes NNPF-related SEI messages (NNPFC SEI and NNPFA The objective is to provide a method for processing SEI.

[0006] In addition, the present disclosure aims to more clearly identify the output picture of the NNPF by NNPF-related SEI messages. More specifically.

[0007] In addition, the present disclosure aims to clarify the meaning of information related to the output picture of the NNPF. Purpose.

[0008] In addition, the present disclosure aims to reduce the decoder error by clarifying the meaning of information related to the output picture of the NNPF. By doing so.

[0009] In addition, the present disclosure aims to improve coding quality and efficiency by clarifying the meaning of information related to the output picture of the NNPF. By doing so.

[0010] In addition, the present disclosure aims to improve efficiency by adjusting the signaling order of information related to the output picture of the NNPF. To do so.

[0011] In addition, the present disclosure aims to provide a non-temporary computer-readable recording medium for storing a bitstream generated by the video encoding method according to the present disclosure. Purpose.

[0012] In addition, the present disclosure aims to provide a non-temporary computer-readable recording medium for storing a bitstream that is received and decoded by the video decoding apparatus according to the present disclosure and is used for video restoration. Purpose. To do so.

[0013] In addition, the present disclosure aims to provide a method for transmitting a bitstream generated by the video encoding method according to the present disclosure. Purpose.

[0014] The technical problems to be achieved in the present disclosure are not limited to the technical problems mentioned above, and mention Another technical problem not addressed will be clearly understood by those having ordinary knowledge in the technical field to which the present disclosure pertains from the following description.

Means for Solving the Problem

[0015] The video decoding method performed by a video decoder according to an embodiment of the present disclosure includes obtaining post-filter-based corresponding output picture information for an input picture from an NNPFC (neural-network post-filter characteristics) SEI (supplemental enhancement information) message, and obtaining a corresponding output picture for the input picture based on the corresponding output picture information, where the corresponding output picture information may include output picture generation information regarding the presence or absence of generation of a corresponding output picture of a post-filter for the input picture (may include; may be provided with; may be configured with; may be constructed with; may be set with; may include by inclusion; may include; may contain; may have).

[0016]

[0017]

[0018] On the other hand, according to an embodiment of the present disclosure, the value of the output picture generation information may be restricted to a specific value based on specific conditions.

[0019] On the other hand, according to an embodiment of the present disclosure, the specific conditions may be associated with the purpose of the post-filter.

[0020] On the other hand, according to an embodiment of the present disclosure, the specific conditions may be further associated with whether the purpose of the post-filter is pixel rate upsampling.

[0021] On the other hand, according to one embodiment of the present disclosure, the specific condition is the picture of the input picture. It may be further associated with an index.

[0020] On the other hand, according to one embodiment of the present disclosure, based on the aforementioned specific conditions, at least within the specified range The value of the output picture generation information for one input picture is limited to 1. stomach.

[0021] On the other hand, according to one embodiment of the present disclosure, the specific range is associated with the number of input pictures. The number of input pictures and the information associated with them may be signaled.

[0022] On the other hand, according to one embodiment of the present disclosure, the picture index of the input picture If it is 0, the value of the output picture generation information for the input picture is limited to 1. It is permissible.

[0023] On the other hand, according to one embodiment of the present disclosure, the specific condition is associated with the number of input pictures. It's okay to be kicked.

[0024] On the other hand, according to one embodiment of the present disclosure, the corresponding output picture for the input picture Based on the existence of the above, the input picture is in the bitstream the corresponding It may be replaced with a response picture.

[0025] On the other hand, according to one embodiment of the present disclosure, the output picture generation information is the postfill Regardless of whether the purpose of the signal includes picture rate upsampling or not, it is signaled. good.

[0026] The video encoding method performed by the video encoding device according to one embodiment of the present disclosure is an input pic The step of determining post-filter-based corresponding output picture information for the char,

[0027] The corresponding output picture information is NNPFC (neural-network post -filter characteristics)SEI(supplemental The stage where the message is signaled (enhancement information) , including, the corresponding output picture information is the post-filter for the input picture. The output picture generation information may include whether or not an output picture was generated in response.

[0028] Furthermore, according to this disclosure, the bitstream generated by the video encoding method relating to this disclosure A non-temporary, computer-readable recording medium may be provided for storing the data.

[0029] Furthermore, according to this disclosure, the video received and decoded by the video decoding device relating to this disclosure is video A non-temporary computer-readable recording medium that stores the bitstream used to restore the image It may be provided.

[0030] Furthermore, according to this disclosure, the bitstream generated by the video encoding method is transmitted A method may be provided.

[0031] The features of this disclosure briefly summarized above are illustrative of the characteristics described in the detailed description of this disclosure below. The mere fact that it is described as such does not limit the scope of this disclosure. [Effects of the Invention]

[0032] This disclosure provides a video encoding / decoding method and apparatus with improved encoding / decoding efficiency. It can be provided.

[0033] Furthermore, according to this disclosure, the semantics of the information within the NNPFC SEI message are corrected. Correcting it makes it possible to communicate meaning more clearly.

[0034] Furthermore, according to this disclosure, the semantics of the information within the NNPFC SEI message are corrected. By correcting this, the decoder's error can be reduced.

[0035] Furthermore, according to this disclosure, the information signaling order within NNPFC SEI messages is determined. By organizing things, efficiency can be improved.

[0036] Furthermore, according to this disclosure, NNPF-related SEI messages can trigger NNPF output pici By more clearly identifying information about the user, efficiency can be improved.

[0037] Furthermore, according to this disclosure, NNPF clarifies the meaning of information related to the output picture. This can improve coding quality and efficiency.

[0038] Furthermore, according to this disclosure, the bitstream generated by the video encoding method relating to this disclosure It is possible to provide a non-temporary computer-readable recording medium for storing data.

[0039] Furthermore, according to this disclosure, the video received and decoded by the video decoding device relating to this disclosure is video A non-temporary computer-readable recording medium for storing the bitstream used to restore the image. It can be provided.

[0040] Furthermore, according to this disclosure, the bitstream generated by the video encoding method is transmitted We can provide a method for doing so.

[0041] The effects obtained from this disclosure are not limited to those mentioned above, but include any other effects not mentioned. From the following statements, it is clear to a person with ordinary skill in the art to which this disclosure pertains. It will be understood. [Brief explanation of the drawing]

[0042] [Figure 1] This is a schematic diagram showing a video coding system to which the embodiments of this disclosure can be applied. [Figure 2] This is a schematic diagram showing a video encoding device to which the embodiments of this disclosure can be applied. [Figure 3] This is a schematic diagram showing an image decoding device to which the embodiments of this disclosure can be applied. [Figure 4] This diagram illustrates the interleaved method for lumen channel induction. [Figure 5] This is a diagram illustrating the output picture of an NNPF (Neural-network post filter). [Figure 6] This is a flowchart illustrating a video decoding method to which the embodiments of this disclosure can be applied. [Figure 7] This is a flowchart illustrating a video encoding method to which the embodiments of this disclosure can be applied. [Figure 8] This figure illustrates a content streaming system to which the embodiments of this disclosure can be applied. [Modes for carrying out the invention]

[0043] In the following, with reference to the attached drawings, embodiments of this disclosure will be described in relation to the part of the technology to which this disclosure belongs. This will be explained in detail so that it can be easily implemented by someone with ordinary knowledge in the field. The disclosure may be embodied in various other forms and is not limited to the embodiments described herein.

[0044] In describing embodiments of this disclosure, specific descriptions of known configurations or functions are omitted. If it is determined that disclosing the gist of the disclosure may obscure the facts, then detailed explanations regarding that will be omitted. In the figures, parts unrelated to the explanation relating to this disclosure have been omitted, and similar parts have been marked with similar reference numerals. Assign a number to it.

[0045] In this disclosure, one component is "connected," "joined," or "linked" with another component. If this is the case, then in addition to direct connections, there are other configurations in between them. This may include indirect connections where elements exist. Furthermore, it may also include cases where one element connects to other elements. When we say "includes" or "possesses," unless otherwise specified, this means that it also includes other components. It doesn't mean excluding something, but rather that it can potentially include other components.

[0046] In this disclosure, terms such as "first," "second," etc., are used to distinguish one component from other components. It is used solely for the purpose of [specific purpose], and unless otherwise specified, it does not limit the order or importance of the constituent elements. It is not. Therefore, within the scope of this disclosure, the first component in one embodiment is not the same as the other embodiment. In the examples, it may also be referred to as the second component, and similarly, the second component in one embodiment may also be In other embodiments, it may be referred to as the first component.

[0047] In this disclosure, components that are distinct from each other are used to clearly explain their respective characteristics. This does not necessarily mean that those components are always separate. A number of components are integrated to form a single hardware or software unit. Often, a single component is distributed and used as multiple hardware or software units. It may be configured in this way. Therefore, without further ado, integrated or distributed The embodiments described herein are also included in the scope of this disclosure.

[0048] In this disclosure, the components described in various embodiments are not necessarily essential components. This does not mean that some components are optional. Therefore, in one embodiment... Embodiments consisting of a subset of the components described below are also included in the scope of this disclosure. Embodiments that further include other components in addition to the components described in various embodiments are also within the scope of this disclosure. It is included.

[0049] This disclosure relates to the encoding and decoding of video, and the terms used in this disclosure are the same as those used in this disclosure. Unless otherwise defined in this disclosure, the terms have the ordinary meaning common to the field of the technology to which this disclosure pertains. That's fine.

[0050] In this disclosure, “picture” generally refers to one video of a specific time period. A slice / tile is a unit that represents an image. A coding unit that constitutes part of a picture; one picture is composed of one or more slices / tiles. It may be done. Also, a slice / tile can be one or more CTUs (coding tree It may include (unit).

[0051] In this disclosure, “pixel” or “pel” means one pic It can mean the smallest unit that makes up a picture (or image). It can also refer to the unit corresponding to a pixel. The word "sample" can be used. A sample generally refers to a pixel or This can represent the pixel value, and only the luma component pixel / pixel value is represented. It is also acceptable to represent only the pixel / pixel value of the chroma component.

[0052] In this disclosure, "unit" may represent a basic unit of video processing. A knit includes at least one of the following: a specific area of ​​a picture and information related to that area. That's fine. Depending on the situation, the unit may be called a "sample array" or a "block". This can be rephrased as "or area" or similar terms. In general, MxN blocks A sample (or sample array) consisting of M columns and N rows, or a transformation coefficient. It may include a set (or array) of (transform coefficients). .

[0053] In this disclosure, "current block" means "current coding block" or "current coding block". "Encoding unit", "Block to be encoded", "Block to be decoded", or "Processing target" It can mean one of the "blocks". When a prediction is made, "current block" means "current This can mean "predicted block" or "target block for prediction". Transformation (inverse transformation) / Quantization (inverse When quantization is performed, the "current block" becomes the "current transformed block" or the "transformed block". It can mean "blocked". When filtering is performed, "currently blocked" means "blocked". This can mean "blocks that are filtered."

[0054] In this disclosure, "current block" means "chroma block" unless otherwise explicitly stated. , a block containing both a luma component block and a chroma component block or "the current block It can mean "mablock". Currently, the luma component block of the block explicitly means "lumablock". Tables that include explicit descriptions of "luma component block," such as "luma block" or "current luma block." It may be displayed. Also, the chroma component block of the current block is explicitly referred to as "chroma block". Includes explicit mention of chroma component block, such as "current chroma block" or "current chroma block". It is acceptable to express it.

[0055] In this disclosure, " / " and "," may be interpreted as "and / or." For example, "A " / B" and "A, B" can be interpreted as "A and / or B". Also, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C."

[0056] In this disclosure, "or" may be interpreted as "and / or". For example, "A or B " means 1) only "A", 2) only "B", or 3) "A and B". It is possible. Or, in this disclosure, “or” means “additionally or as an alternative.” It can mean "only or alternatively."

[0057] Overview of the video coding system

[0058] Figure 1 is a schematic diagram showing a video coding system to which the embodiments of this disclosure can be applied. That is the case.

[0059] A video coding system according to one embodiment comprises an encoding device 10 and a decoding device 20. It may include. The encoding device 10 encodes video and / or image (i mage) information or data in file or streaming form on a digital storage medium or This information can be transmitted to the decoding device 20 via the network.

[0060] An encoding device 10 according to one embodiment includes a video source generation unit 11, an encoding unit 12, and a transmission unit. It may include part 13. The decoding device 20 according to one embodiment includes a receiving unit 21, a decoding unit 22, and It may include a rendering unit 23. The encoding unit 12 is called a video / image encoding unit. The decoding unit 22 may be called the video / image decoding unit. The transmitting unit 13 is The numbering unit 12 may be included. The receiving unit 21 may be included in the decoding unit 22. Rendering unit 23 may include a display unit, and the display unit may be a separate device or external device. It may be composed of components.

[0061] The video source generation unit 11 performs video / image capture, synthesis, or generation processes, etc. This allows you to acquire video / images. The video source generation unit 11 generates video / images May include a capture device and / or a video / image generation device. A capture device is, for example, one or more cameras, previously captured video / This may include video / video archives containing video footage. Examples of video / video generation devices include... For example, this may include computers, tablets and smartphones, etc. (electronically) It can generate audio / video. For example, a computer can generate a virtual video / image. This is often done, and in this case, the video / image capture is performed by the process of generating the related data. The process can be replaced.

[0062] The encoding unit 12 can encode the input video / image. The encoding unit 12 compresses and For improved coding efficiency, a series of procedures such as prediction, transformation, and quantization can be performed. The encoding unit 12 processes the encoded data (encoded video / image information) into a bitstream. It can output in bitstream format.

[0063] The transmission unit 13 outputs encoded video / image information in the form of a bitstream or It can acquire data and store it in file or streaming format on a digital storage medium or network. It can be transmitted via the network to the receiving unit 21 of the decoding device 20 or to other external objects. The storage media include USB, SD, CD, DVD, Blu-ray (registered trademark: same applies hereinafter), and HD. It may include various storage media such as D drives and SSDs. The transmission unit 13 transmits a predetermined file It may include elements for generating media files according to the format, broadcast / May include elements for transmission over a communication network. The transmitting unit 13 is code It may be provided as a transmission device separate from the processing unit 12, in which case the transmission device is a bitst To obtain encoded video / image information or data output in the form of a video, It also includes another processor and a transmission unit that transmits it in file or streaming format. That's fine. The receiving unit 21 extracts the bitstream from the storage medium or network. The data can be transmitted to the decoding unit 22 after being sent / received.

[0064] The decoding unit 22 performs a series of operations corresponding to the operation of the encoding unit 12, such as inverse quantization, inverse transform, and prediction. You can decrypt the video / image by following the procedure.

[0065] The rendering unit 23 can render the decoded video / image. The rendered video / image may be displayed through the display unit.

[0066] Overview of video encoding equipment

[0067] Figure 2 is a schematic diagram showing a video encoding device to which the embodiments of this disclosure can be applied.

[0068] As shown in Figure 2, the video encoding device 100 consists of a video splitting unit 110, a subtraction unit 115, and a conversion unit. Unit 120, quantization unit 130, inverse quantization unit 140, inverse transformation unit 150, addition unit 155, fill Taring unit 160, memory 170, inter prediction unit 180, intra prediction unit 185, and It may include an entropy encoding unit 190. Interpretation unit 180 and The prediction unit 185 may be collectively referred to as the "prediction unit". Transformation unit 120, quantization unit 130, inverse The quantization unit 140 and the inverse transformation unit 150 are included in the residual processing unit. The residual processing unit may further include a subtraction unit 115.

[0069] All or at least some of the multiple components constituting the video encoding device 100 are as follows: Therefore, as a single hardware component (for example, an encoder or processor) It may be implemented in this way. Also, memory 170 is DPB (decoded picture b It may include (uffer) and may be embodied by a digital storage medium.

[0070] The video splitting unit 110 receives the input video (or picture) input to the video encoding device 100. Divide the frame into one or more processing units. This is possible. For example, the processing unit is a coding unit (coding It may be called a unit (CU). A coding unit is a coding tree unit. Coding tree unit (CTU) or maximum coding unit (l QT / BT / TT(Quad-tr Recursively by the structure ( ee / binary-tree / ternary-tree) It can be obtained by (ecursively) splitting. For example, one coding The unit has a quad-tree structure, a binary tree structure, and / or a ternary tree structure. Based on the structure, it is divided into multiple coding units of lower (deeper) depth. Good. For the division of coding units, a quad-tree structure is first applied, and by A nuclear tree structure and / or ternary tree structure may be applied later. The coding procedures relating to this disclosure are performed based on the final coding unit which is not specified. It is fine. The largest coding unit may be used directly as the final coding unit. The coding units of the lower depth obtained by dividing the maximum coding unit are It may be used as the final coding unit. Here, the coding procedure is: This may include procedures such as prediction, transformation and / or restoration, as described later. Another example is the aforementioned code The processing unit for the ding procedure is the prediction unit (PU). t) or a transform unit (TU). The prediction The unit and the conversion unit are each separated from the final coding unit or Partitioning is permitted. The prediction unit may be a unit of sample prediction. The conversion unit generates a unit for which the conversion coefficient is derived, and / or a residual from the conversion coefficient. It may be a unit that induces a residual signal.

[0071] The prediction unit (inter-prediction unit 180 or intra-prediction unit 185) processes the target block (current A prediction is made for the current block, and the prediction includes a prediction sample for the current block. It can generate predicted blocks. The prediction unit currently generates blocks. The decision is made on whether intra-prediction or inter-prediction is applied on a per-unit or per-unit basis. The prediction unit generates various information about the prediction of the current block and calculates the entropy encoding. The information regarding the prediction can be transmitted to the reading unit 190. The data may be encoded in section 190 and output in the form of a bitstream.

[0072] The intra prediction unit 185 predicts the current block by referring to the samples in the current picture. The aforementioned referenced sample is in intra-prediction mode and / or intra The prediction method indicates that the current block may be located in the vicinity (neighbor), or They can be located at a distance. The intra-prediction mode has multiple non-directional modes and multiple directional modes. It may include a non-directional mode, for example, DC mode and planar mode (Pla It may include nar mode. The directional mode depends on the accuracy of the predicted direction, for example, It may include 33 direction prediction modes or 65 direction prediction modes. However, this is not limited to This is an example, and depending on the settings, more or fewer directional prediction modes may be used. Good. The intra prediction unit 185 uses the prediction mode applied to the surrounding block to determine the current block. You can also determine the prediction mode applied to the lock.

[0073] The interpretation unit 180 identifies a reference vector on a reference picture that is determined by the motion vector. Based on the lock (reference sample array), the predicted block for the current block is It can be guided. In this case, the amount of motion information transmitted in interprediction mode can be reduced. To do this, the movement information is broken down based on the correlation of movement information between surrounding blocks and the current block. The motion information can be predicted at the lock, subblock, or sample level. The motion information may include a vector and a reference picture index. Directional information (L0 prediction, L1 prediction, Bi prediction, etc.) may be included. Interpretation In this case, the surrounding blocks are the spatially surrounding blocks that currently exist within the picture. (l neighboring block) and the temporal peripheral block present in the reference picture It may include a lock (temporal neighboring block). A reference picture including a reference block and a reference picture including the temporal peripheral block are They may be identical or different to each other. The aforementioned temporal peripheral blocks are identical positional references. Block (collocated reference block), same-location CU ( It may be called colCU, etc. The reference picture including the aforementioned temporal peripheral block is the same It is called a col-located picture (colPic). Good. For example, the interpretation unit 180 generates a list of motion information candidates based on the surrounding blocks. Configure and derive the motion vector and / or reference picture index of the current block. It is possible to generate information that indicates which candidate should be used for the purpose of various predictions. Interpretation may be performed based on the mode, for example, skip mode and merge mode. In this case, the interpretation unit 180 uses the motion information of the surrounding blocks as the motion information of the current block. It can be used as follows. In skip mode, unlike merge mode, the residual signal is It does not need to be transmitted. Motion vector prediction In n,MVP) mode, the motion vector of the surrounding block is used as the motion vector predictor (mot It is used as an ion vector predictor and motion vector difference (mot Indicators for ion vector difference and motion vector predictors By encoding (indicator), the motion vector of the current block is signaled It is possible to do this. The motion vector difference is the motion vector of the current block and the motion vector This can represent the difference from the predictor.

[0074] The prediction unit generates a prediction signal based on various prediction methods and / or prediction techniques described later. This is possible. For example, the prediction unit can perform intraprediction or intraprediction for the current block prediction. You may apply ter prediction, or you may apply intra prediction and inter prediction simultaneously. A prediction method that applies intra-prediction and inter-prediction simultaneously for block prediction is CI IP (combined inter and intra prediction) You may be called. Also, the prediction unit currently uses an intrablock copy for predicting the block. Intrablock Co. (IBC) can also be performed. P stands for, for example, SCC (screen content coding). It can be used for coding video content such as games. IBC is, A reference block that has already been restored in the current picture, located at a predetermined distance from the current block. This is a method of predicting the current block using locks. When IBC is applied, the current block The position of a reference block within the block is determined by the vector (block vector) corresponding to the predetermined distance. It may be encoded as (Tor). IBC basically makes predictions within the current picture, but In terms of deriving reference blocks within the picture, this can be done similarly to interpretation. In other words, IBC uses at least one of the interpretation methods described in this disclosure. It is possible.

[0075] The predicted signal generated by the prediction unit is used to generate the restored signal, or the residual It may be used to generate a signal. The subtraction unit 115 subtracts the input video signal (original block). From the original sample array, the prediction signal (predicted block, predicted) is output from the prediction unit. Subtract the measured sample array to obtain the residual signal. It can generate residual blocks (residual sample array). The signal may be transmitted to the conversion unit 120.

[0076] The conversion unit 120 applies a conversion method to the residual signal and converts the conversion coefficient (transfor It is possible to generate m coefficients. For example, the transformation method is DCT. (Discrete Cosine Transform), DST(Discrete Sine Transform), KLT(Karhunen-Loeve Tran sform), GBT (Graph-Based Transform), or CNT ( Among the conditionally non-linear transforms, a small number It may include at least one. Here, GBT represents the relationship information between pixels as a graph. In this case, it means the transformation obtained from this graph. CNT is all the previously restored pins Xels (all previously reconstructed pixels) This refers to the transformation obtained by generating a prediction signal using this method. The transformation process is square It may also be applied to pixel blocks of the same size and shape, not just squares but variable sizes. It may be applied to the block of 's.

[0077] The quantization unit 130 quantizes the conversion coefficients and transmits them to the entropy encoding unit 190. It can be transmitted. The entropy encoding unit 190 transmits the quantized signal (quantity). The information regarding the childized conversion coefficients can be encoded and output as a bitstream. The information regarding the quantized transformation coefficients may be called residual information. The unit 130 determines the amount of block form based on the coefficient scan order. The converted transformation coefficients can be rearranged in a one-dimensional vector form, and the quantum in the one-dimensional vector form Based on the converted conversion coefficients, it is also possible to generate information about the quantized conversion coefficients. can.

[0078] The entropy encoding unit 190 is, for example, an exponentiation. al Golomb), CAVLC (context-adaptive variab le length coding), CABAC(context-adaptive Various encodings such as binary arithmetic coding The encoding method can be performed. The entropy encoding unit 190 is quantized. In addition to the conversion coefficients, other information necessary for video / image restoration (for example, syntax elements (s The values ​​of syntax elements (such as `<syntax>` elements) can also be encoded together or individually. Yes. Encoded information (for example, encoded video / image information) is bit In stream form, NAL (Network Abstraction Layer) The video / image information may be transmitted or stored in units of knits. Meter set (APS), Picture parameter set (PPS), Sequence parameter Various parameter sets such as Parameter Sets (SPS) or Video Parameter Sets (VPS) It may also include further information regarding the net. Furthermore, the video / image information may be subject to general restrictions ( (actual constraint information) may be included further. The signaling information, transmitted information, and / or syntax elements referred to in this disclosure. The bitstream is encoded by the encoding procedure described above. It may be included.

[0079] The bitstream may be transmitted over a network or on a digital storage medium. It may be stored in [location]. Here, the network may include broadcasting networks and / or communication networks, etc. Digital storage media include USB drives, SD cards, CDs, DVDs, Blu-ray discs, HDDs, SSDs, etc. This may include various storage media. Output from the entropy encoding unit 190 A transmitting unit (not shown) that transmits signals and / or a storage unit (not shown) that stores them, is a video code It may be provided as an internal / external element of the ionization device 100, or the transmitting unit may be an entropy e It may be provided as a component of the coding unit 190.

[0080] The quantized conversion coefficients output from the quantization unit 130 generate a resistive signal. It may be used for the following purposes. For example, the quantized conversion coefficients are inversely converted by the inverse quantization unit 140 and the inverse conversion. By applying inverse quantization and inverse transform in section 150, the residual signal (residual Blocks or registers (dual samples) can be restored.

[0081] The summing unit 155 processes the restored residual signal into the interprediction unit 180 or intra By adding it to the prediction signal output from the prediction unit 185, the reconstruction is performed. It is possible to generate signals (reconstructed picture, reconstructed block, reconstructed sample array). It can be done. The residual for the block to be processed, as if skip mode were applied. If none exists, the predicted block may be used as the restored block. Addition unit 155 This may be called the restoration unit or restoration block generation unit. The generated restoration signal is currently pict This may be used for intra-predicting the next block to be processed within the process, as described later. This may then be used for interpretation of the next picture after filtering.

[0082] The filtering unit 160 applies filtering to the restored signal to determine subjective / objective image quality. This can improve the results. For example, the filtering unit 160 can improve various aspects of the restored picture. A modified restored picture is generated by applying a filtering method. The corrected restored picture is then placed in memory 170, specifically in the DPB of memory 170. It can be saved. The various filtering methods mentioned above include, for example, deblocking filters. Retaring, sample adaptive offset t), adaptive loop filter, bidirectional It may include a filter (bilateral filter), etc. Filtering section 16 0 is related to filtering, as will be explained later in the description of each filtering method. It can generate various types of information and transmit them to the entropy encoding unit 190. Filtering information is encoded in the entropy encoding unit 190. It may be processed and output in the form of a bitstream.

[0083] The corrected restored picture transmitted to memory 170 is referenced by the interpretation unit 180. It may be used as a picture. This allows the video encoding device 100 to perform interpretation. When this is applied, the predicted mismatch in the video encoding device 100 and the video decoding device This can be avoided and coding efficiency can be improved.

[0084] The DPB in memory 170 is used as a reference picture by the interpretation unit 180. Therefore, the corrected restored picture can be saved. Memory 170 is currently pic Motion information within the block from which motion information was derived (or encoded) and / or It can save the movement information of blocks in the already restored picture. The motion information obtained is either the motion information of the spatially surrounding blocks or the motion information of the temporally surrounding blocks. It may be transmitted to the interpretation unit 180 for use. The memory 170 currently contains The restored sample of the restored block in the chart can be saved and transmitted to the intra prediction unit 185. It is possible.

[0085] Overview of the video decoding device

[0086] Figure 3 is a schematic diagram showing an image decoding device to which the embodiments of this disclosure can be applied.

[0087] As shown in Figure 3, the video decoding device 200 includes an entropy decoding unit 210, Inverse quantization unit 220, inverse transformation unit 230, addition unit 235, filtering unit 240, memory 2 50 may be configured to include an inter-prediction unit 260 and an intra-prediction unit 265. The ter prediction unit 260 and the intra prediction unit 265 can together be called the "prediction unit". The quantization unit 220 and the inverse conversion unit 230 may be included in the resistive processing unit.

[0088] All or at least some of the multiple components constituting the video decoding device 200 are as follows in the embodiment. Therefore, as a single hardware component (for example, a decoder or processor) It may be expressed. Also, memory 170 may include DPB and be expressed by a digital storage medium. It is acceptable to reveal it.

[0089] The video decoding device 200, which receives a bitstream containing video / image information, is shown in Figure 2. The process corresponding to the process performed by the video encoding device 100 is used to restore the video. This is possible. For example, the video decoding device 200 can process the processing units applied by the video encoding device. Decoding can be performed using the unit. Therefore, the decoding process unit The "set" may be, for example, a coding unit. A coding unit is a coding unit. It can be obtained as a coding tree unit, or by splitting the maximum coding unit. Then, the reconstructed video signal decoded and output by the video decoding device 200 is played back by the playback device (Figure (It can be played back without showing.)

[0090] The video decoding device 200 processes the signal output from the video encoding device shown in Figure 2 into a bitstream. It can be received in a blob form. The received signal is processed by the entropy decoding section. It may be decoded by 210. For example, the entropy decoding unit 210 is the B The stream is parsed to obtain the information necessary for video (or picture) restoration (for example). (video / image information) can be derived. The video / image information is adapted Image Parameter Set (APS), Picture Parameter Set (PPS), Sequence Various parameters such as Parameter Set (SPS) or Video Parameter Set (VPS) The video / image information may also include further information regarding the meter set. Further includes general constraint information. That's fine. The video decoding device uses the parameter set to decode the video. Further information and / or the aforementioned general restrictions may be available. The signaling information, received information and / or syntax elements are decoded by the decoding tool. The bitstream may be obtained by decoding it as follows. For example, the entropy decoding unit 210 uses exponential Golomb coding, CAVLC, or C Based on coding methods such as ABAC, the information in the bitstream is decoded. Quantization of the transformation coefficients related to the syntax element values ​​and resistivity required for image restoration. The resulting value can be output. More specifically, the CABAC entropy decoding method is: The bitstream receives a bin corresponding to each syntax element, and then decodes The syntax element information to be decoded, the surrounding blocks, and the decoded blocks. Using the ding information or the symbol / bin information decoded in a previous stage, the context (co Determine the ntext model and predict the probability of bin occurrence based on the determined context model. Perform arithmetic decoding of the bins, and each bin It is possible to generate a symbol corresponding to the value of the tax element. At this time, CABAC The tropy decoding method determines the context model of the next symbol / bin after the context model has been determined. For Dell, update the context model using decoded symbol / bin information. Yes, it is possible. Of the information decoded by the entropy decoding unit 210, the information relating to the prediction The information is provided to the prediction unit (inter-prediction unit 260 and intra-prediction unit 265), and the entro The resistive value obtained by entropy decoding in the P-decoding unit 210, In other words, the quantized conversion coefficients and related parameter information are input to the inverse quantization unit 220. It is acceptable. Also, among the information decoded by the entropy decoding unit 210, fill Information regarding taring may be provided to the filtering unit 240. Meanwhile, video coding A receiving unit (not shown) that receives the signal output from the device is located inside / outside the video decoding device 200. The receiving section may be further provided as an element, or the receiving section may be an entropy decoding It may be provided as a component of part 210.

[0091] On the other hand, the video decoding device relating to this disclosure is called a video / image decoding device. The aforementioned video decoding device is an information decoder (video / image / picture information decoder). ) and / or including a sample decoder (video / image / picture sample decoder) The information decoder may also include an entropy decoding unit 210. The simple decoder consists of an inverse quantization unit 220, an inverse transformation unit 230, an addition unit 235, and a filtering unit. Of the unit 240, memory 250, inter prediction unit 260, and intra prediction unit 265, at least It may include at least one.

[0092] The inverse quantization unit 220 inversely quantizes the quantized conversion coefficients and outputs the conversion coefficients. This is possible. The inverse quantization unit 220 rearranges the quantized transformation coefficients in a two-dimensional block form. This is possible. In this case, the realignment is performed in the order of the coefficient scan performed by the video encoding device. It may be carried out according to the order. The inverse quantization unit 220 sets the quantization parameters (for example, quantization s Inverse quantization is performed on the quantized transformation coefficients using step size information, and the transformation coefficients ( You can obtain the transform coefficients.

[0093] The inverse conversion unit 230 inversely converts the conversion coefficients to obtain a resistive signal (residual block, A residual sample array can be obtained.

[0094] The prediction unit makes a prediction for the current block and generates a prediction sample for the current block. It can generate a predicted block that includes [this]. The prediction part is Based on the information regarding the prediction output from the entropy decoding unit 210 The current block is then determined to determine whether intra-prediction or inter-prediction is applied. This allows for the determination of specific intra / interface prediction modes (prediction methods).

[0095] The ability of the prediction unit to generate prediction signals based on various prediction methods (techniques) described later is important for video. This is as mentioned in the description of the prediction unit of the coding device 100.

[0096] The intra prediction unit 265 predicts the current block by referring to the samples in the current picture. This is possible. The explanation regarding the intra prediction unit 185 is given in relation to the intra prediction unit 265. They may be applied identically.

[0097] The interpretation unit 260 predicts the reference vector identified by the motion vector on the reference picture. Based on the lock (reference sample array), the predicted block for the current block is It can be guided. In this case, the amount of motion information transmitted in interprediction mode can be reduced. To do this, the movement information is based on the correlation of movement information between surrounding blocks and the current block. The motion information can be predicted at the block, subblock, or sample level. The motion information may include vectors and reference picture indices. Directional information (L0 prediction, L1 prediction, Bi prediction, etc.) may be included. Interpretation field In addition, surrounding blocks are spatial surrounding blocks that currently exist within the picture. (neighboring block) and the temporal peripheral block present in the reference picture It may include a temporary neighboring block. For example The interpretation unit 260 then constructs a motion information candidate list based on the surrounding blocks and receives Based on the selected candidate information, the motion vector and / or reference picture of the current block are determined. An index can be derived. Interpretation based on various prediction modes (methods) Measurement may be performed, and the information regarding the prediction is the interpretation of the current block. It may include information indicating the mode (method).

[0098] The summing unit 235 adds the acquired residual signal to the prediction unit (interpretation unit 260 and / or including intra prediction unit 265.) Prediction signal output from (predicted block, By adding it to the predicted sample array, the reconstructed signal (reconstructed picture, reconstructed block, reconstructed) is obtained. It can generate the original sample array. If there is no residual for the lock, the predicted block is used as the restore block. It is permissible to leave it there. The explanation regarding the addition unit 155 also applies equally to the addition unit 235. Good. The addition unit 235 may be called the restoration unit or the restoration block generation unit. The generated restoration The signal is currently used for intra-prediction of the next block to be processed in the picture. Furthermore, as will be described later, it is filtered and used for predicting the next picture. You may do so.

[0099] The filtering unit 240 applies filtering to the restored signal to determine subjective / objective image quality. This can improve the results. For example, the filtering unit 240 can improve various aspects of the restored picture. Apply a filtering method to create a modified restored picture. The corrected restored picture can be stored in memory 250, specifically in the DPB of memory 250. It can be saved in. The various filtering methods mentioned above include, for example, deblocking Filtering, sample adaptive offsets et al., adaptive loop filter, bidirectional It may include filters (such as a bilateral filter).

[0100] The (corrected) restored picture stored in the DPB of memory 250 is in the interpretation unit. Memory 260 may be used as a reference picture. Memory 250 is currently used for moving images within the picture. The motion information of the block from which the information was derived (or decoded) and / or already restored The motion information of the blocks within the picture can be saved. The information will be used as motion information for spatially surrounding blocks or for temporally surrounding blocks. Therefore, it can be transmitted to the interpretation unit 260. The memory 250 currently contains the picture The restored sample of the restored block can be saved and transmitted to the intra prediction unit 265. It is possible.

[0101] In this specification, the filtering unit 160 and the interpretation unit of the video encoding device 100 are described. The embodiments described in 180 and the intra-prediction unit 185 respectively are video decoding device 20 The same applies to the filtering unit 240, the inter-prediction unit 260, and the intra-prediction unit 265 of 0. It may be applied in one or a corresponding manner.

[0102] Neural network post-filter characteristics (characteristics, NNPFC)

[0103] The combinations in Tables 1 to 3 represent the NNPFC syntax structure.

[0104] [Table 1]

[0105] [Table 2]

[0106] [Table 3]

[0107] The NNPFC syntax structures in Tables 1-3 are SEI (supplemental e (Insurance information) Signaled in the form of a message. i. The SEI messages that signal the NNPFC syntax structure in Tables 1 to 3 are NN This can be called a PFC SEI message.

[0108] NNPFC SEI messages are post-processing fill It is possible to identify neural networks that can be used as targets. Post-processing filters identified for specific pictures. The use of Luta is neural network post-filter activation. -filter activation,NNPFA) Shown using SEI messages This can be done. Here, "post-processing filter" and "post-filter" may have the same meaning. .

[0109] To use such SEI messages, you need to define the following variables: good.

[0110] - The width and height of the input picture may be cropped in lumens, and this width and Heights should be indicated by CroppedWidth and CroppedHeight, respectively. It is possible.

[0111] - CroppedYPic[idx] is a lumens sample array of the input picture. and the chroma sample array CroppedCbPic[idx] and Cropped dCrPic[idx] is used as input to NNPF if they exist. The index idx is in the range of 0 to numInputPics-1. You may do so.

[0112] - BitDepth Y This is the bitwise operation of the lumens sample array of the input picture. It can show a pusu.

[0113] - BitDepth C This is the chroma sample array of the input picture (if any) The bit depth can be shown.

[0114] - ChromaFormatIdc can indicate a chroma format identifier. ru.

[0115] - If the value of nnpfc_auxiliary_inp_idc is 1, fill The taring strength control value StrengthControlVal must be a real number in the range of 0 to 1. It must be done.

[0116] The input picture with index 0 is determined by the NNPFC SEI message. The NNPF that was activated by the NNPFA SEI message corresponds to the picture It is acceptable. The input where index i is within the range of 1 to numInputPics-1 The input picture takes precedence over the input picture with index i-1 in the output order. That's fine.

[0117] nnpfc_purpose & 0x08 is not the same as 0, and the index is 0. The input picture has the same fp_arrangement_type as 5. When the packing sequence is associated with an SEI message, all input pictures This is a frame packing array with the same fp_arrangement_type as 5. Often associated with SEI messages, fp_current_frame_is_f It may have the same value as rame0_flag.

[0118] There may be two or more NNPFC SEI messages for the same picture. Two or more NNPFC SEI messages with different nnpfc_id values If two or more NNPFC SEIs exist or are activated for one picture, The messages are either identical or different nnpfc_purpose and nnpf It may have a c_mode_idc value.

[0119] nnpfc_purpose can indicate the purpose of NNPF as shown in Table 4. The value of pfc_purpose is within the range of 0 to 63 in the bitstream. It may be restricted to the range of 64 to 65535 for nnpfc_purpose. The values ​​may be reserved for future use. The decoder is nn in the range of 64 to 65535. NNPFC SEI messages with pfc_purpose must be ignored. i. If the value of nnpfc_purpose is reserved for future use, this SE The syntax element of the I message is such that nnpfc_purpose is the same as the value in question. It may be extended to syntax elements that exist under the condition ChromaFormatId If c is the same as 3, then nnpfc_purpose & 0x02 is not the same as 0. It must be ChromaFormatIdc or nnpfc_purpose & When 0x02 is not the same as 0, nnpfc_purpose & 0x20 must be the same as 0. must be the same.

[0120]

Table 4

[0121] nnpfc_id may include an identification number that can be used to identify the NNPF. nn The pfc_id value must exist within the range of 0 to 2 32 -2. It must not exist outside the range of 256 to 511 and 2 range and 2 31 ~2 32 The nnpfc_id values in the range of -2 are reserved for future use. The decoder must ignore NNPFC SEI messages having nnpfc_id 31 in the range of 256 to 511 or 2 32 ~2 -2.

[0122] If the NNPFC SEI message is the first NNPFC SEI message in the current coding order having a specific nnpfc_id value within the current CLVS, then the following may apply. - The SEI message may indicate the basic (base) NNPF. may apply.

[0123] - The SEI message may be associated with all subsequent decoded pictures of the currently coded picture and the current layer until the current CLVS ends in the output order.

[0124] - The SEI message may be associated with all subsequent decoded pictures of the currently coded picture and the current layer until the current CLVS ends in the output order. associated with all subsequent decoded pictures of the currently coded picture and the current layer until the current CLVS ends in the output order. be associated.

[0125] The NNPFC SEI message is the previous one within the current CLVS in the decoding order. It may be a repetition of the NNPFC SEI message, and the subsequent semantics are this S The EI message is currently the only NNPFC SEI message with the same content within CLVS. It may be applied as if it were a gi.

[0126] A value of 0 for nnpfc_mode_idc indicates that the SEI message is a bit indicating basic NNPF. Includes a stream or is associated with a basic NNPF having the same nnpfc_id value. It can indicate an update.

[0127] NNPFC SEI messages currently have a specific nnnpfc_id value within CLVS. If it is the first NNPFC SEI message in the coding order, nnpf A value of 1 for c_mode_idc indicates that the basic NNPF associated with the nnpfc_id value is in the neural network. It can be shown that the neural network is tagged with the URI nnpfc_tag_uri Nerves identified by the URI displayed in nnpfc_uri using the identified format It can be a net.

[0128] NNPFC SEI messages currently have a specific nnpfc_id value within CLVS. In the decoding order, not the first NNPFC SEI message, but the first NNP If it is not an FC SEI message repetition, the value of nnpfc_mode_idc is 1. Updates related to the basic NNPF that have the same nnpfc_id value are tag URI Displayed in nnpfc_uri using the format identified by nnpfc_tag_uri It can be shown that it is defined by a URI.

[0129] The value of nnpfc_mode_idc is in the range of 0 to 1 in the bitstream. may be restricted as follows. Values in the range of 2 to 255 for nnpfc_mode_idc may be reserved for future use and may not be present in the bitstream. The decoder shall ignore NNPFC SEI messages having nnpfc_mode_idc in the range of 2 to 255. Values of nnpfc_mode_idc greater than 255 are not present in the bitstream and need not be reserved for future use.

[0130] When the SEI message is the first NNPFC SEI message in the decoding order having a specific nnpfc_id value within the current CLVS, NNPF PostProcessingFilter() may be assigned to be the same as the basic NNPF.

[0131] When the SEI message is not the first NNPFC SEI message in the decoding order having a specific nnpfc_id value within the current CLVS and is not a repetition of the first NNPFC SEI message, NNPF PostProcessingFilter() may be obtained by applying the update defined by the SEI message to the basic NNPF.

[0132] Updates are not accumulated. Rather, each update may be applied to the basic NNPF, which is the NNPF specified by the first NNPFC SEI message in the decoding order having a specific nnpfc_id value within the current CLVS.

[0133] nnpfc_reserved_zero_bit_a is subject to bitstream restrictions It may be restricted to have the same value as 0. The decoder is nnpfc_reserv NNPFC SEI messages with a non-zero value for ed_zero_bit_a will be ignored. It is acceptable to restrict it.

[0134] nnpfc_tag_uri is a neural network or nnpfc used as the basic NNPF. Upgrading to basic NNPF using the nnpfc_id value identified by _uri Syntax and semantics identified in IETF RFC 4151 for identifying dates It may include a tag URI with a .Using nnpfc_tag_uri, Even without a central registry, the format of neural network data specified by nnrpf_uri can be uniquely identified. It is the same as "tag:iso.org,2023:15938-17" nnpfc_ta g_uri is identified by nnpfc_uri, and the neural network data is ISO / IEC 1593 This demonstrates compliance with 8-17.

[0135] nnpfc_uri is the neural network used as the basic NNPF or the same nnpfc_i IETF Interns identify updates related to the basic NNPF that use d values. It has the syntax and semantics specified in et Standard 66. URIs may be included.

[0136] A value of 1 for nnpfc_property_present_flag indicates the purpose of the filter. Input formatting, output formatting, and syntax related to complexity It can be shown that a sub-element exists. nnpfc_property_pr A value of 0 for esent_flag indicates the purpose of the filter, input formatting, and output formatting. This demonstrates the absence of syntax elements related to matting and complexity. This is possible. The SEI message is the first NNPFC SEI in the decoding order. This is a message, and if you currently have a specific nnpfc_id value in CLVS, nnp The value of fc_property_present_flag is restricted to being the same as 1. It is acceptable. The value of nnpfc_property_present_flag is the same as 0. In some cases, the value of nnpfc_property_present_flag is 1. All syntax elements that exist only in this case and for which no inference value has been specified. The value is the NNPFC SEI message, which includes the basic NNPF that SEI provides updates for. It can be inferred that it is identical to the corresponding syntax element within the page.

[0137] A value of 1 for nnpfc_base_flag indicates that the SEI message is based on NNPF. This can be shown. A value of 0 for nnpfc_base_flag indicates that the SEI message This can show updates related to the basic NNPF. nnpfc_ba If se_flag does not exist, the value of nnpfc_base_flag is inferred to be 0. You may do so.

[0138] The following restrictions may apply to the value of nnpfc_base_flag:

[0139] - NNPFC SEI messages currently have a special feature in the CLVS in terms of decoding order. If it is the first NNPFC SEI message with a fixed nnpfc_id value, nn The value of pfc_base_flag can be the same as 1.

[0140] - NNPFC SEI message nnpfcB is currently C in the decoding order. Not the first NNPFC SEI message with a specific nnpfc_id value within LVS If the value of nnpfc_base_flag is the same as 1, NNPFC SEI The message is the first NNP that has the same nnpfc_id in the decoding order. This may correspond to an iteration of the FC SEI message nnpfcA. That is, nnpfcB The payload condensates may be the same as those of nnpfcA.

[0141] NNPFC SEI messages currently have a specific n in the CLVS in the decoding order. Not the first NNPFC SEI message with an npfc_id value, but a specific nmpfc If it does not correspond to the first NNPFC SEI message iteration with the _id value, the next The content may be applied.

[0142] - SEI messages have the same nnpfc_id value and are in the decoding order. This allows you to define updates related to the preceding basic NNPF.

[0143] - SEI messages in output order include the current restored picture of the current layer and all The subsequent restored picture is currently restored up to the end of CLVS or within CLVS. Only the picture up to the next restored picture is relevant, and in the current decoding order Within CLVS, the subsequent NNPFC SEI has an earlier value among the specified nnpfc_id values. Related to the message.

[0144] The NNPFC SEI message nnpfcCurr is currently in the decoding order. In the first NNPFC SEI message with a specific nnpfc_id value within CLVS Not in the first NNPFC SEI message iteration with a specific nnpfc_id value If not (i.e., if the value of nnpfc_base_flag is 0), nnpfc_ When the value of property_present_flag is 1, the following restrictions apply: It may be applied.

[0145] - The value of nnpfc_purpose in the NNPFC SEI message is decoded In the current order, the first NNPF in CLVS that has a specific nnpfc_id value The value of nnpfc_purpose in the C SEI message must be the same.

[0146] - nnpfc_base is a syntax element within NNPFC SEI messages. _flag and preceding nnpfc_complexity_info_p The value of resent_flag is a specific n in the current decoding order within CLVS. The corresponding syntax in the first NNPFC SEI message that has an npfc_id value It must be the same as the value of the element.

[0147] - In the decoding order, currently has a specific nnpfc_id value in CLVS nnpfc_complexity_info in the first NNPFC SEI message _present_flag must be identical to 0, or both must be identical to 1. The following may apply:

[0148] (1) nnpfc_parameter_parameter in nnpfcCurr _type_idc is nnpfc_parameter_pa in nnpfcBase It must be identical to rameter_type_idc.

[0149] (2) nnpfc_log2_parameter_bit_ in nnpfcCurr If length_minus3 exists, nnpfc_l in nnpfcCurr og2_parameter_bit_length_minus3 is nnpfcBa nnpfc_log2_parameter_bit_length_minu in se It must be s3 or lower.

[0150] (3) nnpfc_num_parameters_idc in nnpfcBase is If it is the same as 0, nnpfc_num_paramete in nnpfcCurr rs_idc must be equal to 0.

[0151] (4) Otherwise (nnpfc_num_parameter in nnpfcBase) (If ers_idc is greater than 0) nnpfcCurr contains nnpfc_num_p arameters_idc is greater than 0 or nnpfc in nnpfcBase It must be less than or equal to _num_parameters_idc.

[0152] (5) nnpfc_num_kmac_operations in nnpfcBase If _idc is the same as 0, then nnpfc_num_kma in nnpfcCurr c_operations_idc must be equal to 0.

[0153] (6) Otherwise (nnpfc_num_kmac_op in nnpfcBase) (If erations_idc is greater than 0) nnpfc Curr _num_kmac_operations_idc is greater than 0 and nnpfcB Unless it is under nnpfc_num_kmac_operations_idc within ase It must be done.

[0154] (7) nnpfc_total_kilobyte_size in nnpfcBase If it is the same as 0, then nnpfc_total_kilob in nnpfcCurr yte_size must be equal to 0.

[0155] (8) Otherwise (nnpfc_total_kilob in nnpfcBase) (If yte_size is greater than 0) nnpfc_tot in nnpfcCurr al_kilobyte_size is greater than 0 or nn in nnpfcBase It must be less than or equal to pfc_total_kilobyte_size.

[0156] nnpfc_out_sub_c_flag is nnpfc_purpose & 0 If x02 is not equal to 0, then variables outSubWidthC and outSubHei The value of ghtC can be shown. The value of nnpfc_out_sub_c_flag is 1. The value of outSubWidthC is 1, and the value of outSubHeightC is 1. It can be shown that. A value of 0 for nnpfc_out_sub_c_flag means ou The value of tSubWidthC is 2, and the value of outSubHeightC is 1. This can be shown that the value of ChromaFormatIdc is 2, and nnpfc_o If ut_sub_c_flag exists, nnpfc_out_sub_c_fl The value of ag must be equal to 1.

[0157] nnpfc_out_colour_format_idc is nnpfc_purp If ose & 0x20 is not the same as 0, the color format of the NNPFC output and The values ​​of the variables outSubWidthC and outSubHeightC are shown below. It is possible. A value of 1 for nnpfc_out_colour_format_idc is N The NPFC output color format is 4:2:0 format, and outSubWi It can be shown that both dthC and outSubHeightC are the same as 2. It works. The value of nnpfc_out_colour_format_idc is 2, NNPFC The output color format is 4:2:2 format, and outSubWidthC It can be shown that is 2 and outSubHeightC is 1. (nnpf) A value of 3 for c_out_colour_format_idc indicates the color format of the NNPFC output. The format is 4:2:4, and outSubWidthC and outSu It can be shown that bHeightC is 1 in all cases. nnpfc_out_c The value of olour_format_idc may be restricted to not be equal to 0.

[0158] nnpfc_purpose & 0x02 and nnpfc_purpose & 0 When x20 is the same as 0, outSubWidthC and outSubH Each of the eightC values ​​is identical to SubWidthC and SubHeightC. This can be inferred.

[0159] nnpfc_pic_width_in_luma_samples and nnpfc_ pic_height_in_luma_samples is the cropped and decoded This occurs as a result of applying NNPF, identified by nnpfc_id, to the output picture. The width and height of the lumens sample array of the picture can be shown, respectively. fc_pic_width_in_luma_samples and nnpfc_pic_ If height_in_luma_samples does not exist, each will be Cropped It can be inferred that this is the same as edWidth and CroppedHeight. nnpf The value of c_pic_width_in_luma_samples is CroppedWi It should be within the range of dth to CroppedWidth*16-1. nnpf The value of c_pic_height_in_luma_samples is CroppedH It should be within the range of eight to CroppedHeight*16-1.

[0160] nnpfc_num_input_pics_minus1+1 is the input to NNPF and This indicates the number of decoded output pictures used. nnpfc_n The value of um_input_pics_minus1 should be within the range of 0 to 63. It's okay to be restricted.

[0161] nnpfc_interpolated_pics[i] is used as input to NNPF The NNPF is generated between the i-th picture and the (i+1)th picture. The number of interpolated pictures created can be shown. nnpfc_interpolate The value of d_pics[i] may be limited to the range of 0 to 63. nnpfc_inte The value of rpolated_pics[i] is between 0 and nnpfc_num_input_pi Restrict to at least one i within the range cs_minus1-1 to be greater than 0. It is permissible.

[0162] The value 1 of nnpfc_input_pic_output_flag[i] is the i-th This demonstrates that NNPF generates a corresponding output picture for an input picture. This is possible. The value of nnpfc_input_pic_output_flag[i] is 0. NNPF should not generate a corresponding output picture for the i-th input picture. It can demonstrate this.

[0163] The variable numInputPics indicates the number of pictures used as input to NNPF, and numOutputPics indicates the total number of pictures generated as a result of NNPF. The variables may be derived as shown in Table 5.

[0164] [Table 5]

[0165] A value of 1 for nnpfc_component_last_flag is an input to NNPF. The last dimension of the force tensor inputTensor and the output tensor o, which is the result of NNPF. This can indicate that outputTensor is currently being used as the channel. (nnpf) A value of 0 for c_component_last_flag is the input tensor for NNPF. The third dimension of the input Tensor and the output tensor, which is the result of NNPF. This indicates that tTensor is currently being used for the channel.

[0166] The first dimension of the input tensor and output tensor is, in some neural network frameworks This SE may be used for the batch index. The formula within the semantics of an I message corresponds to a batch index like 0. The batch size is used, but the batch size used as input for neural network inference is later This may be determined by the implementation of the process.

[0167] For example, if the value of nnpfc_inp_order_idc is the same as 3, nnpfc When the value of _auxiliary_inp_idc is the same as 1, the input tensor has 4 The seven channels consist of 1 lumer matrix, 2 chroma matrices, and 1 auxiliary input matrix. It may exist. In this case, the DeriveInputTensors() process is input Each of the 7 channels of the tensor can be guided one by one, and among these channels, a specific channel When a channel is being processed, it may be referred to as the current channel during the process.

[0168] nnpfc_inp_format_idc is the truncated decoded output pic This document demonstrates how to convert sample values ​​from a chart into input values ​​for NNPF. If _inp_format_idc is 0, the input value to NNPF is a real number. The InpY() and InpC() functions may be specified as shown in Formula 1.

[0169]

number

[0170] If the value of nnpfc_inp_format_idc is 1, the input value of NNPF is unsigned integer numbers, and the InpY () and InpC() functions may be derived as shown in Table 6.

[0171]

Table 6

[0172] The variable inpTensorBitDepth Y may be derived from the syntax element nn pfc_inp_tensor_luma_bitdepth_minus8 described below. inpTensorBitDepth C may be derived from the syntax element nnpfc_inp_tensor_chroma_bitdepth_minus8 described below.

[0173] Values of nnpfc_inp_format_idc greater than 1 may be reserved for future use and need not be present in the bitstream. The decoder must ignore NNPFC SEI messages containing reserved values of nnpfc_i np_format_idc.

[0174] nnpfc_inp_tensor_luma_bitlength_minus8 + 8 can indicate the bit depth of the luma sample values in the input integer tensor. inpTensorBitDepth Y The value of may be derived as shown in Equation 2.

[0175]

Equation

[0176] nnpfc_inp_tensor_luma_bitlength_minus8 The value may be restricted to exist within the range of 0 to 24.

[0177] nnpfc_inp_tensor_chroma_bitdepth_minus8 +8 can indicate the bit depth of the chroma sample values ​​in the input integer tensor. inpTensorBitDepth C The value of can be derived as shown in equation 3.

[0178]

number

[0179] nnpfc_inp_tensor_chroma_bitdepth_minus8 The value of may be restricted to exist within the range of 0 to 24.

[0180] nnpfc_inp_order_idc is the truncated decoded output picture A method for aligning the sample array of a character to one of the input pictures for NNPF. It can be shown.

[0181] The value of nnpfc_inp_order_idc is between 0 and 3 in the bitstream. It must exist within the range. 4~ for nnpfc_inp_order_idc The value 255 does not exist in the bitstream. The decoder exists in the range of 4 to 255. Ignore NNPFC SEI messages with nnpfc_inp_order_idc. It must be done. Values ​​of nnpfc_inp_order_idc greater than 255 It does not exist in the bitstream and is not reserved for future use.

[0182] If the value of ChromaFormatIdc is not 1, nnpfc_inp_orde The value of r_idc must not be 3.

[0183] Table 7 contains explanations regarding the nnpfc_inp_order_idc value.

[0184] [Table 7]

[0185] The patch is the length of the sample from the picture components (e.g., lumens or chroma components). It may be a shape array.

[0186] If nnpfc_auxiliary_inp_idc is greater than 0, the NNPF input terminal This indicates that auxiliary input data exists in the nsol. nnpfc_auxili A value of 0 for ary_inp_idc indicates that auxiliary input data is not present in the input tensor. It is possible. The value of nnpfc_auxiliary_inp_idc is 1, auxiliary input The data can be shown in Tables 8 to 10 by the method of disclosure.

[0187] The value of nnpfc_auxiliary_inp_idc is in the bitstream. It must exist in the range of 0 to 1. The values ​​between 2 and 255 do not exist in the bitstream. The decoder is within the range of 2 to 255. NNPFC SEI messages with nnpfc_inp_order_idc enclosed This must be ignored. nnpfc_inp_order_idc greater than 255 The value does not exist in the bitstream and is not reserved for future use.

[0188] If the value of nnpfc_auxiliary_inp_idc is the same as 1, then the variable strengthControlScaledVal can be derived as shown in Equation 4. .

[0189]

number

[0190] Given vertical sample coordinates cTop and the left of the sample patch contained in the input tensor Input tensor for the horizontal sample coordinate cLeft, which specifies the upper end sample position, inpu The process DeriveInputTensors() for inducing tTensor is This can be shown as a combination of Tables 8 to 10.

[0191] [Table 8]

[0192] [Table 9]

[0193] [Table 10]

[0194] nnpfc_separate_colour_description_prese The value of nt_flag 1 is the base color (colour p) for the picture according to NNPF. The unique combinations of rimaries, transformation characteristics, and matrix coefficients are the SEI message syntax. This indicates that it will be specified as a S structure. nnfpc_separate_colo A value of 0 for ur_description_present_flag indicates that NNPF has a P The combination of the base color, transformation characteristics, and matrix coefficients of the texture is expressed in the VUI parameters of CLVS. It can be shown that it is identical to what has been shown.

[0195] nnpfc_colour_primaries are vui_colo except for the following. The same semantics defined for the ur_primaries syntax element. It may have a "S".

[0196] - nnpfc_colour_primaries are the primary colors used in CLVS. Rather, the picture that appears as a result of applying the NNPF specified in the SEI message. It can indicate the basic color.

[0197] - NNPFC SEI messages include nnpfc_colour_primaries If it does not exist, the value of nnpfc_colour_primaries will be vui_co It can be inferred that this is the same as the value of lour_primaries.

[0198] nnpfc_transfer_characteristics is, except for the following: For the vui_transfer_characteristics syntax element It may have the same semantics as defined.

[0199] - nnpfc_transfer_characteristics is set to CLVS The result of applying the NNPF specified to the SEI message, rather than the conversion characteristics used. This allows us to demonstrate the transformation characteristics of the resulting picture.

[0200] - NNPFC SEI message contains nnpfc_transfer_charact If eristics does not exist, nnpfc_transfer_character The value of istics is the value of vui_transfer_characteristics It can be inferred that this is the same thing.

[0201] nnpfc_matrix_coeffs is vui_matrix_ except for the following: The coeffs syntax element should have the same semantics as specified. stomach.

[0202] - nnpfc_matrix_coeffs are matrix coefficients used in CLVS. The line of the picture that appears as a result of applying the NNPF specified in the SEI message. Column coefficients can be shown.

[0203] - The NNPFC SEI message contains nnpfc_matrix_coeffs. Otherwise, the value of nnpfc_matrix_coeffs will be vui_matrix_c It can be inferred that this is the same as the oeffs value.

[0204] - The acceptable values ​​for nnpfc_matrix_coeffs are the VUI parameters Decoded and displayed as a ChromaFormatIdc value for semantics It does not need to be limited by the chroma format of the video picture.

[0205] - If the value of nnpfc_matrix_coeffs is the same as 0, nnpfc The value of _out_order_idc must not be the same as 1 or 3.

[0206] The value 0 of nnpfc_out_format_idc can indicate that for the subsequent post-processing or display the sample values output by NNPF are real numbers linearly mapped to the range of values from 0 to 1 within the range of unsigned integer values from 0 to (1 << bitDepth) - 1 for the bit depth bitDepth required for subsequent processing or display. The value 1 of nnpfc_ out_format_idc can indicate that the luma sample values output by NNPF are unsigned integers in the range of 0 to (1 << (nnpfc_out_tensor_luma_bitlength _minus8 + 8)) - 1, and the chroma sample values output by NNP F can be shown to be unsigned integers in the range of 0 to (1 << (nnpfc_out_tens or_chroma_bitlength_minus8 + 8)) - 1.

[0207] Values of nnpfc_out_format_idc greater than 1 may be reserved for future use and do not exist in the bitstream. The decoder must not ignore NNPFC SEI messages containing reserved values of nnpfc_out_ format_idc.

[0208] nnpfc_out_tensor_luma_bitdepth_minus8 + 8 can indicate the bit depth of the luma sample values in the output integer tensor. The value of n npfc_out_tensor_luma_bitdepth_minus8 must be in the range of 0 to 24.

[0209] nnpfc_out_tensor_chroma_bitdepth_minus8 +8 can indicate the bit depth of the chroma sample values ​​in the output integer tensor. . nnpfc_out_tensor_chroma_bitdepth_minus8 The value must be within the range of 0 to 24.

[0210] If nnpfc_purpose & 0x10 is not the same as 0, then nnpfc_o The value of ut_format_idc must be the same as 1, and at least one of the following restrictions At least one of them may be true.

[0211] - nnpfc_out_tensor_luma_bitdepth_minus8 +8 is BitDepth Y bigger

[0212] - nnpfc_out_tensor_chroma_bitdepth_minu s8+8 is BitDepth C bigger

[0213] nnpfc_out_order_idc is the output of the samples output from NNPF. The order can be indicated. The value of nnpfc_out_order_idc is the bitst The `nnpfc_out_order_i` must exist within the range of 0 to 3. The values ​​4 to 255 for DC do not exist in the bitstream. The decoder is 4 to 2 NNPFC S with nnpfc_out_order_idc in the range of 55 EI messages must be ignored. nnpfc_out_or greater than 255 The value of der_idc does not exist in the bitstream and is not reserved for future use. If the value of nnpfc_purpose & 0x02 is 0, then nnpfc_ou The value of t_order_idc must not be the same as 3.

[0214] Table 11 provides explanations for the values ​​of nnpfc_out_order_idc.

[0215] [Table 11]

[0216] Given vertical sample coordinates cTop and the patch of samples contained in the input tensor The output tensor ou corresponds to the horizontal sample coordinate cLeft, which indicates the upper left corner sample position. FilteredY is the output sample array filtered from tputTensor. Sample values ​​within Pic, FilteredCbPic, and FilteredCrPic The StoreOutputTensors() process for inducing this is shown in Table 12 and Table It can be expressed as 13 combinations.

[0217] [Table 12]

[0218] [Table 13]

[0219] nnpfc_overlap is an NNPF method for overwrapping adjacent input tensors. The number of horizontal and vertical samples can be shown (overlapping). nnpfc The value of _overlap must be within the range of 0 to 16383.

[0220] A value of 1 for nnpfc_constant_patch_size_flag is NNPF nnpfc_patch_width_minus1 and nnpfc_patch_h The patch size displayed by eight_minus1 is accepted exactly as input. This can be shown to be true. nnpfc_constant_patch_siz A value of 0 for e_flag means that NNPF has width inpPatchWidth and height inpPatch This indicates that it accepts any patch size with chHeight as input. This is possible. Here, inpPatchWidth+2*nnpfc_overlap The same extended patch width (i.e., the patch plus the overlapping area) ) is nnpfc_extended_patch_width_cd_delta_m It is a positive integer multiple of inus1+1+2*nnpfc_overlap, and inpPatc The height of the extended patch is the same as hHeight+2*nnpfc_overlap, n npfc_extended_patch_height_cd_delta_minu It is a positive integer multiple of s1 + 1 + 2 * nnpfc_overlap.

[0221] npfc_patch_width_minus1+1 is npfc_consta When the value of nt_patch_size_flag is 1, the number of patches required for NNPF input The horizontal sample size can be indicated. nnpfc_patch_width_m The value of inus1 is in the range of 0 to Min(32766,CroppedWidth-1). It must be there.

[0222] npfc_patch_height_minus1+1 is nnpfc_const When the value of ant_patch_size_flag is 1, the number of patches required for input to NNPF The vertical sample size can be shown using nnpfc_patch_height. The value of _minus1 is in the range of 0 to Min(32766,CroppedHeight-1). It must exist within the enclosure.

[0223] nnpfc_extended_patch_width_cd_delta_min us1+1+2*nnpfc_overlap is equal to nnpfc_constant_pa When the value of tch_size_flag is 0, the extension required for input to NNPF is It is possible to show the common divisor of the allowed values ​​of the patch width. To do. nnpfc_extended_patch_width_cd_delta_m The value of inus1 is in the range of 0 to Min(32766,CroppedWidth-1). It must exist.

[0224] nnpfc_extended_patch_height_cd_delta_mi nus1+1+2*nnpfc_overlap is equal to nnpfc_constant_p When the value of `atch_size_flag` is 0, the extensions required for input to NNPF are... To show the common divisor of the allowed values ​​of the patch heights. This is possible. nnpfc_extended_patch_height_cd_delt The value of a_minus1 is between 0 and Min(32766,CroppedHeight-1). It must exist within the range.

[0225] The inpPatchWidth and inpPatchHeight variables are, respectively, The width and height of the patch can be set accordingly.

[0226] If the value of nnpfc_constant_patch_size_flag is 0, The following may apply.

[0227] - The values ​​of inpPatchWidth and inpPatchHeight are determined by external means It may be provided by or set by the post-processor itself.

[0228] - The value of inpPatchWidth+2*nnpfc_overlap is nnpf c_extended_patch_width_cd_delta_minus1+1 It must be a positive integer multiple of +2*nnpfc_overlap, and inpPatchW idth must be less than or equal to CroppedWidth. The value of pPatchHeight+2* nnpfc_overlap is nnpfc_e xtended_patch_height_cd_delta_minus1+1+2 * must be a positive integer multiple of nnpfc_overlap, and inpPatchHei ght must be less than or equal to CroppedHeight.

[0229] Otherwise, (nnpfc_constant_patch_size_flag If the value is 1, then the value of inpPatchWidth is nnpfc_patch_wi It is set to be the same as dth_minus1+1, and the value of inpPatchHeight is nn It may be set to the same value as pfc_patch_height_minus1+1.

[0230] Variables outPatchWidth, outPatchHeight, horCScal ing, verCScaling, outPatchCWidth, and outPatc hCHeight may be derived as shown in Table 14.

[0231] [Table 14]

[0232] outPatchWidth*CroppedWidth is nnpfc_pic_wi dth_in_luma_samples*inpPatchWidth must be identical It must be, and outPatchHeight*CroppedHeight is nnpf c_pic_height_in_luma_samples*inpPatchHei The requirement that it be identical to ght is a requirement for bitstream conformance.

[0233] nnpfc_padding_type is the cut-off data as explained in Table 15. Padding is applied when referencing sample positions outside the boundaries of the coded output picture. This can show the process. The value of nnpfc_padding_type is 0 to 15. It must exist within the range.

[0234] [Table 15]

[0235] nnpfc_luma_padding_val is nnpfc_padding_t When the value of `ype` is 4, it can indicate the rumor value to be used for padding.

[0236] nnpfc_cb_padding_val is nnpfc_padding_typ When the value of e is 4, the Cb value to be used for padding can be indicated.

[0237] nnpfc_cr_padding_val is nnpfc_padding_typ When the value of e is 4, the Cr value to be used for padding can be indicated.

[0238] Input is vertical sample position y, horizontal sample position x, picture height picHeight In, the picture width (picWidth) and the sample array (CroppedPic) are specified. pSampleVal(y,x,picHeight,picWidth,Croppe The dPic) function can return the induced SampleVal value as shown in Table 16. ru.

[0239] For inputs to the InpSampleVal() function, some inference engines' input tests For compatibility with the Nsol rule, the vertical position may be placed before the horizontal position.

[0240] [Table 16]

[0241] The processes in Table 17 use NNPF PostProcessingFilter(). Then, the picture is filtered using a patch method, and the filtered and / or interpolated picture Often used to generate a filtered and / or interpolated picture. The y is as shown by nnpfc_out_order_idc, Y s Sample array FilteredYPic, Cb sample array FilteredCbPi c may include a Cr sample array FilteredCrPic.

[0242] [Table 17]

[0243] The order of the pictures in the saved output tensor may be the output order, and the output order may be N The output order generated by applying NPF will not conflict with the output order of the input picture. This can be analyzed as an introduction.

[0244] A value of 1 for nnpfc_complexity_info_present_flag is, There is one or more syntax elements that indicate the complexity of NNPF, associated with nnpfc_id. It can be shown that... A value of 0 for ent_flag indicates the complexity of NNPF as associated with nnpfc_id. This can indicate that no x elements exist.

[0245] A value of 0 for nnpfc_parameter_type_idc indicates that the neural network uses integer parameters. It can be indicated that only the 'ta' parameter is used. nnpfc_parameter_type A value of 1 for _flag indicates that the neural network can use floating-point or integer parameters. This is possible. A value of 2 for nnpfc_parameter_type_idc means that the neural network is two It can be shown that only linear parameters are used. nnpfc_parameter The value 3 for _type_idc may be reserved for future use and will be used in the bitstream. It does not exist. The decoder's value is nnpfc_parameter_type_idc The NNPFC SEI message, which is value 3, must be ignored.

[0246] nnpfc_log2_parameter_bit_length_minus3 Values ​​0, 1, 2, and 3 indicate that the neural network has more than 8, 16, 32, and 64 bits, respectively. It can be indicated that the length parameter will not be used. (nnpfc_parameter) _type_idc exists, and nnpfc_log2_parameter_bit_l If ength_minus3 does not exist, the neural network will not have parameters with a bit length greater than 1. You don't need to use "ta".

[0247] nnpfc_num_parameters_idc is the neural network parameter for NNPF. The maximum number of meters can be expressed in units of 2048. nnpfc_num_par A value of 0 for ameters_idc indicates that the maximum number of neural network parameters is unknown. This is possible. The value of nnpfc_num_parameters_idc is in the range of 0 to 52. It must exist in nnpfc_num_parameters greater than 52. The value of _idc does not exist in the bitstream. The decoder is greater than 52. NNPFC SEI messages with fc_num_parameters_idc It must be ignored.

[0248] If the value of nnpfc_num_parameters_idc is greater than 0, max The NumParameters variable can be derived as shown in equation 5.

[0249]

number

[0250] The number of neural network parameters in NNPF is less than maxNumParameters. The number may be restricted to equal numbers.

[0251] If nnpfc_num_kmac_operations_idc is greater than 0, then NN The maximum number of multiplication operations per sample in PF of multiply-accumulate operations per sa mple) is nnpfc_num_kmac_operations_idc * 100 It is possible to show that a number is less than or equal to 0. nnpfc_num_kmac_op A value of 0 for erations_idc indicates that the maximum number of network multiplication operations is unknown. This can be shown by nnpfc_num_kmac_operations_idc The value is 0-2 32 It must exist within the range of -2.

[0252] nnpfc_total_kilobyte_size greater than 0 indicates neural network compression. This shows the total size (in kilobytes) required to store parameters that have not been specified. This is possible. The total size in bits is the number of bits used to store each parameter. The sum can be a number greater than or equal to the sum. (nnpfc_total_kilobyte) _size may be the total size (in bits) divided by 8000 and rounded to the nearest integer. A value of 0 for nnpfc_total_kilobyte_size indicates that the parameter for the neural network... This indicates that the total size required to save the data is unknown. nnpfc The value of _total_kilobyte_size is between 0 and 2 32 If it does not exist within the range of -2 It must be done.

[0253] nnpfc_reserved_zero_bit_b is in the bitstream. It must be the same as 0. The decoder is nnpfc_reserved_zero_ NNPFC SEI messages where bit_b is not 0 must be ignored.

[0254] nnpfc_payload_byte[i] is the i-th byte of the bitstream. It may include the byte sequence nnpfc_paylo for all existing values ​​of i. ad_byte[i] is a complete bit number compliant with ISO / IEC 15938-17. It must be a trim.

[0255] Neural network post-filter activation (r activation, NNFPA)

[0256] Table 18 shows the syntax structure for NNFPA.

[0257] [Table 18]

[0258] The NNPFA syntax structure in Table 18 is signaled in the form of an SEI message. i. SEI messages that signal the NNPFA syntax structure in Table 18 are NNPF This can be called an SEI message.

[0259] The NNPFA SEI message is used for post-processing filtering of the picture set. The target neural network post-processing filter (NNPF) identified by nnpfa_target_id Possible uses can be activated or deactivated. NNPF is activated for specific pics. For the char, the target NNPF is the same as nnpfc_ nnpfa_target_id. An NNPF identified by the last NNPFC SEI message having an ID Good. Here, the last NNPFC SEI message is in the current decoding order. The first VCL NAL unit in the picture may precede the NN containing the basic NNPF. It does not need to be a repetition of the PFC SEI message.

[0260] When NNPF is used for another purpose or to filter out other color components, Multiple NNPFA SEI messages may exist for the same picture.

[0261] nnpfa_target_id is currently associated with nnfpa_targ in relation to the picture. One or more NNPFC SEI messages that have the same nnpfc_id as et_id Therefore, the specified NNPF can be shown.

[0262] The value of nnpfa_target_id is 0-2 32 It must not exist within the range of -2 No. The range is 256~511 and 2 31 ~2 32 nnpfa_targe within the range of -2 The t_id value may be reserved for future use. The decoder is 256~511 or 2 31 ~2 32 NNPFA SE with nnpfa_target_id within the range of -2 I must ignore "I" messages.

[0263] NNPFA SEI messages with a specific value for nnpfa_target_id are: Unless one or both of the following conditions are true, the PU must not currently exist.

[0264] - Currently, within CLVS, there are units that exist in the PU that precedes the current PU in the decoding order. NNPFCs that have the same nnpfc_id as the specific value of nnpfa_target_id SEI messages exist.

[0265] - Currently, the PU has a specific value for nnpfa_target_id that is the same as nnpfc_id There is an NNPFC SEI message that does this.

[0266] The PU has an NNPFC SEI message with a specific value nnpfc_id, and a specific value NNPFA SEI members have the same nnpfa_target_id as nnpfc_id. If all messages are included, the NNPFC SEI messages are in the decoding order. Therefore, it must precede the NNPFA SEI message.

[0267] A value of 1 for nnpfa_cancel_flag is currently the same as nnpf in SEI messages. Set by any previous NNPFA SEI message that has a_target_id It can be shown that the persistence of the target NNPF is canceled. That is, target N NPF currently uses the same nnpfa_target_id and 0 as SEI messages. Other NNPFA SEI messages containing the npfa_cancel_flag It will not be used further unless activated. (Value of nnpfa_cancel_flag) A value of 0 can indicate that nnpfa_persistence_flag will continue.

[0268] nnpfa_persistence_flag is the target NNPF for the current layer. This can demonstrate persistence. A value of 0 for nnpfa_persistence_flag is The target NNPF can currently only be used for post-processing filtering on pictures. This can be shown. A value of 1 for nnpfa_persistence_flag is the target N NPF outputs the current picture and current ray in order until one or more of the following conditions are true. This indicates that it can be used for post-processing filtering of all subsequent pictures within the file. It is possible.

[0269] - A new CLVS for the current layer is started.

[0270] - Bitstream ends

[0271] - Currently, the SEI message has the same nnpfa_target_id and the same nnp as 1. Current status related to NNPFA SEI messages with fa_cancel_flag The picture inside the ear will be output after the current picture in the output order.

[0272] The target NNPF currently uses the same nnpfa_target_id and 1 as the SEI message. NNPFA SEI messages with the same nnpfa_cancel_flag are related This does not apply to subsequent pictures within the current layer.

[0273] nnpfcTargetPictures is currently NNP in terms of decoding order. nnpf, which precedes the FA SEI message and has the same nnpf_target_id A set of pictures associated with the last NNPFC SEI message that has a c_id It's fine. nnpfaTargetPictures is currently NNPFA SEI Mess. This can be a set of pictures in which the target NNPF is activated by Sage. All any pictures included in aTargetPictures are nnpfcTa It should also be included in rgetPictures.

[0274] Post-filter hint

[0275] Table 19 shows the syntax structure for post-filter hints.

[0276] [Table 19]

[0277] The post-filter hint syntax structure in Table 19 is in the form of a SEI message signal. It may be used. The post-filter hint syntax structure in Table 19 signals the SEI mechanism. This message can be called a post-filter hint SEI message.

[0278] Post-filter hint SEI messages are used to improve display quality. The coding and output picture set can potentially be used for post-processing. This can provide post-filter coefficients or correlation information for the design of a post-filter.

[0279] The value of filter_hint_cancel_flag is 1, and the SEI message is currently The output order applied to the layer retains the previous post-filter hint SEI message. This indicates that the continuation can be canceled. (filter_hint_cancel_fl) A value of 0 for `ag` can indicate that post-filter hint information follows.

[0280] filter_hint_persistence_flag is for the current layer Post-filter hints can indicate the persistence of SEI messages. filter_h The value of int_persistence_flag is 0, and the post-filter hint is currently decorated. This indicates that the filter will only apply to the selected picture. A value of 1 for t_persistence_flag means that the post-filter hint SEI message The current decoded picture is applied if one or more of the following conditions are true. This indicates that the output order will persist for all subsequent pictures within the current layer. It is possible.

[0281] - A new CLVS for the current layer is started.

[0282] - Bitstream ends

[0283] - Post-filter hints SEI message associated with the picture in the AU's current layer - This will be output after the current picture in the output order.

[0284] filter_hint_size_y is the vertical size of the filter coefficient or correlation array. This can be shown. The value of filter_hint_size_y is in the range of 1 to 15. It must be there.

[0285] filter_hint_size_x is the horizontal size of the filter coefficient or correlation array. This can be shown. The value of filter_hint_size_x is in the range of 1 to 15. It must be there.

[0286] filter_hint_type is the transmitted filter hint as shown in Table 20. It is possible to indicate the type of [something]. The value of filter_hint_type is in the range of 0 to 2. It must exist. The same filter_hint_type value as 3 is bitst It does not exist in Ream. The decoder is a post with filter_hint_type 3. The filter hint SEI message must be ignored.

[0287] [Table 20]

[0288] value of filter_hint_chroma_coeff_present_flag 1 can be used to show that a filter coefficient exists for the chroma. A value of 0 for hint_chroma_coeff_present_flag means that the chroma is not present. It can be shown that no such filter coefficients exist.

[0289] filter_hint_value[cIdx][cy][cx] are the filter coefficients. Alternatively, the cross-correlation matrix elements between the original signal and the decoded signal can be shown with 16-bit precision. This is possible. The value of filter_hint_value[cIdx][cy][cx] is , -2 31 +1~2 31 It must be within the range of -1. cIdx is the associated color element This indicates that cy represents the vertical counter and cx represents the horizontal counter. This is possible. Depending on the value of filter_hint_type, the following may be applied.

[0290] - If the value of filter_hint_type is 0, filter_hint_ 2D FIR(Fini) of size_y*filter_hint_size_x The coefficients of the (Impulse Response) filter may be transmitted.

[0291] - On the other hand, if the value of filter_hint_type is 1, then two 1D FIR The filter coefficients of the filter may be transmitted. In this case, filter_hint_siz The value of e_y must be 2. An index cy that is 0 is a horizontal filter The cy coefficient, which is 1, represents the filter coefficient of the vertical filter. In the filtering process, a horizontal filter is applied first, and the result is then filtered by a vertical filter. Filtering is acceptable.

[0292] - Otherwise (if the value of filter_hint_type is 2), transmission The hint provided can show the cross-correlation matrix between the original signal s and the decoded signal s'. ru.

[0293] filter_hint_size_y*filter_hint_size_x size The normalized cross-correlation matrix for the related color components identified by cIdx is given by equation 6. It may be defined as follows.

[0294]

number

[0295] In equation 6, s represents the sample array of the color component cIdx of the original picture, and s' is, This represents the corresponding decoded picture array, where h is the vertical height of the associated color component. Here, w represents the horizontal width of the associated color component, and bitDepth represents the bit depth of the color component. This represents that. Also, OffsetY is (filter_hint_size_y>>1) and They are identical, and OffsetX is the same as (filter_hint_size_x>>1). The range of cy is 0 <= cy <filter_hint_size_yであり、 The range of cx is 0 <= cx <filter_hint_size_xである。

[0296] The decoder generates a cross-correlation matrix of the original signal and the decoded signal. (auto- Wiener post-filter from cross-correlation matrix It can induce this.

[0297] Problems with conventional technology

[0298] Currently, the design for input and output pictures in NNPFC SEI messages is as follows: If so, the following problems may arise.

[0299] 1. Output picture related information (for example, nnpfc_input_pic_outpu In relation to the semantics of t_flag[i])

[0300] NNPF (neural-network post-filter), that is, nerve The purpose of a mesh post-filter is picture rate upsampling. When including upsampling, any i-th and i+1-th input picture - Identify the number of interpolated pictures generated in between using nnpfc_interpolated In addition to signaling with the _pics[i] flag, NNPFC SEI messages also include i n is output picture generation information that determines whether or not the nth input picture will be output. npfc_input_pic_output_flag[i] may be included. On the other hand, Current semantics of nnpfc_input_pic_output_flag[i] It is unclear and may lead to one or more different analyses, but in particular, Phil The following problems may arise in relation to the bitstream after the tarning process. This will be explained with reference to Figure 5. Figure 5 is a diagram illustrating the NNPF output picture. That is the case.

[0301] (1) For example, considering Figure 5, if the NNPFC identifier, i.e., nnpfc_id is 0 Alternatively, there may be two NNPFC SEI messages that are both set to 1, and each is activated These two NNPFs are virtually the same filter, but the input or output picture There may be some differences in the signaling of the signal. Here, the (i-1)th input The first picture and the i-3rd input picture are nnpfc_input_pic_o Assuming that utput_flag[i] is associated with 0, then the following two Analysis is possible.

[0302] 1) The input picture is nnpfc_input_pic_output with a value of 0. If associated with _flag[i], this means that the picture is an NNPF file. The final bit remains uncorrected or filtered by the retaring process. This means it is part of the stream (i.e., the filtered bitstream). Yes, it is possible. As can be seen in Figure 5, according to this type of analysis, after the filtering process The bitstream has been corrected by the filtering process, with zero or more pictures The filter, and one or more new pictures generated by the filtering process ( In other words, it may be composed of (interpolated pictures).

[0303] 2) The input picture is nnpfc_input_pic_output with a value of 0. When associated with _flag[i], this is the NNPFC filtering process. Seth indicates that the picture is not output and is not part of the final bitstream. This is possible. In other words, the input picture is deleted after the filtering process. This can be shown. In Figure 5, according to this analysis, filtering process The bitstream after Seth is modified by the filtering process, one or more One or more pictures and filtering filters generated by the filtering process. It may consist of pictures (i.e., interpolated pictures).

[0304] What is explained in the table of contents 1) above is, relatively speaking, the true meaning of the semantics of the current syntax. It is considered close, but currently, in the text, it cannot be asserted that scenario 1) in the table of contents will always apply. It is not possible to do so, therefore, the output picture generation information is nnpfc_input_p It would be necessary to clarify the meaning of `ic_output_flag[i].`.

[0305] 2. In relation to the numerous output pictures unrelated to the purpose of NNPF

[0306] In NNPFC SEI messages, the output picture refers to the modified input picture. It can represent a filtered or selected version of the picture, and the final bit It may be used to substitute a specific input picture in the stream. If one or more input pictures exist, filtering will occur regardless of the purpose of NNPF. There may be one or more pictures.

[0307] On the other hand, when one or more input pictures are input to NNPF, the purpose of NNPF is to pic If it's not charrate upsampling, a large number of filtered pictures The scenarios we have may not be clearly supported by current technology. Even though the scenario is quite possible, it is restricted without a clear reason. A slightly more flexible alternative is needed.

[0308] 3. Output picture when the purpose of NNPF is not picture rate upsampling. - Related to

[0309] The output picture in the NNPFC SEI message is a modified / false version of the input picture. A filtered version of the picture, and a specific input within the final bitstream. Assuming it is intended to replace a picture, there is one or more input pictures. In some cases, there are one or more filtered pictures regardless of the purpose of NNPF. That's fine.

[0310] On the other hand, if the purpose of NNPF does not include picture rate upsampling, the output Signaling for the output picture despite the number of pictures being set to 1 This does not need to be done. With this design, NNPF can process one or more input pictures. The processing when this occurs can be unclear. In this case, the output picture will utilize NNPF. Picture within the same access unit for the transformed NNPFA SEI message This can correspond to the first input picture. However, without a clear explanation, It can be difficult to determine if this was intended, and problems may arise.

[0311] Summary of Examples

[0312] This disclosure proposes various embodiments that can solve the problems of conventional designs, including those mentioned above. The embodiments described may be used independently, or in combination with each other or with other embodiments. This may also be done, and these can be said to be included in this disclosure.

[0313] 1. The following description will be added to the output picture corresponding to the input picture. stomach:

[0314] (1) Input picture having a corresponding picture output from the filtering process This was replaced with the output picture in the bitstream after the filtering process. That's fine.

[0315] (2) Input picture for which there is no corresponding picture output from the NNPF filtering process The char may exist in the bitstream after the filtering process.

[0316] 2. In relation to output picture signaling, the following improvements may be proposed:

[0317] (1) Regardless of whether the purpose of NNPF includes picture rate upsampling or not The code may be modified so that a signal is always provided indicating whether or not the input picture will be output.

[0318] (2) If picture rate upsampling is not included in the purpose of NNPF, The constraint is that there must be at least one output picture corresponding to the power picture. You may add more items.

[0319] 3. As an alternative, if the purpose of NNPF does not include picture rate upsampling In addition, the output picture generation information is nnpfc_input_pic_output_ The value of flag[0] is inferred to be the first value (e.g., 1), and nnpfc_i nput_pic_output_flag[i](where i ranges from 1 to nnpfc_n) Values ​​within um_input_pics_minus1 are assumed to be the second value (for example, 0). It may be revised as discussed.

[0320] 4. As an alternative, the signaling of the output picture may include the following improvements:

[0321] (1) Regardless of whether the purpose of NNPF includes picture rate upsampling or not Whether the input picture is output or not (i.e., whether it is filtered or modified) It may be modified so that (this can be done) is signaled.

[0322] (2) When the purpose of NNPF does not include picture rate upsampling, The value of pfc_input_pic_output_flag[0] is the first value (for example, 1 ) must be equal to the The condition `first value(for example,1))` has been added. stomach.

[0323] 5. As an alternative, the following improvements may be made to the signaling of the output picture:

[0324] (1) Regardless of whether the purpose of NNPF includes picture rate upsampling or not Whether the input picture is output or not (i.e., whether it is filtered or modified) It may be modified so that (this can be done) is signaled.

[0325] (2) When the purpose of NNPF does not include picture rate upsampling, the output The picture generation information is nnpfc_input_pic_output_flag[ The value of [0] must be the first value (e.g., 1) regardless of the purpose of NNPF (sh could be equal to the first value(for exa The condition `mple, 1))` may be added.

[0326] Below, we will provide an example of the NNPF(neu) encoding of the video bitstream. ral-network post-filter) Input PIX in SEI messages Improvements to the char and output picture are presented. The embodiments described below are standard chars. Video codecs (for example, VVC (Versatile Video Coding)) (and VSEI (Versatile Supplemental Enhancement)) and VSEI (Versatile Supplemental Enhancement) ent Information Messages for Coded Video Created based on Bitstreams, but also suitable for other video coding techniques. It is self-evident that it may be used, and this can be said to be included in the content of this disclosure.

[0327] On the other hand, the syntax names used when describing the examples below are for clarity of explanation. This is arbitrarily designated for that purpose, and therefore the syntax name may be changed. This is self-evident, and it can be said that even if the syntax name is changed, it will still be included in this disclosure.

[0328] The embodiments of this application will be described in detail below with reference to the drawings.

[0329] Example 1

[0330] Example 1 will provide a detailed explanation of the example described in Table of Contents 1 above. The following section describes the VSEI message syntax and semantics.

[0331] One example is NNPFC (Neural-network post-filter). Characteristics) SEI message syntax and semantics are The following modifications may be made.

[0332] For example, the NNPFC SEI message contains information about the output picture. It is often necessary to include information about the output picture, which corresponds to the input picture. Information regarding whether or not an output picture will be generated, i.e., output picture generation information, may be included. For example, output picture generation information is nnpfc_input_pic_output_ It may be flag[i], where i may be an index indicating a picture. For example, the output picture generation information is nnpfc_input_pic_outpu t_flag[i] is the i-th input picture in NNPF(neural-network It can indicate whether or not a corresponding output picture is generated based on a k post-filter. On the other hand, the value of nnpfc_input_pic_output_flag[i] is the first A value (for example, 1) indicates that NNPF will generate a corresponding output picture. The value of nnpfc_input_pic_output_flag[i] is the second value ( For example, if it is 0), it can indicate that NNPF will not generate a corresponding output picture. Cut.

[0333] On the other hand, input pictures with corresponding output pictures are output after the filtering process. It can be replaced with a picture, and the output picture exists within the bitstream. Good. On the other hand, input pictures that do not have a corresponding output picture will be filtered after the filtering process. It may reside directly within the bitstream.

[0334] Example 2

[0335] Example 2 will provide a detailed explanation of the example described in Table of Contents 2 in the overview of the above example. The following section describes the VSEI message syntax and semantics.

[0336] One example is NNPFC (Neural-network post-filter). Characteristics) SEI message syntax and semantics are The following modifications may be made. The NNPFC SEI message syntax is as follows: It may be signaled as follows. According to the present invention, the output picture generation information is nnpfc The signaling order of _input_pic_output_flag may be changed.

[0337] [Table 21]

[0338] For example, the nnpfc_purpose syntax indicates the purpose of NNPFC. This may be considered a report. For example, for NNPFC purposes, picture rate upsample This may include things like picture rate upsampling. As stated above, I will omit further explanation.

[0339] For example, the nnpfc_id syntax corresponds to information indicating an NNPFC identifier. Yes, you may. This is as explained above, and I will omit the repeated explanation.

[0340] For example, the nnpfc_property_present_flag syntax This refers to filter attributes (e.g., filtering purpose, input formatting, output formatting) Information relating to syntax elements (such as syntax and / or complexity) This is correct. This is as explained above, and a redundant explanation will be omitted.

[0341] For example, the nnpfc_base_flag syntax relates to basic NNPF. Information can be presented, as described above, and therefore, a redundant explanation will be omitted.

[0342] For example, the nnpfc_num_input_pics_minus1 syntax This relates to the number of decoded output pictures used as input to NNPF. It can be information. For example, nnpfc_num_input_pics_minu The value obtained by adding 1 to the value indicated by s1 is the decoded value used as input to NNPF. The number of output pictures can be shown. On the other hand, as an example, nnpfc_num_in The value of put_pics_minus1 may be a value in the range of 0 to 63, and nnp If the value of fc_purpose&0x08 is not 0, nnpfc_num_input The value of _pics_minus1 must be greater than 0.

[0343] For example, the NNPFC SEI message contains information about the output picture. It is often necessary to include information about the output picture, which corresponds to the input picture. Information regarding whether or not an output picture will be generated, i.e., output picture generation information, may be included. For example, output picture generation information is nnpfc_input_pic_output It may be _flag[i], where i may be an index representing a picture. As an example, the output picture generation information is nnpfc_input_pic_outp ut_flag[i] is NNPF(neural-n) in the i-th input picture. Indicates whether or not a corresponding output picture is generated based on etwork post-filter. This is possible. On the other hand, nnpfc_input_pic_output_flag[i] If the value is the first value (for example, 1), NNPF will generate the corresponding output picture. It can be shown, but nnpfc_input_pic_output_flag[i] If the value is a secondary value (for example, 0), NNPF will not generate a corresponding output picture. This can be shown. On the other hand, the output picture generation information is whatever the purpose of NNPF may be. In other words, it is acceptable to include and signal this in SEI messages regardless of the purpose of NNPF. Furthermore, information about the number of input pictures (for example, nnpfc_num_input_p It may be signaled based on ics_minus1). On the other hand, as an example, the output picture The nnpfc_input_pic_output_flag[i] is the generator information. The value may be restricted to a specific value based on other information. For example, nnpfc_inpu The value of t_pic_output_flag[i] is determined based on the purpose of NNPFC, etc. It may be restricted to a fixed value. Here, nnpfc_purpose, i.e., the purpose of NNPFC If the value of the information about the target is a specific value, and 0x08 is that specific value, then the output picture generation information The value of the report nnpfc_input_pic_output_flag[i] is, It must be a predefined value based on the chart index i. The purpose of NNPFC is picture rate upsampling. If it does not include e upsampling, for example, nnpfc_purpose&0 If x08 is 0, the output picture generation information is nnpfc_input_pic_ The output_flag[i] value is a specific value (for example, 1) for a given range of i. It should be equal to 1. For example, at NNPFC The purpose is picture rate upsampling. If it does not include ling, for example, if nnpfc_purpose&0x08 is 0 The output picture generation information is nnpfc_input_pic_output_f The lag[i] value is a specific value (e.g., 1) for at least one i within a certain range of i. It must be determined to be equal to 1. Here is an example. i, which is information used to identify a picture, is a value ranging from 0 to the number of input pictures. (For example, values ​​within the range of nnpfc_num_input_pics_minus1) That's fine.

[0344] For example, the nnpfc_out_sub_c_flag syntax is used under certain conditions. This is related to outSubWidthC and outSubHeightC, and this is, As stated above, I will omit further explanation.

[0345] For example, the nnpfc_out_colour_format_idc syntax This is related to the color format of NNPFC output, as mentioned above. Yes, and I will omit the redundant explanation.

[0346] As an example, nnpfc_pic_width_in_luma_samples Tax and nnpfc_pic_height_in_luma_samples Each tax will be added to the cropped, decoded output picture, along with the nnpfc_id. The resulting lumens sample array of pictures obtained from applying the NNPF identified by The width and height can be shown separately, as explained above, and therefore the redundant explanation is omitted. do.

[0347] For example, the nnpfc_interpolated_pics[i] syntax is Between the i-th picture and the (i+1)-th picture used as input to NNPF This may be information indicating the number of interpolated pictures generated by NNPF. As stated above, I will omit further explanation.

[0348] On the other hand, the variable numInputPics, which indicates the number of pictures input to NNPF, and The variable numOutputP indicates the total number of pictures output as a result from NNPF. ics can be guided as follows:

[0349] [Table 22]

[0350] Example 3

[0351] Example 3 will provide a detailed explanation of the example described in Table of Contents 3 above. The following section describes the VSEI message syntax and semantics.

[0352] One example is NNPFC (Neural-network post-filter). Characteristics) SEI message syntax and semantics are The following modifications may be made.

[0353] For example, the NNPFC SEI message contains information about the output picture. It is often necessary to include information about the output picture, which corresponds to the input picture. Information regarding whether or not an output picture will be generated, i.e., output picture generation information, may be included. For example, output picture generation information is nnpfc_input_pic_output It may be _flag[i], where i may be an index representing a picture. As an example, the output picture generation information is nnpfc_input_pic_outp ut_flag[i] is the NNPF (neural-network) input picture for the i-th input picture. It indicates whether or not a corresponding output picture is generated based on the ork post-filter. Yes, it is possible. On the other hand, the value of nnpfc_input_pic_output_flag[i] If the first value (for example, 1) is selected, it indicates that NNPF will generate a corresponding output picture. This can be done, and the value of nnpfc_input_pic_output_flag[i] is the second A value (e.g., 0) indicates that NNPF will not generate a corresponding output picture. This is possible. On the other hand, the output picture generation information is nnpfc_input_pic_ou The value of tput_flag[i] may be derived based on other information. For example, n The value of npfc_input_pic_output_flag[i] is the output of the NNPFC. It may be guided based on targets, etc. Here, nnpfc_purpose, i.e., NN If the value of the information regarding the purpose of PFC is a specific value, and 0x08 is that specific value, then the output picture The nnpfc_input_pic_output_flag[i] is the generator information. The value may be derived from a predefined value based on the picture index i. As an example, the value of nnpfc_input_pic_output_flag[0] and Other values ​​of nnpfc_input_pic_output_flag[i] are relative to each other. It is acceptable to be led to different values. For example, the purpose of NNPFC is to increase the picture rate. If it does not include picture rate upsampling, for example If nnpfc_purpose&0x08 is 0, then it is the output picture generation information. The value of nnpfc_input_pic_output_flag[0] is a specific value (e.g., For example, it is inferred as 1), and the remaining nnpfc_input_pic_out The value of put_flag[i] may be inferred to be another value (for example, 0). For example, i, which is information used to identify a picture, ranges from 1 to the number of input pictures. Within the range of the corresponding value (for example, nnpfc_num_input_pics_minus1) The value can be within a range.

[0354] Example 4

[0355] Example 4 will provide a detailed explanation of the example described in Table of Contents 4 above. The following section describes the VSEI message syntax and semantics.

[0356] One example is NNPFC (Neural-network post-filter). Characteristics) SEI message syntax and semantics are The following modifications may be made. The NNPFC SEI message syntax is as follows: It may be signaled as follows. According to the present invention, the output picture generation information is nnpfc The signaling order of _input_pic_output_flag may be changed.

[0357] [Table 23]

[0358] Examples include the nnpfc_purpose syntax and the nnpfc_id syntax. S, nnpfc_property_present_flag syntax, nnpf c_base_flag syntax, nnpfc_num_input_pics_m inus1 syntax, nnpfc_out_sub_c_flag syntax, n npfc_out_colour_format_idc syntax, nnpfc_p ic_width_in_luma_samples syntax, nnpfc_pic _height_in_luma_samples syntax, nnpfc_inte The syntax for rpolated_pics[i] is as described above, and I will omit redundant explanations.

[0359] For example, the NNPFC SEI message contains information about the output picture. It is often necessary to include information about the output picture, which corresponds to the input picture. Information regarding whether or not an output picture will be generated, i.e., output picture generation information, may be included. For example, output picture generation information is nnpfc_input_pic_output It may be _flag[i], where i may be an index representing a picture. As an example, the output picture generation information is nnpfc_input_pic_outp ut_flag[i] is the NNPF (neural-network) input picture for the i-th input picture. It indicates whether or not a corresponding output picture is generated based on the ork post-filter. Yes, it is possible. On the other hand, the value of nnpfc_input_pic_output_flag[i] If the first value (for example, 1) is selected, it indicates that NNPF will generate a corresponding output picture. This can be done, and the value of nnpfc_input_pic_output_flag[i] is the second A value (e.g., 0) indicates that NNPF will not generate a corresponding output picture. This is possible. On the other hand, the output picture generation information is, whatever the purpose of NNPF may be, that is, NN Regardless of the purpose of the PF, it may be included in the SEI message and signaled. However, input Information about the number of pictures (for example, nnpfc_num_input_pics_m It may be signaled based on inus1). On the other hand, as an example, the output picture generation information The value of the report nnpfc_input_pic_output_flag[i] is different from other It is acceptable to specify a particular value based on the information. For example, nnpfc_input_pic The value of _output_flag[i] is determined to a specific value based on the purpose of NNPFC, etc. It is permissible to do so. Here, regarding nnpfc_purpose, that is, the purpose of NNPFC If the information value is a specific value, and 0x08 is that specific value, then n is the output picture generation information. The value of npfc_input_pic_output_flag[i] is the picture input It should be a predefined value based on the dex value i. For example, NNP. The purpose of FC is picture rate upsampling. If it does not include ampling, for example nnpfc_purpose&0x08 is 0 If so, the output picture generation information is nnpfc_input_pic_outpu The value of t_flag[0] should be a specific value (for example, 1) (shall b e equals 1). In other words, it may be limited to a specific value. On the other hand, as another example... Here, the value of nnpfc_purpose, i.e., information about the purpose of NNPFC, is If it is a specific value, and 0x08 is the specific value, then the output picture generation information is nnpfc_ The value of input_pic_output_flag[i] is the picture index. It may be inferred (inferred) to a predefined value based on a certain i. For example, NN The purpose of PFC is picture rate upsampling. If it does not include sampling, for example nnpfc_purpose&0x08 If it is 0, the output picture generation information is nnpfc_input_pic_outp The value of ut_flag[0] may be inferred to a specific value (for example, 1). In other words, it is acceptable to be guided to a specific value.

[0360] On the other hand, the variable numInputPics, which indicates the number of pictures input to NNPF, and The variable numOutputP indicates the total number of pictures output as a result from NNPF. The ICS may be induced as shown in Table 22 above.

[0361] According to the embodiments described in this disclosure, in addition to solving the aforementioned prior art, information By adjusting the signaling order and clarifying the semantics of the information, This can reduce decoder errors and improve coding quality and efficiency.

[0362] Examples of video decoding methods

[0363] The following describes video encoding and video decoding methods according to various embodiments of the present invention. The video decoding method in 6 may be performed by the video decoding device 200, as shown in Figure 7. The method may be performed by the video encoding device 100. Also, the video decoding in Figures 6 and 7 The encoding method is performed based on the above-described examples (including Examples 1 to 4). good.

[0364] Figure 6 illustrates a video decoding method performed by a video decoding device according to one embodiment of the present disclosure. This is a diagram for clarity.

[0365] First, as an example, a post-filter-based corresponding output picture for an input picture. - Information is NNPFC (neural-network post-filter cha (acteristics)SEI(supplemental enhancement) (t information) Can be obtained from the message (S610). Postfile The corresponding output picture information in Rutabase may refer to NNPF-related information. The linked information is as explained in the table above.

[0366] Subsequently, based on the acquired corresponding output picture information, the corresponding input picture is processed. The output picture may be acquired (S620). In this case, the corresponding output picture information is entered. If it indicates that there is no corresponding output picture for the power picture, the corresponding output picture The 'char' does not need to be acquired, and the corresponding output picture information is the corresponding output for the input picture. If it indicates that a picture exists, the corresponding output picture may be retrieved. Here, The corresponding output picture information is the corresponding output picture of the post-filter for the input picture. - May include output picture generation information regarding whether or not it was generated. This may include the nnpfc_input_pic_output_flag mentioned above. On the other hand, output picture generation information may be obtained based on specific conditions. For example, The values ​​of the output picture generation information may be restricted to specific values ​​based on certain conditions. Another example is that the values ​​of the output picture generation information are inferred to be specific values ​​based on certain conditions. It is also possible that certain conditions are associated with the purpose of post-filtering. However, under specific conditions, the purpose of the post-filter is picture rate upsampling. It can be associated with whether or not it is the case. For example, a specific condition is the input picture pic It may be associated with the marker index. On the other hand, based on specific conditions, a small number within a specific range However, the value of the output picture generation information for a single input picture is restricted to a specific value. This is also good. For example, based on specific conditions, at least one input picture within a specific range The value of the output picture generation information may be limited to 1. On the other hand, a specific range is the input picture Information related to the number of churners, but also related to the number of input pictures, may be signaled. On the other hand, as an example, if the picture index of the input picture is a specific value (for example, 0) If so, the value of the output picture generation information for the input picture is a specific value (for example, 1) may be limited. On the other hand, as an example, certain conditions may be associated with the number of input pictures. It may also be possible to have a corresponding output picture for an input picture. Based on this, the input picture will be converted to the corresponding output picture (in the picture stream). (It may be replaced.) On the other hand, as an example, the output picture generation information Regardless of whether the purpose of the post-filter includes picture rate upsampling or not. It's okay to signal.

[0367] Subsequently, although not shown in the diagram, the picture is restored based on the corresponding output picture information. It's okay.

[0368] On the other hand, the video decoding method shown in Figure 6 corresponds to one embodiment of this disclosure, so certain steps are changed. The order of the stages may be changed, or some stages may be added or deleted. It should be obvious that any such modifications are also included in this disclosure.

[0369] Examples of video encoding methods

[0370] Figure 7 illustrates a video encoding method performed by a video encoding device according to one embodiment of the present disclosure. This is a diagram for clarity.

[0371] First, post-filter-based corresponding output picture information is determined for the input picture. (S710) This may be done. Post-filter-based corresponding output picture information is NNP This can refer to F-related information. NNPF-related information is as explained by referring to the table above, etc. be.

[0372] Subsequently, the corresponding output picture information is NNPFC (neural-network p ost-filter characteristics)SEI(supplemen (tal enhancement information) is signaled by a message (S720) In this case, the corresponding output picture information corresponds to the input picture. It can indicate that no output picture exists, and conversely, it can indicate the corresponding input picture. It is also possible to indicate that an output picture exists. Here, the corresponding output picture information is Output indicating whether or not a corresponding output picture is generated for the input picture after applying a post-filter. It may include picture generation information. The output picture generation information includes the nnpfc_in mentioned above. The put_pic_output_flag may be included. On the other hand, the output picture generation information The information may be signaled based on specific conditions. On the other hand, as an example, output picture generation The values ​​of the information may be restricted to specific values ​​based on certain conditions. However, other examples include, The values ​​of the force picture generation information can also be inferred to specific values ​​based on certain conditions. For example, a specific condition may be associated with the purpose of post-filtering, but the specific condition may be pre This relates to whether or not the purpose of the post-filter is picture rate upsampling. It may be done. For example, certain conditions are related to the picture index of the input picture. It may be attached. On the other hand, based on certain conditions, at least one input pic within a certain range The values ​​of the output picture generation information for the char may be restricted to specific values. For example, specific conditions Based on the case, output picture raw for at least one input picture within a specific range The value of the information may be limited to 1. On the other hand, as an example, the picture-in of the input picture. If DEX is a specific value (for example, 0), then the output picture for the input picture The values ​​of the generated information may be restricted to specific values ​​(for example, 1). On the other hand, as an example, specific conditions This may be associated with the number of input pictures. Also, as an example, the input pictures may be Based on the existence of a corresponding output picture, the input picture is (Pictures (In the stream) it may be replaced with the corresponding output picture. On the other hand, as an example, output The picture generation information indicates that the purpose of the post-filter is to upsample the picture rate. It may be signaled regardless of whether it is included or not.

[0373] Subsequently, although not shown in the diagram, the picture is restored based on the corresponding output picture information. It's okay.

[0374] As another example, the bitstream generated by the video encoding method is recorded. A computer-readable medium may be provided, and the bitstream generated by the video encoding method A method for transmitting the signal may be provided.

[0375] On the other hand, the video encoding method shown in Figure 7 corresponds to one embodiment of the present disclosure, so certain steps are changed. The order of the stages may be changed, or some stages may be added or deleted. It should be obvious that any such modifications are also included in this disclosure.

[0376] According to this application, the meaning of the information that can be included in VSEI can be clarified, and therefore, Deco This reduces errors in coding, allows for more accurate scenario representation, and improves coding quality. This is possible. Moreover, according to this application, it is possible to change the signaling order of specific information. Therefore, coding efficiency can be improved.

[0377] Figure 8 illustrates a content streaming system to which the embodiments of this disclosure can be applied. This is a diagram.

[0378] As shown in Figure 8, a content streaming system to which an embodiment of the present disclosure is applied is Broadly speaking, encoding servers, streaming servers, web servers, media It may include a storage device, user equipment, and multimedia input devices.

[0379] The aforementioned encoding server is used by smartphones, cameras, camcorders, etc. The content input from the multimedia input device is compressed into digital data and bit It is responsible for generating the stream and sending it to the streaming server. Other examples For example, multimedia input devices such as smartphones, cameras, and camcorders When generating a bitstream directly, the encoding server can be omitted. stomach.

[0380] The bitstream is a video encoding method and / or video to which an embodiment of the present disclosure is applied. It may be generated by an encoding device, and the streaming server will use the bitstream The bitstream can be temporarily stored during the transmission or reception process.

[0381] The aforementioned streaming server performs a mal based on a user request via the web server. The web server transmits the media data to the user's device and tells the user any It can act as an intermediary to inform users whether the service is available. If you request the desired service from the bar, the web server will then stream it to the streaming server. - The data is transmitted to the streaming server, which then transmits the multimedia data to the user. This is possible. In this case, the content streaming system is a separate control server This may include, in which case the control server is the content streaming system It can play a role in controlling the commands and responses between the various devices within the system.

[0382] The aforementioned streaming server includes media storage and / or encoding servers. Content can be received from the bar. For example, from the encoding server When receiving content, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server The bitstream can be stored for a certain period of time.

[0383] Examples of user devices include mobile phones, smartphones, Notebook computers, digital broadcasting terminals, PDAs personal digital assistants), PMP(portable) e multimedia player), navigation, slate PC (slat ePC, tablet PC, ultrabook OK), wearable device (for example, watch type) Smartwatch, smart glasses, HMD head-mounted display, digital TV, desktop computer This could include things like digital signage.

[0384] Each server within the aforementioned content streaming system operates as a distributed server. In this case, the data received by each server may be processed in a distributed manner.

[0385] The scope of this disclosure is limited to cases where the operation according to the methods of various embodiments is performed on a device or computer. Software or machine-executable instructions (e.g., operational system, application) that enable this. Software, firmware, programs, etc., and such software A non-temporary computer that stores software or instructions and can be executed on a device or computer. Computer-readable media (non-transitory computer-readable) (Includes medium)

[0386] [Industrial applicability] The embodiments described herein can be used for encoding / decoding video.

[0387] [Claims when filing an international application] [Claim 1] A video decoding method performed by a video decoding device, Post-filter-based correspondence for input picture information to output picture information NNPFC (neural-network post-filter characterist ics)SEI(supplemental enhancement informa The stage of obtaining information from the message (tion), Based on the corresponding output picture information, the corresponding output picture for the input picture This includes the step of obtaining the marker, The aforementioned corresponding output picture information is the corresponding output of the post-filter for the input picture. A video decoding method that includes output picture generation information regarding whether or not a picture was generated. [Claim 2] The value of the output picture generation information is limited to a specific value based on specific conditions, claim The video decoding method described in 1. [Claim 3] The aforementioned specific conditions are associated with the purpose of the post-filter, as described in claim 1. Decryption method. [Claim 4] The aforementioned specific conditions are that the purpose of the post-filter is picture rate upsampling. The video decoding method according to claim 3, which is further associated with whether or not it exists. [Claim 5] The aforementioned specific conditions are further associated with the picture index of the input picture. The video decoding method according to claim 3. [Claim 6] Based on the aforementioned specific conditions, for at least one of the input pictures within a specific range The video decoding method according to claim 5, wherein the value of the output picture generation information is limited to 1. [Claim 7] The aforementioned specific range relates to the number of input pictures signaled in the bitstream. A video decoding method according to claim 6, determined based on information. [Claim 8] If the picture index of the input picture is 0, the input picture The value of the output picture generation information for the video decoding described in claim 5 is limited to 1. method. [Claim 9] The aforementioned specific condition is associated with the number of input pictures, as described in claim 1. Encoding method. [Claim 10] The output picture generation information is determined by the purpose of the post-filter, which is to increase the picture rate. The video decoding method according to claim 1, which is signaled regardless of whether or not it includes sampling. . [Claim 11] A video encoding method performed by a video encoding device, The stage for determining post-filter-based corresponding output picture information for the input picture. Floor and The corresponding output picture information is NNPFC (neural-network post -filter characteristics)SEI(supplemental The stage where the message is signaled (enhancement information) , including, The aforementioned corresponding output picture information is the corresponding output of the post-filter for the input picture. A video encoding method that includes output picture generation information regarding whether or not a picture is generated. [Claim 12] The bitstream generated by the video encoding method described in claim 11 is recorded. Computer-readable media. [Claim 13] A method for transmitting a bitstream generated by a video encoding method, The aforementioned video encoding method is The stage for determining post-filter-based corresponding output picture information for the input picture. Floor and The corresponding output picture information is NNPFC (neural-network post -filter characteristics)SEI(supplemental The stage where the message is signaled (enhancement information) , including, The aforementioned corresponding output picture information is the corresponding output of the post-filter for the input picture. A method that includes output picture generation information regarding whether or not a picture was generated.

Claims

1. A method for image decoding, The stage for obtaining NNPFC (Neural-Network Post-Filter Characteristics) SEI (Supplemental Enhancement Information) messages; The NNPFC SEI message comprises one or more picture filtering flags corresponding to one or more input pictures, and purpose information indicating the purpose of the post-filter. A step of determining whether a corresponding output picture for an input picture is generated by the post-filter based on the picture filtering flags corresponding to the input picture included in the one or more picture filtering flags; and The value of the picture filtering flag being equal to 1 indicates that the post-filter generates the corresponding output picture for the input picture, and A value equal to 0 for the picture filtering flag indicates that the post-filter does not generate the corresponding output picture for the input picture. The step of generating the corresponding output picture with respect to the input picture by the post-filter based on the above determination; A video decoding method wherein, regardless of whether the objective information relates to picture rate upsampling, one or more picture filtering flags are included in the acquired NNPFC SEI message.

2. The video decoding method according to claim 1, wherein at least one of the one or more picture filtering flags has a value of 1, based on the fact that the objective information is not related to picture rate upsampling.

3. The video decoding method according to claim 1, wherein the number of one or more input pictures and the number of one or more picture filtering flags are determined based on numerical information included in the NNPFC SEI message.

4. The video decoding method according to claim 1, wherein, based on the number of one or more input pictures being equal to 1, the one or more picture filtering flags include corresponding picture filtering flags having a value equal to 1.

5. A video encoding method, The steps for generating NNPFC (Neural-Network Post-Filter Characteristics) SEI (Supplemental Enhancement Information) messages; and The NNPFC SEI message comprises one or more picture filtering flags corresponding to one or more input pictures, and purpose information indicating the purpose of the post-filter. The picture filtering flag is used to determine whether a corresponding output picture for an input picture is generated by the post-filter. The value of the picture filtering flag being equal to 1 indicates that the post-filter generates the corresponding output picture for the input picture, and A value equal to 0 for the picture filtering flag indicates that the post-filter does not generate the corresponding output picture for the input picture. The process includes the step of encoding video information including the aforementioned NNPFC SEI message; A video encoding method wherein, regardless of whether the objective information is related to picture rate upsampling, one or more picture filtering flags are included in the NNPFC SEI message.

6. A method for transmitting image data, The stage of generating a bitstream; and The step of transmitting the bitstream; The aforementioned bitstream is The steps for generating NNPFC (Neural-Network Post-Filter Characteristics) SEI (Supplemental Enhancement Information) messages; and The process involves encoding video information including the aforementioned NNPFC SEI message; and the result is generated by this process. The NNPFC SEI message comprises one or more picture filtering flags corresponding to one or more input pictures, and purpose information indicating the purpose of the post-filter. The picture filtering flag is used to determine whether a corresponding output picture for an input picture is generated by the post-filter. The value of the picture filtering flag being equal to 1 indicates that the post-filter generates the corresponding output picture for the input picture, and A value equal to 0 for the picture filtering flag indicates that the post-filter does not generate the corresponding output picture for the input picture. A transmission method wherein, regardless of whether the objective information is related to picture rate upsampling, one or more picture filtering flags are included in the NNPFC SEI message.