Image encoding / decoding method and device, and recording medium on which bitstream is stored

WO2026197727A1PCT designated stage Publication Date: 2026-09-24LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/004220
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-16
Filing Date
2026-03-16
Publication Date
2026-09-24

Smart Images

  • Figure KR2026004220_24092026_PF_FP_ABST
    Figure KR2026004220_24092026_PF_FP_ABST
Patent Text Reader

Abstract

An image decoding method and device, according to the present disclosure, can receive a bitstream including encoded video pictures and reconstruct the encoded video pictures included in the bitstream. Here, the bitstream can include an encoder optimization information (EOI) supplemental enhancement information (SEI) message. The EOI SEI message can include EOI type information indicating the type of optimization method. A quantization parameter value of a picture for applying the EOI SEI message can be derived on the basis of initial QP information and QP delta information.
Need to check novelty before this filing date? Find Prior Art

Description

Video encoding / decoding method and device, and a recording medium storing a bitstream

[0001] The present invention relates to a video encoding / decoding method and apparatus, and a recording medium storing a bitstream.

[0002] Recently, the demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition) video, has been increasing across various application fields, and accordingly, high-efficiency video compression technologies are being discussed.

[0003] Various image compression technologies exist, such as inter-prediction technology that predicts pixel values ​​in the current picture from previous or subsequent pictures, intra-prediction technology that predicts pixel values ​​in the current picture using pixel information within the current picture, and entropy coding technology that assigns short codes to values ​​with high frequency and long codes to values ​​with low frequency; by utilizing these image compression technologies, image data can be effectively compressed for transmission or storage.

[0004] The present disclosure provides a method and apparatus for configuring encoder optimization information.

[0005] The present disclosure provides a method and apparatus for signaling encoder optimization information.

[0006] The present disclosure provides a method and apparatus for deriving quantization parameter values ​​for applying encoder optimization information.

[0007] A video decoding method and apparatus according to the present disclosure receive a bitstream including an encoded video picture and can restore the encoded video picture included in the bitstream. The bitstream may include an Encoder Optimization Information (EOI) SEI (Supplemental Enhancement Information) message. The EOI SEI message may include EOI type information indicating the type of optimization method. The quantization parameter value of the picture for applying the EOI SEI message may be derived based on initial QP (Quantization Parameter) information and QP delta information.

[0008] In the image decoding method and apparatus according to the present disclosure, the initial QP information can specify an initial value of QP for each slice referencing a picture parameter set.

[0009] In the image decoding method and apparatus according to the present disclosure, the initial QP information may be signaled from a set of picture parameters of the bitstream.

[0010] In the image decoding method and apparatus according to the present disclosure, the QP delta information can specify an initial value of QP to be used for blocks within a slice.

[0011] In the image decoding method and apparatus according to the present disclosure, the QP delta information may relate to the first slice segment among a plurality of slice segments belonging to the picture.

[0012] In the image decoding method and apparatus according to the present disclosure, the QP delta information can be signaled in the slice segment header of the bitstream.

[0013] In the image decoding method and apparatus according to the present disclosure, the QP delta information may relate to the first slice among a plurality of slices belonging to the picture.

[0014] In the image decoding method and apparatus according to the present disclosure, the QP delta information can be signaled in the slice header of the bitstream.

[0015] A video encoding method and apparatus according to the present disclosure may receive a video picture to be encoded, encode the received video picture to generate video information regarding the video picture, generate an Encoder Optimization Information (EOI) SEI (Supplemental Enhancement Information) message, and generate a bitstream including the video information and the EOI SEI message. The EOI SEI message may include EOI type information indicating the type of optimization method. The quantization parameter value of the picture to which the EOI SEI message is applied may be derived based on initial Quantization Parameter (QP) information and QP delta information.

[0016] A computer-readable digital storage medium is provided that stores encoded video / image information that causes an image decoding method to be performed by a decoding device according to the present disclosure.

[0017] A computer-readable digital storage medium is provided that stores video / image information generated according to the image encoding method according to the present disclosure.

[0018] A method and apparatus for transmitting video / image information generated according to the image encoding method according to the present disclosure are provided.

[0019] According to the present disclosure, by defining encoder optimization information, it is possible to easily check whether optimization is performed for human viewing or machine analysis, and whether optimization is applied during the pre-processing or encoding process.

[0020] According to the present disclosure, encoder optimization information can be implemented to meet the intended functions and purposes of each standard.

[0021] FIG. 1 illustrates a video / image coding system according to the present disclosure.

[0022] FIG. 2 shows a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and to which encoding of a video / image signal is performed.

[0023] FIG. 3 shows a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and to which decoding of a video / image signal is performed.

[0024] FIG. 4 illustrates a method for restoring a video picture performed in a decoding device (300) according to the present disclosure.

[0025] FIG. 5 illustrates a schematic configuration of a decoding device (300) that performs a method for restoring a video picture according to the present disclosure.

[0026] FIG. 6 illustrates a method for generating a bitstream performed in an encoding device (200) according to the present disclosure.

[0027] FIG. 7 illustrates a schematic configuration of an encoding device (200) that performs a method for generating a bitstream according to the present disclosure.

[0028] FIG. 8 shows an example of a content streaming system to which embodiments of the present disclosure can be applied.

[0029] The present disclosure is susceptible to various modifications and may have various embodiments; specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, and it should be understood that it includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the present disclosure. Similar reference numerals have been used for similar components in the description of each drawing.

[0030] Terms such as "first," "second," etc., may be used to describe various components, but said components should not be limited by said terms. Such terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present disclosure, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.

[0031] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.

[0032] The terms used in this application are used merely to describe specific embodiments and are not intended to limit the disclosure. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, terms such as “comprising” or “having” are intended to specify the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0033] The present disclosure relates to video / video coding. For example, the methods / embodiments disclosed herein may be applied to methods disclosed in the VVC (versatile video coding) standard. Additionally, the methods / embodiments disclosed herein may be applied to methods disclosed in the EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard), or next-generation video / video coding standards (e.g., H.267 or H.268).

[0034] This specification presents various embodiments regarding video / image coding, and unless otherwise noted, said embodiments may be performed in combination with one another.

[0035] In this specification, "video" may refer to a set of images over time. "Picture" generally refers to a unit representing a single image of a specific time period, and "slice" or "tile" is a unit that constitutes a part of a picture in coding. A slice or tile may contain one or more coding tree units (CTUs). A picture may consist of one or more slices or tiles. A tile is a rectangular area composed of multiple CTUs within a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of ​​CTUs having a height equal to the height of the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of ​​CTUs having a height specified by the picture parameter set and a width equal to the width of the picture. CTUs within a tile are arranged continuously according to the CTU raster scan, whereas tiles within a picture may be arranged continuously according to the tile's raster scan. A single slice may include an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that can be exclusively contained in a single NAL unit. Meanwhile, a single picture may be divided into two or more subpictures. A subpicture may be a rectangular area of ​​one or more slices within a picture.

[0036] A pixel, or pel, can refer to the smallest unit that constitutes a picture (or image). Additionally, the term 'sample' may be used as a counterpart to pixel. A sample generally represents a pixel or its value, and it may represent only the pixel / pixel value of the luminance (luma) component or only the pixel / pixel value of the chroma component.

[0037] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of ​​a picture and information related to that area. A unit may include one luminance block and two chroma (e.g., cb, cr) blocks. Depending on the case, the term unit may be used interchangeably with terms such as block or area. In general, an MxN block may include samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows.

[0038] In this specification, "A or B" may mean "only A," "only B," or "both A and B." Alternatively, in this specification, "A or B" may be interpreted as "A and / or B." For example, in this specification, "A, B or C" may mean "only A," "only B," "only C," or "any combination of A, B and C."

[0039] A slash ( / ) or a comma used in this specification may mean "and / or." For example, "A / B" may mean "A and / or B." Accordingly, "A / B" may mean "only A," "only B," or "both A and B." For example, "A, B, C" may mean "A, B or C."

[0040] In this specification, "at least one of A and B" may mean "only A," "only B," or "both A and B." Additionally, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted as synonymous with "at least one of A and B."

[0041] Additionally, in this specification, "at least one of A, B and C" may mean "only A," "only B," "only C," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" may mean "at least one of A, B and C."

[0042] Additionally, parentheses used in this specification may mean "for example." Specifically, where indicated as "prediction (intra-prediction)," "intra-prediction" may be proposed as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra-prediction," and "intra-prediction" may be proposed as an example of "prediction." Furthermore, even where indicated as "prediction (i.e., intra-prediction)," "intra-prediction" may be proposed as an example of "prediction."

[0043] Technical features described individually within a single drawing in this specification may be implemented individually or simultaneously.

[0044] FIG. 1 illustrates a video / image coding system according to the present disclosure.

[0045] Referring to FIG. 1, the video / image coding system may include a first device (source device) and a second device (receiving device).

[0046] A source device can transmit encoded video / image information or data in the form of a file or streaming to a receiving device via a digital storage medium or a network. The source device may include a video source, an encoding device, and a transmission unit. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. A transmitter may be included in the encoding device. A receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.

[0047] A video source may acquire video / images through processes such as video / image capture, synthesis, or generation. The video source may include a video / image capture device and / or a video / image generation device. A video / image capture device may include one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device may include a computer, a tablet, a smartphone, etc., and may generate video / images (electronically). For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process in which related data is generated.

[0048] The encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0049] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device in the form of a file or streaming via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit may include elements for creating a media file through a predetermined file format and elements for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.

[0050] The decoding device can decode video / images by performing a series of procedures such as inverse quantization, inverse transform, and prediction corresponding to the operation of the encoding device.

[0051] The renderer can render the decoded video / image. The rendered video / image can be displayed through the display unit.

[0052] FIG. 2 shows a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and to which encoding of a video / image signal is performed.

[0053] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a predictor (220), a residual processor (230), an entropy encoder (240), an adder (250), a filter (260), and a memory (270). The predictor (220) may include an inter-predictor (221) and an intra-predictor (222). The residual processor (230) may include a transformer (232), a quantizer (233), a dequantizer (234), and an inverse transformer (235). The residual processor (230) may further include a subtractor (231). The addition unit (250) may be referred to as a reconstructor or a reconstructed block generator. The above-described image segmentation unit (210), prediction unit (220), residual processing unit (230), entropy encoding unit (240), addition unit (250), and filtering unit (260) may be configured by one or more hardware components (e.g., an encoding device chipset or processor) according to the embodiment. Additionally, the memory (270) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.

[0054] The image segmentation unit (210) can divide an input image (or picture, frame) input to an encoding device (200) into one or more processing units. For example, the processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure.

[0055] For example, a single coding unit may be divided into multiple coding units with a deeper depth based on a quad tree structure, a binary tree structure, and / or a terrestrial structure. In this case, for example, the quad tree structure may be applied first and the binary tree structure and / or terrestrial structure may be applied later. Alternatively, the binary tree structure may be applied before the quad tree structure. A coding procedure according to the present specification may be performed based on a final coding unit that is no longer divided. In this case, based on coding efficiency according to image characteristics, the maximum coding unit may be used directly as the final coding unit, or, if necessary, the coding unit may be recursively divided into coding units of a lower depth so that a coding unit of the optimal size may be used as the final coding unit. Here, the term "coding procedure" may include procedures such as prediction, transformation, and restoration described below.

[0056] As another example, the processing unit may further include a Prediction Unit (PU) or a Transform Unit (TU). In this case, the Prediction Unit and the Transform Unit may each be divided or partitioned from the aforementioned final coding unit. The Prediction Unit may be a unit for sample prediction, and the Transform Unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from transformation coefficients.

[0057] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the context. In general, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chroma component. A sample may be used to refer to a single picture (or image) as a term corresponding to a pixel or pel.

[0058] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (prediction block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input video signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input video signal (original block, original sample array) within the encoding device (200) may be called a subtraction unit (231).

[0059] The prediction unit (220) performs a prediction for a block to be processed (hereinafter referred to as the current block) and can generate a predicted block containing prediction samples for the current block. The prediction unit (220) can determine whether intra prediction is applied or inter prediction is applied at the current block or CU level. The prediction unit (220) can generate various information regarding the prediction, such as prediction mode information, as described below in the description of each prediction mode, and transmit it to the entropy encoding unit (240). The information regarding the prediction can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.

[0060] The intra prediction unit (222) can predict the current block by referring to samples within the current picture. The referenced samples may be located near the current block or at a certain distance from the current block depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The non-directional mode may include at least one DC mode or a planar mode. The directional mode may include 33 directional modes or 65 directional modes depending on the degree of fineness of the prediction direction. However, this is merely an example, and depending on the settings, more or fewer directional modes may be used. The intra prediction unit (222) may determine the prediction mode applied to the current block by using the prediction mode applied to the surrounding blocks.

[0061] The inter prediction unit (221) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture containing the reference blocks and the reference picture containing the temporal neighboring blocks may be the same or different. The above temporal surrounding blocks may be referred to by names such as collocated reference block, collocated CU (colCU), etc., and the reference picture containing the above temporal surrounding blocks may be referred to as a collocated picture (colPic). For example, the inter prediction unit (221) may construct a list of motion information candidates based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of surrounding blocks as motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vectors of surrounding blocks are used as motion vector predictors, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0062] The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit may apply intra prediction or inter prediction for prediction of a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called a combined inter and intra prediction (CIIP) mode. Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in screen content coding (SCC) for games. IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may utilize at least one of the inter prediction techniques described in this specification. The palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, sample values ​​within the picture can be signaled based on information regarding the palette table and palette index. The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or to generate a residual signal.

[0063] The transformation unit (232) can generate transform coefficients by applying a transformation technique to a residual signal. For example, the transformation technique may include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transformation obtained based on a prediction signal generated using all previously restored pixels. Additionally, the transformation process may be applied to a pixel block of the same size in a square, or to a block of variable size that is not square.

[0064] The quantization unit (233) quantizes the transformation coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information regarding the quantized transformation coefficients) and output it as a bitstream. The information regarding the quantized transformation coefficients may be called residual information. The quantization unit (233) can rearrange the block-shaped quantized transformation coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information regarding the quantized transformation coefficients based on the one-dimensional vector-shaped quantized transformation coefficients.

[0065] The entropy encoding unit (240) can perform various encoding methods such as exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit (240) may encode information required for video / image restoration (e.g., values ​​of syntax elements, etc.) together or separately, in addition to quantized transform coefficients.

[0066] Encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream at the level of a Network Abstraction Layer (NAL) unit. The video / image information may further include information regarding various parameter sets, such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Additionally, the video / image information may further include general constraint information. In this specification, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the encoding procedure described above and included in the bitstream. The bitstream may be transmitted over a network or stored on a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) that transmits the signal output from the entropy encoding unit (240) and / or a storage unit (not shown) that stores it may be configured as internal / external elements of the encoding device (200), or the transmission unit may be included in the entropy encoding unit (240).

[0067] Quantized transformation coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, a residual signal (residual block or residual samples) can be restored by applying inverse quantization and inverse transformation to the quantized transformation coefficients through the inverse quantization unit (234) and the inverse transformation unit (235). An adder (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter-prediction unit (221) or the intra-prediction unit (222). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The adder (250) may be called a reconstruction unit or a reconstruction block generation unit. The generated restoration signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after filtering as described below. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.

[0068] The filtering unit (260) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (260) can generate a modified restored picture by applying various filtering methods to the restored picture, and can store the modified restored picture in memory (270), specifically in the DPB of memory (270). The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (240). The information regarding filtering can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.

[0069] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter-prediction unit (221). Through this, when inter-prediction is applied, the encoding device can avoid prediction mismatches between the encoding device (200) and the decoding device, and can also improve encoding efficiency.

[0070] The DPB of the memory (270) can store the modified restored picture to be used as a reference picture in the inter-prediction unit (221). The memory (270) can store motion information of blocks from which motion information is derived (or encoded) within the current picture and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (221) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (270) can store restoration samples of the blocks restored within the current picture and transmit them to the intra-prediction unit (222).

[0071] FIG. 3 shows a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and to which decoding of a video / image signal is performed.

[0072] Referring to FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-predictor (332) and an intra-predictor (331). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321).

[0073] The aforementioned entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) may be configured by a single hardware component (e.g., a decoding device chipset or processor) according to an embodiment. Additionally, the memory (360) may include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.

[0074] When a bitstream containing video / image information is input, the decoding device (300) can restore the image in correspondence with the process in which the video / image information is processed by the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied by the encoding device. Accordingly, the processing unit for decoding may be a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a binary tree structure. One or more conversion units may be derived from the coding unit. And, the restored image signal decoded and output through the decoding device (300) can be played back through a playback device.

[0075] The decoding device (300) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through an entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may further include general constraint information. The decoding device can decode the picture based on information regarding the parameter sets and / or the general constraint information. The signaling / receiving information and / or syntax elements described below in this specification may be decoded through the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit (310) can decode information within the bitstream based on coding methods such as exponential chord coding, CAVLC, or CABAC, and output the values ​​of syntax elements required for image restoration and the quantized values ​​of transformation coefficients regarding residuals. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, determine a context model using information on the syntax element to be decoded and decoding information of surrounding and decoding target blocks or information on symbols / bins decoded in the previous step, predict the probability of occurrence of the bin according to the determined context model, and perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin after determining the context model. Among the information decoded in the entropy decoding unit (310), information regarding prediction is provided to the prediction unit (inter prediction unit (332) and intra prediction unit (331)), and the residual value for which entropy decoding was performed in the entropy decoding unit (310), i.e., quantized transformation coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive residual signals (residual blocks, residual samples, residual sample array). Additionally, among the information decoded in the entropy decoding unit (310), information regarding filtering can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (300), or the receiving unit may be a component of the entropy decoding unit (310).

[0076] Meanwhile, the decoding device according to the present specification may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoding device (video / image / picture information decoding device) and a sample decoding device (video / image / picture sample decoding device). The information decoding device may include the entropy decoding unit (310), and the sample decoding device may include at least one of the inverse quantization unit (321), inverse transform unit (322), adder (340), filtering unit (350), memory (360), inter prediction unit (332), and intra prediction unit (331).

[0077] In the inverse quantization unit (321), the quantized transformation coefficients can be inversely quantized to output transformation coefficients. The inverse quantization unit (321) can rearrange the quantized transformation coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain transformation coefficients.

[0078] In the inverse conversion unit (322), the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).

[0079] The prediction unit (320) can perform a prediction for the current block and generate a predicted block containing prediction samples for the current block. The prediction unit (320) can determine whether an intra prediction or an inter prediction is applied to the current block based on information regarding the prediction output from the entropy decoding unit (310), and can determine a specific intra / inter prediction mode.

[0080] The prediction unit (320) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit (320) may apply intra prediction or inter prediction for prediction of a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called a combined inter and intra prediction (CIIP) mode. Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, such as SCC (screen content coding). IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may utilize at least one of the inter prediction techniques described in this specification. The palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, information regarding the palette table and palette index can be included in the above video / image information and signaled.

[0081] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block according to the prediction mode, or may be located at a certain distance from the current block. In intra prediction, the prediction modes may include one or more non-directional modes and a plurality of directional modes. The intra prediction unit (331) may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.

[0082] The inter prediction unit (332) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) may construct a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-prediction can be performed based on various prediction modes, and information regarding the prediction may include information indicating the inter-prediction mode for the current block.

[0083] The adder (340) can generate a restoration signal (restoration picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter prediction unit (332) and / or the intra prediction unit (331)). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the prediction block can be used as the restoration block.

[0084] The addition unit (340) may be called a restoration unit or a restoration block generation unit. The generated restoration signal may be used for intra-predicting the next block to be processed within the current picture, may be output after filtering as described below, or may be used for inter-predicting the next picture. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.

[0085] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can generate a modified restored picture by applying various filtering methods to the restored picture, and can transmit the modified restored picture to memory (360), specifically to the DPB of memory (360). The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0086] The (modified) restored picture stored in the DPB of the memory (360) can be used as a reference picture in the inter prediction unit (332). The memory (360) can store motion information of blocks from which motion information within the current picture has been derived (or decoded) and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter prediction unit (332) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (360) can store restoration samples of blocks restored within the current picture and transmit them to the intra prediction unit (331).

[0087] In this specification, the embodiments described in the filtering unit (260), inter prediction unit (221), and intra prediction unit (222) of the encoding device (200) may be applied to the filtering unit (350), inter prediction unit (332), and intra prediction unit (331) of the decoding device (300) in the same or corresponding manner.

[0088] FIG. 4 illustrates a method for restoring a video picture performed in a decoding device (300) according to the present disclosure.

[0089] A bitstream containing an encoded video picture can be received (S400).

[0090] The encoded video picture of the bitstream can be restored (S410).

[0091] Video information regarding an encoded video picture can be extracted from the bitstream. Based on the extracted video information, the encoded video picture can be restored.

[0092] The bitstream may contain Encoder Optimization Information (EOI). Encoder Optimization Information can be used to indicate whether the video is optimized for human viewing or machine analysis, and what type of optimization was applied during the pre-processing or encoding process.

[0093] Encoder optimization information can be configured in the supplemental enhancement information (SEI) message of the bitstream. In this case, the encoder optimization information can be named the EOI SEI message.

[0094] The EOI SEI message according to the present disclosure can be configured as shown in Table 1 below.

[0095] encoder_optimization_info(payloadSize ) {Descriptoreoi_cancel_flagu(1)if( !eoi_cancel_flag ) {eoi_persistence_flagu(1)eoi_for_human_viewing_idcu(2)eoi_for_machine_analysis_idcu(2)eoi_reserved_zero_2bitsu(2)eoi_typeu(16)if( EoiObjectBasedFlag ) {eoi_object_based_idcu(16)if( eoi_object_based_idc & 0x02 ) {eoi_quant_threshold_deltaue(v)if( eoi_quant_threshold_delta > 0 )eoi_pic_quant_object_flagu(1)}}if( EoiTemporalResamplingFlag ) {eoi_temporal_resampling_type_flagu(1)eoi_num_int_picsue(v)if( eoi_temporal_resampling_type_flag &&eoi_num_int_pics > 0 )eoi_src_pic_flagu(1)}if( EoiSpatialResamplingFlag ) {eoi_orig_pic_dimensions_flagu(1)if( eoi_orig_pic_dimensions_flag ) {eoi_orig_pic_widthu(16)eoi_orig_pic_heightu(16)} elseeoi_spatial_resampling_type_flagu(1)}if( EoiPrivacyProtectionFlag ) {eoi_privacy_protection_method_idcu(16)eoi_privacy_info_typeu(8)}}}

[0096] EOI SEI 메시지는 EOI 취소 플래그(eoi_cancel_flag)를 포함할 수 있다.

[0097] eoi_cancel_flag may be related to the persistence of encoder optimization information. If the value of eoi_cancel_flag is 1, it may indicate that the persistence of EOI SEI messages included in the previous PU (Picture Unit) in the output order is canceled. If the value of eoi_cancel_flag is 0, it may indicate that optimization-related information applied during pre-processing or encoding follows.

[0098] EOI SEI messages may include an EOI persistence flag (eoi_persistence_flag).

[0099] eoi_persistence_flag may indicate the persistence of optimization information provided in the corresponding SEI message. If the value of eoi_persistence_flag is 0, it may indicate that the optimization information is applied only to the current picture. If the value of eoi_persistence_flag is 1, it may indicate that the optimization information is applied to the current picture and all subsequent pictures of the current layer in output order until any of the following conditions become True.

[0100] - A new CLVS (Coded Layer Video Sequence) of the current layer starts.

[0101] - The bitstream ends.

[0102] - The picture of the current layer, which is located after the current picture in the output order and is associated with the EOI SEI message, is output.

[0103] The EOI SEI message may include an EOI identifier for human viewing (eoi_for_human_viewing_idc).

[0104] If the value of eoi_for_human_viewing_idc is 3, it may indicate that human viewing is included in the purpose of the applied optimization. If the value of eoi_for_human_viewing_idc is 2, it may indicate that the video is suitable for human viewing but is not specifically optimized. If the value of eoi_for_human_viewing_idc is 1, it may indicate that the video is unsuitable for human viewing. If the value of eoi_for_human_viewing_idc is 0, it may indicate that it is unknown whether the video is suitable for human viewing.

[0105] The EOI SEI message may include an EOI indicator (eoi_for_machine_analysis_idc) for machine analysis.

[0106] If the value of eoi_for_machine_analysis_idc is 3, it may indicate that the purpose of the applied optimization includes machine analysis. If the value of eoi_for_machine_analysis_idc is 2, it may indicate that the video is suitable for machine analysis but is not specifically optimized. If the value of eoi_for_machine_analysis_idc is 1, it may mean that the video is unsuitable for machine analysis. If the value of eoi_for_machine_analysis_idc is 0, it may mean that it is unknown whether the video is suitable for machine analysis.

[0107] The values ​​of eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc can be restricted so that they do not both become 1.

[0108] The EOI SEI message may include EOI type information (eoi_type) indicating the type of optimization method.

[0109] For example, eoi_type can be defined as shown in Table 2 below.

[0110] bitMaskInterpretation0x01Object-based optimization; the pictures for which this SEI message persists have been pre-processed or encoded so that detected objects in the pictures are optimized with respect to other parts of the pictures for the indicated optimization purposes0x02Temporal resampling optimization0x04Spatial resampling optimization0x08Temporal quality optimization in a manner that quality fluctuates temporally0x10Spatial quality optimization; the pictures for which this SEI message persists have been pre-processed or encoded to reduce unnecessary information or improve the quality of necessary information.(e.g to reduce the amount of noise and remove speckles at the picture-level)0x20Privacy protection optimization; the pictures for which this SEI message persists have been pre-processed or encoded to protect personal information. (e.g. removal or replacing of personal identifiable information, pseudonymization, anonymization)

[0111] If (eoi_type & bitMask) is not 0, it may indicate that an optimization type with the bitMask value in Table 1 has been applied. If eoi_type is greater than 0 and (eoi_type & bitMask) is 0, it may indicate that an optimization type with the corresponding bitMask value has not been applied. If eoi_type is 0, it may indicate that an optimization determined by the application has been used.

[0112] The variable EoiObjectBasedFlag, which indicates whether object-based optimization is included in the type of optimization represented by eoi_type, can be derived as shown in the following mathematical formula 1.

[0113] [Mathematical Formula 1]

[0114] EoiObjectBasedFlag = ( ( eoi_type & 0x01 ) > 0 ) ? 1:0

[0115] The variable EoiTemporalResamplingFlag, which indicates whether temporal resampling optimization is included in the type of optimization represented by eoi_type, can be derived as shown in the following mathematical equation 2.

[0116] [Mathematical Formula 2]

[0117] EoiTemporalResamplingFlag = ( ( eoi_type & 0x02 ) > 0 ) ? 1:0

[0118] The variable EoiSpatialResamplingFlag, which indicates whether spatial resampling optimization is included in the type of optimization represented by eoi_type, can be derived as shown in the following mathematical equation 3.

[0119] [Mathematical Formula 3]

[0120] EoiSpatialResamplingFlag = ( ( eoi_type & 0x04 ) > 0 ) ? 1:0

[0121] The variable EoiTemporalQualityFlag, which indicates whether temporal quality optimization is included in the type of optimization represented by eoi_type, can be derived as shown in the following mathematical equation 4.

[0122] [Mathematical Formula 4]

[0123] EoiSpatialResamplingFlag = ( ( eoi_type & 0x04 ) > 0 ) ? 1:0

[0124] The variable EoiSpatialQualityFlag, which indicates whether spatial quality optimization is included in the type of optimization represented by eoi_type, can be derived as shown in the following mathematical equation 5.

[0125] [Mathematical Formula 5]

[0126] EoiSpatialQualityFlag = ( ( eoi_type & 0x10 ) > 0 ) ? 1:0

[0127] The variable EoiPrivacyProtectionFlag, which indicates whether privacy protection optimization is included in the type of optimization represented by eoi_type, can be derived as shown in the following mathematical equation 6.

[0128] [Mathematical Formula 6]

[0129] EoiPrivacyProtectionFlag = ( ( eoi_type & 0x20 ) > 0 ) ? 1:0

[0130] For example, if certain top-level temporal sublayers are encoded with very coarse quantization so that human viewers find the quality degradation unpleasant but machine work performance is not degraded, eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc can be set to 0 and 1, respectively, and the value of eoi_type can be set so that EoiTemporalQualityFlag becomes 1.

[0131] If the value of eoi_persistence_flag is 0, EoiTemporalResamplingFlag and EoiTemporalQualityFlag may be restricted to a value of 0.

[0132] The EOI SEI message may include an object identifier (eoi_object_based_idc).

[0133] eoi_object_based_idc may represent the type of object-based optimization defined in Table 3. Here, if (eoi_object_based_idc & bitMask) is not 0, it may indicate that the type of object-based optimization associated with the bitMask value has been applied. If eoi_object_based_idc is greater than 0 and (eoi_object_based_idc & bitMask) is 0, it may indicate that the type of object-based optimization associated with the bitMask value has not been applied. If eoi_object_based_idc is 0, it may indicate that an object-based optimization of the type defined by the application has been applied. The value of eoi_object_based_idc may be restricted to the range of 0 to 31.

[0134] eoi_object_based_idc는 EoiObjectBasedFlag의 값이 1인 경우에 시그날링될 수 있고, 그렇지 않은 경우에 시그날링되지 않을 수 있다.

[0135] bitMaskInterpretation0x01Areas outside the detected objects have been blurred prior to encoding.0x02Areas outside the detected objects have been encoded with coarser transform-domain quantization than the quantization used for the detected objects.0x04Areas outside the detected objects have been overwritten with a constant sample value.0x08Areas outside the detected objects have been overwritten in some form but not with a constant sample value.0x10Areas in the objects have been treated differently based on the object size. For example, objects are pre-sorted in size and larger objects are coded with coarser quality than smaller objects during encoding.

[0136] EOI SEI 메시지는 양자화 임계값 정보(eoi_quant_threshold_delta)를 포함할 수 있다.

[0137] If the value of eoi_quant_threshold_delta is 0, this may indicate that the difference in quantization parameters (QP) between the region outside the detected object and the region containing one or more objects is unknown or unspecified. An eoi_quant_threshold_delta greater than 0 may be used to represent a quantization parameter threshold that determines the region classified outside the object or the region containing one or more objects according to the value of eoi_pic_quant_object_flag described below.

[0138] eoi_quant_threshold_delta may be signaled when the value of EoiObjectBasedFlag is 1, and may not be signaled otherwise. eoi_quant_threshold_delta may be signaled when the value of (eoi_object_based_idc & 0x02) is 1, and may not be signaled otherwise.

[0139] The EOI SEI message may include a quantization object flag (eoi_pic_quant_object_flag).

[0140] If the value of eoi_pic_quant_object_flag is 1, it may indicate that the region coded by the quantization parameter value (PicQuant) is a region containing one or more detected objects. If the value of eoi_pic_quant_object_flag is 0, it may indicate that the region coded by PicQuant is a region outside of the detected objects.

[0141] When eoi_pic_quant_object_flag is 1 and eoi_quant_threshold_delta is greater than 0, regions with quantization parameters greater than or equal to (PicQuant+eoi_quant_threshold_delta) may represent regions outside of detected objects. On the other hand, when eoi_pic_quant_object_flag is 0 and eoi_quant_threshold_delta is greater than 0, regions with quantization parameters less than or equal to (PicQuant-eoi_quant_threshold_delta) may represent regions containing one or more detected objects.

[0142] eoi_pic_quant_object_flag may be signaled when the value of EoiObjectBasedFlag is 1, and may not be signaled otherwise. eoi_pic_quant_object_flag may be signaled when the value of (eoi_object_based_idc & 0x02) is 1, and may not be signaled otherwise. eoi_pic_quant_object_flag may be signaled when the value of eoi_quant_threshold_delta is not 0, and may not be signaled otherwise.

[0143] EOI SEI messages may include a temporal resampling flag (eoi_temporal_resampling_type_flag).

[0144] If the value of eoi_temporal_resampling_type_flag is 0, it may indicate that the temporal resampling optimization is a subsampling operation. If the value of eoi_temporal_resampling_type_flag is 1, it may indicate that the temporal resampling optimization is an upsampling operation.

[0145] eoi_temporal_resampling_type_flag may be signaled when the value of EoiTemporalResamplingFlag is 1, and may not be signaled otherwise.

[0146] The EOI SEI message may include picture count information (eoi_num_int_pics).

[0147] If the value of eoi_num_int_pics is greater than 0, this may indicate that within the persistence of the SEI message, the number of pictures that the encoding system excludes between each pair of coded pictures in the output order (when eoi_temporal_resampling_type_flag is 0) or adds between each pair of original pictures for encoding (when eoi_temporal_resampling_type_flag is 1) is constant.

[0148] If eoi_temporal_resampling_type_flag is 0 and eoi_num_int_pics is greater than 0, eoi_num_int_pics may represent the number of pictures excluded by the encoding system between each pair of coded pictures in the output order. If eoi_temporal_resampling_type_flag is 1 and eoi_num_int_pics is greater than 0, eoi_num_int_pics may represent the number of pictures added by the encoding system between each pair of source pictures for encoding. If the value of eoi_num_int_pics is 0, this may indicate that the number of pictures excluded or added by the encoding system within the persistence of the corresponding SEI message is unknown or variable. The value of eoi_num_int_pics may be restricted to the range of 0 to 63.

[0149] eoi_num_int_pics may be signaled when the value of EoiTemporalResamplingFlag is 1, and may not be signaled otherwise.

[0150] The EOI SEI message may include a source picture flag (eoi_src_pic_flag).

[0151] If the value of eoi_src_pic_flag is 1, it may indicate that a picture within the same access unit containing the EOI SEI message is the source picture for temporal upsampling optimization. If the value of eoi_src_pic_flag is 0, it may indicate that no such indication is provided.

[0152] eoi_src_pic_flag may be signaled when the value of EoiTemporalResamplingFlag is 1, and may not be signaled otherwise. eoi_src_pic_flag may be signaled when the value of eoi_temporal_resampling_type_flag is 1 and the value of eoi_num_int_pics is greater than 0, and may not be signaled otherwise.

[0153] EOI SEI messages may include a picture dimension flag (eoi_orig_pic_dimensions_flag).

[0154] If the value of eoi_orig_pic_dimensions_flag is 1, it may indicate that picture size information (i.e., eoi_orig_pic_width and eoi_orig_pic_height) described later exists. If the value of eoi_orig_pic_dimensions_flag is 0, it may indicate that eoi_orig_pic_width and eoi_orig_pic_height do not exist.

[0155] eoi_orig_pic_dimensions_flag may be signaled when the value of EoiSpatialResamplingFlag is 1, and may not be signaled otherwise.

[0156] The EOI SEI message may include picture size information (eoi_orig_pic_width, eoi_orig_pic_height). eoi_orig_pic_width and eoi_orig_pic_height may represent the width and height of the original source picture, respectively.

[0157] Picture size information may be signaled when the value of EoiSpatialResamplingFlag is 1, and may not be signaled otherwise. Picture size information may be signaled when the value of eoi_orig_pic_dimensions_flag is 1, and may not be signaled otherwise.

[0158] EOI SEI messages may include a spatial resampling flag (eoi_spatial_resampling_type_flag).

[0159] If the value of eoi_spatial_resampling_type_flag is 0, it may indicate that the spatial resampling optimization is a subsampling operation. If the value of eoi_spatial_resampling_type_flag is 1, it may indicate that the spatial resampling optimization is an upsampling operation.

[0160] eoi_spatial_resampling_type_flag may be signaled when the value of EoiSpatialResamplingFlag is 1, and may not be signaled otherwise. eoi_spatial_resampling_type_flag may be signaled when the value of eoi_orig_pic_dimensions_flag is 0, and may not be signaled otherwise.

[0161] The EOI SEI message may include a privacy protection indicator (eoi_privacy_protection_method_idc).

[0162] eoi_privacy_protection_method_idc may represent a method or algorithm used to apply privacy protection optimization. eoi_privacy_protection_method_idc may be defined as shown in Table 4. If eoi_privacy_protection_method_idc is greater than 0 and (eoi_privacy_protection_method_idc & bitMask) is not 0, this may indicate that the method / algorithm associated with the bitMask value was used for privacy protection. If the value of eoi_privacy_protection_method_idc is 0, this may indicate that the method / algorithm used for privacy protection is unknown or determined by the application. The value of eoi_privacy_protection_method_idc may be restricted to a range of 0 to 15.

[0163] eoi_privacy_protection_method_idc may be signaled when the value of EoiPrivacyProtectionFlag is 1, and may not be signaled otherwise.

[0164] bitMaskInterpretation0x01Blurring; personal information is blurred to make it unidentifiable.0x02Replacing; personal information is replaced with something different from the original to make it unidentifiable.0x04Masking; personal information is masked so that it cannot be identified0x08Pixelation; personal information is pixelated to make it undiscernible

[0165] The EOI SEI message may include privacy type information (eoi_privacy_info_type).

[0166] eoi_privacy_info_type can indicate the type of protected information. eoi_privacy_info_type can be defined as shown in Table 5. Here, if eoi_privacy_info_type is greater than 0 and ( eoi_privacy_info_type & bitMask ) is not 0, this may indicate that the type of information associated with the bitMask value is protected. If eoi_privacy_info_type is 0, this may indicate that information of a type defined by the application is protected. The value of eoi_privacy_info_type may be restricted to a range of 0 to 7.

[0167] eoi_privacy_info_type may be signaled when the value of EoiPrivacyProtectionFlag is 1, and may not be signaled otherwise.

[0168] bitMaskInterpretation0x01Information that identifies a person is protected. For example, the face of the person.0x02Information that can identify vehicles is protected. For example, the license plate of the vehicle.0x04Information that can infer locations is protected. For example text or images on signs.

[0169] For the interpretation of an EOI SEI message, the quantization parameter value (PicQuant) for each picture to which the EOI SEI message is applied can be derived based on a flag (pps_qp_delta_info_in_ph_flag).

[0170] pps_qp_delta_info_in_ph_flag may indicate whether QP (quantization parameter) delta information exists in the picture header. For example, if the value of pps_qp_delta_info_in_ph_flag is 1, it may indicate that QP delta information exists in the picture header but not in the slice header. If the value of pps_qp_delta_info_in_ph_flag is 0, it may indicate that QP delta information does not exist in the picture header but exists in the slice header. If pps_qp_delta_info_in_ph_flag does not exist, the value of pps_qp_delta_info_in_ph_flag can be inferred to be 0.

[0171] If the value of pps_qp_delta_info_in_ph_flag is 1, PicQuant can be derived to a value of (26 + pps_init_qp_minus26 + ph_qp_delta). Otherwise (i.e., if the value of pps_qp_delta_info_in_ph_flag is 0), PicQuant can be derived to a value of (26 + pps_init_qp_minus26 + sh_qp_delta). Here, sh_qp_delta may be for the first slice among multiple slices belonging to the picture.

[0172] Initial QP information (pps_init_qp_minus26) may represent the initial value of the QP for each slice referencing the PPS. The initial value of the QP may be modified at the picture level based on ph_qp_delta or at the slice level based on sh_qp_delta. Here, the QP delta information (ph_qp_delta) in the picture header may specify the initial value of the QP to be used for coding blocks within the picture. The QP delta information (sh_qp_delta) in the slice header may specify the initial value of the QP to be used for coding blocks within the slice.

[0173] In the VVC standard, when applying EOI SEI messages, PicQuant can be calculated using ph_qp_delta defined in the picture header. However, since the AVC and HEVC standards do not define picture headers, ph_qp_delta cannot be used. Therefore, it is necessary to derive PicQuant using the information defined in each standard.

[0174] To this end, we propose a method to derive PicQuant for applying EOI SEI messages in the HEVC standard.

[0175] For the interpretation of the EOI SEI message, the PicQuant for each picture to which the EOI SEI message is applied can be derived based on at least one of the initial QP information (init_qp_minus26) or the QP delta information (slice_qp_delta). For example, the PicQuant for each picture to which the EOI SEI message is applied can be derived individually as shown in the following Equation 1.

[0176] [Mathematical Formula 1]

[0177] PicQuant = 26 + init_qp_minus26 + slice_qp_delta

[0178] In Equation 1, init_qp_minus26 can specify the initial value of the QP (hereinafter referred to as SliceQp) for each slice referencing the PPS. init_qp_minus26 can be signaled in the picture parameter set (PPS). The initial value of SliceQp can be modified / updated in the slice segment layer based on slice_qp_delta.

[0179] In Equation 1, slice_qp_delta can specify the initial value of the QP to be used for the coding blocks within the slice. slice_qp_delta can be signaled in the slice segment header. A picture may consist of multiple slices, and each slice may consist of one or more slice segments. Multiple slice segments may include at least one of independent slice segments or dependent slice segments. If a picture contains multiple slice segments, slice_qp_delta may be related to any one of the multiple slice segments. For example, slice_qp_delta may be signaled for the first slice segment among the multiple slice segments belonging to a picture.

[0180] In other words, the PicQuant for a picture to which the EOI SEI message is applied can be derived based on init_qp_minus26 and slice_qp_delta for the first slice segment among multiple slice segments belonging to the picture. The first slice segment may refer to a slice segment containing the top-left coding block within the picture. Alternatively, the first slice segment may refer to an independent slice segment containing the top-left coding block within the picture.

[0181] In addition, we propose a method to derive PicQuant for applying EOI SEI messages in the AVC standard.

[0182] For the interpretation of the EOI SEI message, the PicQuant for each picture to which the EOI SEI message is applied can be derived based on at least one of the initial QP information (pic_init_qp_minus26) or the QP delta information (slice_qp_delta). For example, the PicQuant for each picture to which the EOI SEI message is applied can be derived individually as shown in the following Equation 2.

[0183] [Mathematical Formula 2]

[0184] PicQuant = 26 + pic_init_qp_minus26 + slice_qp_delta

[0185] In Equation 2, pic_init_qp_minus26 can specify the initial value of the QP (hereinafter referred to as SliceQp) for each slice. pic_init_qp_minus26 can be signaled in the Picture Parameter Set (PPS). The initial value of SliceQp can be modified / updated in the slice layer based on slice_qp_delta.

[0186] In Equation 2, slice_qp_delta can specify the initial value of the QP to be used for the macroblocks within the slice. slice_qp_delta can be signaled in the slice header. A picture may consist of one or more slices. If a picture contains multiple slices, slice_qp_delta may relate to any one of the multiple slices. For example, slice_qp_delta may be signaled for the first slice among the multiple slices belonging to a picture.

[0187] In other words, PicQuant for a picture to which EOI SEI messages are applied can be derived based on pic_init_qp_minus26 and slice_qp_delta for the first slice among multiple slices belonging to the picture. The first slice may refer to a slice containing the top-left macroblock within the picture.

[0188] The EOI SEI message according to the present disclosure may be included in a network abstraction layer (NAL) unit of the bitstream. Alternatively, the encoder optimization information according to the present disclosure may be configured in a high-level syntax of the bitstream. Here, the high-level syntax may be at least one of a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), or a slice header (SH). Alternatively, the encoder optimization information according to the present disclosure may be defined as a separate NAL unit type within the bitstream.

[0189] FIG. 5 illustrates a schematic configuration of a decoding device (300) that performs a method for restoring a video picture according to the present disclosure.

[0190] Referring to FIG. 5, the decoding device (300) may include a receiving unit (500), a video information extraction unit (510), and a video restoration unit (520).

[0191] The receiver (500) can receive a bitstream including an encoded video picture.

[0192] The video information extraction unit (510) can extract video information regarding an encoded video picture from the bitstream. Additionally, the video information extraction unit (710) can extract encoder optimization information from the bitstream, as seen with reference to FIG. 4.

[0193] The video restoration unit (520) can restore an encoded video picture based on extracted video information.

[0194] FIG. 6 illustrates a method for generating a bitstream performed in an encoding device (200) according to the present disclosure.

[0195] A video picture being encoded can be received (S600).

[0196] Video information regarding the video picture can be generated by encoding the received video picture (S610).

[0197] A bitstream containing video information about a video picture can be generated (S620).

[0198] In addition, encoder optimization information applied to the bitstream can be generated, as seen with reference to FIG. 4. The generated encoder optimization information can be included in the bitstream. In this case, the encoder optimization information may be configured in the SEI message of the bitstream or in the high-level syntax of the bitstream.

[0199] In addition, the quantization parameter value (PicQuant) of the picture for applying the EOI SEI message can be derived based on at least one of the initial QP information or QP delta information, as seen with reference to Fig. 4.

[0200] FIG. 7 illustrates a schematic configuration of an encoding device (200) that performs a method for generating a bitstream according to the present disclosure.

[0201] Referring to FIG. 7, the encoding device (200) may include a receiving unit (700), a video compression unit (710), and a bitstream generation unit (720).

[0202] The receiver (700) can receive one or more video pictures that are encoded.

[0203] The video compression unit (710) can generate video information regarding the video picture by encoding one or more received video pictures. The video compression unit (710) can generate encoder optimization information applied to the bitstream.

[0204] The bitstream generation unit (720) can generate a bitstream including the video information. The bitstream generation unit (720) can generate a bitstream including the generated encoder optimization information.

[0205] In the embodiments described above, methods are described based on flowcharts as a series of steps or blocks; however, the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps as described above. Furthermore, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and other steps may be included, or one or more steps of the flowcharts may be omitted without affecting the scope of the embodiments of this document.

[0206] The method according to the embodiments of the present document described above may be implemented in the form of software, and the encoding device and / or decoding device according to the present document may be included in a device that performs image processing, such as a TV, computer, smartphone, set-top box, display device, etc.

[0207] When the embodiments described in this document are implemented in software, the method described above may be implemented as a module (process, function, etc.) that performs the function described above. The module may be stored in memory and executed by a processor. The memory may be located inside or outside the processor and may be connected to the processor by various well-known means. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this document may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each figure may be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for implementation may be stored on a digital storage medium.

[0208] In addition, the decoding device and encoding device to which the embodiment(s) of the present specification are applied may be included in multimedia broadcasting transmission and reception devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video conversation devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, Video on Demand (VoD) service providers, Over-the-top video (OTT) devices, internet streaming service providers, 3D video devices, virtual reality (VR) devices, augmented reality (AR) devices, video phone video devices, transportation terminals (e.g., vehicle terminals (including autonomous vehicles), airplane terminals, ship terminals, etc.), and medical video devices, and may be used to process video signals or data signals. For example, Over-the-top video (OTT) devices may include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, Digital Video Recorders (DVRs), etc.

[0209] Additionally, the processing method to which the embodiment(s) of this specification are applied may be produced in the form of a program that is executed by a computer and may be stored on a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of this specification may also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium may include, for example, a Blu-ray disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. Additionally, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission over the Internet). Additionally, a bitstream generated by an encoding method may be stored on a computer-readable recording medium or transmitted via a wired or wireless communication network.

[0210] Additionally, the embodiments of this specification may be implemented as a computer program product by program code, and said program code may be executed on a computer by the embodiments of this specification. said program code may be stored on a computer-readable carrier.

[0211] FIG. 8 shows an example of a content streaming system to which embodiments of the present disclosure can be applied.

[0212] Referring to FIG. 8, a content streaming system to which the embodiment(s) of the present specification are applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0213] The above encoding server compresses content input from multimedia input devices, such as smartphones, cameras, and camcorders, into digital data to generate a bitstream and transmits it to the streaming server. As another example, if multimedia input devices, such as smartphones, cameras, and camcorders, generate the bitstream directly, the encoding server may be omitted.

[0214] The bitstream above may be generated by an encoding method or a bitstream generation method to which the embodiment(s) of the present specification are applied, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0215] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as a medium to inform the user of available services. When a user requests a desired service from the web server, the web server transmits it to the streaming server, and the streaming server transmits the multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server plays the role of controlling commands and responses between each device within the content streaming system.

[0216] The streaming server may receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content may be received in real time. In this case, to provide a seamless streaming service, the streaming server may store the bitstream for a certain period of time.

[0217] Examples of the above user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs (head-mounted displays)), digital TVs, desktop computers, digital signage, etc.

[0218] Each server within the above-mentioned content streaming system can be operated as a distributed server, and in this case, data received from each server can be processed in a distributed manner.

[0219] The claims described in this specification may be combined in various ways. For example, the technical features of the method claims in this specification may be combined to be implemented as a device, and the technical features of the device claims in this specification may be combined to be implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims in this specification may be combined to be implemented as a device, and the technical features of the method claims and the technical features of the device claims in this specification may be combined to be implemented as a method.

Claims

1. A step of receiving a bitstream including an encoded video picture; and The method includes the step of restoring an encoded video picture contained in the bitstream, The above bitstream includes Encoder Optimization Information (EOI) and Supplemental Enhancement Information (SEI) messages, and The above EOI SEI message includes EOI type information indicating the type of optimization method, and The above EOI SEI message is obtained from the NAL (Network Abstraction Layer) unit of the bitstream, and A method in which the quantization parameter value of a picture for applying the above EOI SEI message is derived based on initial QP (Quantization Parameter) information and QP delta information.

2. In Paragraph 1, The above initial QP information is a method for specifying the initial value of the QP for each slice that references the picture parameter set.

3. In Paragraph 2, The above initial QP information is signaled from the picture parameter set of the bitstream, a method.

4. The method of claim 1, wherein the QP delta information specifies the initial value of the QP to be used for blocks within a slice.

5. In Paragraph 4, A method in which the above QP delta information relates to the first slice segment among a plurality of slice segments belonging to the above picture.

6. In Paragraph 5, The above QP delta information is signaled in the slice segment header of the bitstream.

7. In Paragraph 4, A method in which the above QP delta information relates to the first slice among a plurality of slices belonging to the above picture.

8. In Paragraph 7, The above QP delta information is signaled in the slice header of the bitstream.

9. A step of receiving the video picture to be encoded; A step of encoding the received video picture to generate video information regarding the video picture; A step of generating Encoder Optimization Information (EOI) and Supplemental Enhancement Information (SEI) messages; and The method includes the step of generating a bitstream including the video information and the EOI SEI message, wherein The above EOI SEI message includes EOI type information indicating the type of optimization method, and The above EOI SEI message is encoded in the NAL (Network Abstraction Layer) unit of the bitstream, and A method in which the quantization parameter value of a picture for applying the above EOI SEI message is derived based on initial QP (Quantization Parameter) information and QP delta information.

10. A non-transient computer-readable storage medium for storing a bitstream generated by the method according to paragraph 9.

11. A step of generating a bitstream; wherein the bitstream is generated based on the steps of receiving a video picture to be encoded, encoding the received video picture to generate video information regarding the video picture, and generating an Encoder Optimization Information (EOI) SEI (Supplemental Enhancement Information) message, and The method includes the step of transmitting data including the above bitstream, The above EOI SEI message includes EOI type information indicating the type of optimization method, and The above EOI SEI message is encoded in the NAL (Network Abstraction Layer) unit of the bitstream, and A method in which the quantization parameter value of a picture for applying the above EOI SEI message is derived based on initial QP (Quantization Parameter) information and QP delta information.