Image encoding / decoding method and device, and recording medium storing bitstream

The method addresses inefficiencies in image compression by signaling encoder optimization information, enabling tailored encoding for human and machine analysis, thereby enhancing encoding efficiency and quality.

WO2025206841A1PCT designated stage Publication Date: 2025-10-02LG ELECTRONICS INC

Patent Information

Application Number
PCT/KR2025/004093
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-28
Filing Date
2025-03-28
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing image compression technologies struggle to efficiently accommodate diverse optimization needs for human viewing and machine analysis, leading to inefficiencies in encoding and decoding processes.

Method used

A method and apparatus for signaling encoder optimization information, including EOI identifiers for human viewing and machine analysis, with specific constraints and optimizations, such as object-based and privacy optimizations, to enhance encoding efficiency.

Benefits of technology

Enhances encoding efficiency by allowing for optimized video/image encoding and decoding processes tailored to specific viewing purposes, improving compression and transmission quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025004093_02102025_PF_FP_ABST
    Figure KR2025004093_02102025_PF_FP_ABST
Patent Text Reader

Abstract

An image encoding method and device according to the present disclosure may receive a video picture to be encoded, encode the received video picture to generate a compressed video picture, generate encoder optimization information, and generate a bitstream including the compressed video picture and the encoder optimization information. Here, the encoder optimization information may include at least one of an EOI identifier for human viewing and an EOI identifier for machine analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Video encoding / decoding method and device, and recording medium storing bitstream

[0001] The present invention relates to a video encoding / decoding method and device, and a recording medium storing a bitstream.

[0002] Recently, the demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, is increasing in various application fields, and accordingly, high-efficiency image compression technologies are being discussed.

[0003] There are various technologies for image compression, such as inter prediction technology that predicts pixel values ​​included in the current picture from pictures before or after the current picture, intra prediction technology that predicts pixel values ​​included in the current picture using pixel information within the current picture, and entropy encoding technology that assigns short codes to values ​​with high frequency of appearance and long codes to values ​​with low frequency of appearance, and these technologies can be used to effectively compress and transmit or store image data.

[0004] The present disclosure provides a method and apparatus for configuring encoder optimization information.

[0005] The present disclosure provides a method and apparatus for signaling encoder optimization information.

[0006] A video encoding method and device according to the present disclosure can receive a video picture to be encoded, encode the received video picture to generate a compressed video picture, generate encoder optimization information, and generate a bitstream including the compressed video picture and the encoder optimization information. The encoder optimization information can be encoded in a network abstraction layer (NAL) unit of the bitstream. The encoder optimization information can include at least one of an EOI identifier for human viewing or an EOI identifier for machine analysis.

[0007] In the image encoding method and device according to the present disclosure, both the EOI identifier for human viewing and the EOI identifier for machine analysis can have any one value within the range of 0 to 3.

[0008] In the video encoding method and device according to the present disclosure, a constraint may be applied that both the EOI identifier for human viewing and the EOI identifier for machine analysis must not be 1.

[0009] In the video encoding method and device according to the present disclosure, if the value of the EOI identifier for human viewing is 1, this may mean that the video is unsuitable for the human viewing, and if the value of the EOI identifier for machine analysis is 1, this may mean that the video is unsuitable for the machine analysis.

[0010] In the video encoding method and device according to the present disclosure, the number of values ​​that can be set for the other of the EOI identifier for human viewing and the EOI identifier for machine analysis may be determined differently based on the value of one of the EOI identifier for human viewing and the EOI identifier for machine analysis.

[0011] In the video encoding method and device according to the present disclosure, a constraint may be applied that both the EOI identifier for human viewing and the EOI identifier for machine analysis must not be 0.

[0012] In the video encoding method and device according to the present disclosure, a constraint may be applied that the value of at least one of the EOI identifier for human viewing and the EOI identifier for machine analysis must be 3.

[0013] In the video encoding method and device according to the present disclosure, the encoder optimization information may include object-based optimization-related information. Here, whether the object-based optimization-related information is encoded may be determined based on the value of the EOI identifier for human viewing.

[0014] In the video encoding method and device according to the present disclosure, the encoder optimization information may include privacy optimization-related information. Here, whether the privacy optimization-related information is encoded may be determined based on the value of the EOI identifier for human viewing.

[0015] In the video encoding method and device according to the present disclosure, the encoder optimization information may include bit depth optimization-related information, and the bit depth optimization-related information may include information indicating the type of bit depth optimization. Here, a constraint may be applied that the value of the information indicating the type of bit depth optimization must not be 0 based on the value of the EOI identifier for human viewing.

[0016] A video decoding method and device according to the present disclosure may include the steps of receiving a bitstream including an encoded video picture and restoring the encoded video picture included in the bitstream. Here, the bitstream includes encoder optimization information, and the encoder optimization information may be obtained from a network abstraction layer (NAL) unit of the bitstream. The encoder optimization information may include at least one of an EOI identifier for human viewing or an EOI identifier for machine analysis.

[0017] A computer-readable digital storage medium is provided, which stores encoded video / image information that causes a decoding device according to the present disclosure to perform a video decoding method.

[0018] A computer-readable digital storage medium storing video / image information generated by a video encoding method according to the present disclosure is provided.

[0019] A method and device for transmitting video / image information generated by a video encoding method according to the present disclosure are provided.

[0020] According to the present disclosure, encoder optimization information can be signaled more efficiently by encoding optimization-related information while taking into account constraints for defining a clear optimization target.

[0021] According to the present disclosure, encoder optimization information can be signaled more efficiently by encoding detailed information about optimization properties in consideration of the optimization purpose.

[0022] FIG. 1 illustrates a video / image coding system according to the present disclosure.

[0023] FIG. 2 is a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and in which encoding of a video / image signal is performed.

[0024] FIG. 3 is a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and in which decoding of a video / image signal is performed.

[0025] FIG. 4 illustrates a method for generating a bitstream performed in an encoding device (200) according to the present disclosure.

[0026] FIG. 5 illustrates a schematic configuration of an encoding device (200) that performs a method for generating a bitstream according to the present disclosure.

[0027] FIG. 6 illustrates a method for restoring a video picture performed in a decoding device (300) according to the present disclosure.

[0028] FIG. 7 illustrates a schematic configuration of a decoding device (300) that performs a method for restoring a video picture according to the present disclosure.

[0029] FIG. 8 illustrates an example of a content streaming system to which embodiments of the present disclosure can be applied.

[0030] The present disclosure may be modified in various ways and encompasses numerous embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present disclosure. Throughout the description of each drawing, similar reference numerals have been used to designate similar components.

[0031] While terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present disclosure, a first component could be referred to as a "second component," and similarly, a second component could also be referred to as a "first component." The term "and / or" includes a combination of multiple related items described herein or any of multiple related items described herein.

[0032] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.

[0033] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present disclosure. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0034] The present disclosure relates to video / image coding. For example, the methods / embodiments disclosed in this specification can be applied to methods disclosed in the versatile video coding (VVC) standard. In addition, the methods / embodiments disclosed in this specification can be applied to methods disclosed in the essential video coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation of audio video coding standard (AVS2), or the next generation of video / image coding standards (e.g., H.267 or H.268).

[0035] This specification presents various embodiments of video / image coding, and unless otherwise stated, the embodiments may be performed in combination with each other.

[0036] In this specification, a video may refer to a set of images over time. A picture generally refers to a unit representing one image at a specific time point, and a slice / tile is a unit that constitutes part of a picture in coding. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A tile is a rectangular area consisting of multiple CTUs within a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of ​​CTUs that has a height equal to the height of the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of ​​CTUs that has a height specified by the picture parameter set and a width equal to the width of the picture. CTUs within a tile are arranged consecutively according to the CTU raster scan, while tiles within a picture may be arranged consecutively according to the tile raster scan. A slice may contain an integer number of complete tiles or an integer number of contiguous complete CTU rows within a picture, which may be exclusively contained within a single NAL unit. Meanwhile, a picture may be divided into two or more subpictures. A subpicture may be a rectangular region of one or more slices within a picture.

[0037] A pixel, or pel, can refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component.

[0038] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.

[0039] As used herein, "A or B" can mean "only A," "only B," or "both A and B." In other words, as used herein, "A or B" can be interpreted as "A and / or B." For example, as used herein, "A, B or C" can mean "only A," "only B," "only C," or "any combination of A, B and C."

[0040] As used herein, a slash ( / ) or a comma can mean "and / or." For example, "A / B" can mean "A and / or B." Accordingly, "A / B" can mean "only A," "only B," or "both A and B." For example, "A, B, C" can mean "A, B, or C."

[0041] In this specification, "at least one of A and B" may mean "only A", "only B" or "both A and B". Additionally, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted identically to "at least one of A and B".

[0042] Additionally, in this specification, “at least one of A, B and C” can mean “only A,” “only B,” “only C,” or “any combination of A, B and C.” Additionally, “at least one of A, B or C” or “at least one of A, B and / or C” can mean “at least one of A, B and C.”

[0043] Additionally, parentheses used herein may mean "for example." Specifically, when "prediction (intra-prediction)" is indicated, "intra-prediction" may be suggested as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra-prediction," and "intra-prediction" may be suggested as an example of "prediction." Furthermore, even when "prediction (i.e., intra-prediction)" is indicated, "intra-prediction" may be suggested as an example of "prediction."

[0044] Technical features individually described in a single drawing in this specification may be implemented individually or simultaneously.

[0045] FIG. 1 illustrates a video / image coding system according to the present disclosure.

[0046] Referring to FIG. 1, a video / image coding system may include a first device (source device) and a second device (receiving device).

[0047] A source device can transmit encoded video / image information or data to a receiving device via a digital storage medium or a network in the form of a file or streaming. The source device may include a video source, an encoding device, and a transmitting device. The receiving device may include a receiving device, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be configured as a separate device or an external component.

[0048] A video source may obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include a computer, a tablet, a smartphone, etc., and may (electronically) generate video / images. For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process of generating related data.

[0049] An encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, to improve compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0050] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit can include an element for generating a media file via a predetermined file format and an element for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.

[0051] The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device.

[0052] The renderer can render decoded video / images. The rendered video / images can be displayed through the display unit.

[0053] FIG. 2 is a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and in which encoding of a video / image signal is performed.

[0054] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a prediction unit (predictor) 220, a residual processor (residual processor) 230, an entropy encoder (entropy encoder) 240, an adder (adder) 250, a filter (filter) 260, and a memory (memory) 270. The prediction unit (220) may include an inter prediction unit (221) and an intra prediction unit (222). The residual processor (230) may include a transformer (transformer) 232, a quantizer (quantizer) 233, a dequantizer (dequantizer) 234, and an inverse transformer (inverse transformer) 235. The residual processing unit (230) may further include a subtractor (231). The addition unit (250) may be called a reconstructor or a recontructed block generator. The image segmentation unit (210), the prediction unit (220), the residual processing unit (230), the entropy encoding unit (240), the addition unit (250), and the filtering unit (260) described above may be configured by one or more hardware components (e.g., an encoding device chipset or processor) according to an embodiment. In addition, the memory (270) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.

[0055] The image segmentation unit (210) can segment an input image (or picture, frame) input to the encoding device (200) into one or more processing units. For example, the processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively segmented from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure.

[0056] For example, a single coding unit may be split into multiple coding units with deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied before the quad-tree structure. The coding procedure according to the present specification may be performed based on the final coding unit that is no longer split. In this case, based on coding efficiency according to image characteristics, etc., the largest coding unit may be used directly as the final coding unit, or, if necessary, the coding unit may be recursively split into coding units of lower depths, and the coding unit with the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later.

[0057] As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be split or partitioned from the final coding unit described above. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from a transform coefficient.

[0058] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the case. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. A sample can be used as a term corresponding to a pixel or pel in a picture (or image).

[0059] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (prediction block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input video signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, a unit that subtracts a prediction signal (prediction block, prediction sample array) from an input video signal (original block, original sample array) within the encoding device (200) may be called a subtraction unit (231).

[0060] The prediction unit (220) can perform a prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit (220) can determine whether intra prediction or inter prediction is applied on a current block or CU basis. The prediction unit (220) can generate various information related to prediction, such as prediction mode information, as described later in the description of each prediction mode, and transmit the information to the entropy encoding unit (240). The information related to prediction can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.

[0061] The intra prediction unit (222) can predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block, or may be located a certain distance away from the current block, depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The non-directional mode may include at least one of a DC mode or a planar mode. The directional mode may include 33 directional modes or 65 directional modes depending on the degree of detail in the prediction direction. However, this is only an example, and a greater or lesser number of directional modes may be used depending on the settings. The intra prediction unit (222) may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.

[0062] The inter prediction unit (221) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The above temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and the reference pictures including the temporal neighboring blocks may be called collocated pictures (colPic). For example, the inter prediction unit (221) may construct a motion information candidate list based on the neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of the neighboring blocks as the motion information of the current block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0063] The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP) mode. In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described herein. Palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, sample values ​​within a picture can be signaled based on information about the palette table and palette index. The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or a residual signal.

[0064] The transform unit (232) can apply a transform technique to the residual signal to generate transform coefficients. For example, the transform technique can include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is expressed as a graph. CNT refers to a transform obtained based on generating a prediction signal using all previously restored pixels. In addition, the transform process can be applied to a pixel block having a square size and the same size, or can be applied to a block of a non-square variable size.

[0065] The quantization unit (233) quantizes the transform coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients can be called residual information. The quantization unit (233) can rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.

[0066] The entropy encoding unit (240) can perform various encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit (240) can also encode information necessary for video / image restoration (e.g., values ​​of syntax elements, etc.) together or separately from quantized transform coefficients.

[0067] Encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. In the present specification, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the above-described encoding procedure and included in the bitstream. The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media, such as a USB, SD, CD, DVD, Blu-ray, HDD, or SSD. The signal output from the entropy encoding unit (240) may be configured as an internal / external element of the encoding device (200) by a transmitting unit (not shown) and / or a storing unit (not shown), or the transmitting unit may be included in the entropy encoding unit (240).

[0068] The quantized transform coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (234) and the inverse transform unit (235), a residual signal (residual block or residual samples) can be restored. The addition unit (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit (221) or the intra prediction unit (222). When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as a reconstructed block. The addition unit (250) may be called a restoration unit or a reconstructed block generation unit. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, and can also be used for inter prediction of the next picture after filtering as described below. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.

[0069] The filtering unit (260) can improve subjective / objective picture quality by applying filtering to the restoration signal. For example, the filtering unit (260) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and store the modified restoration picture in the memory (270), specifically, in the DPB of the memory (270). The various filtering methods can include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (240). The information regarding filtering can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.

[0070] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter prediction unit (221). Through this, when inter prediction is applied, the encoding device can avoid prediction mismatch between the encoding device (200) and the decoding device, and can also improve encoding efficiency.

[0071] The DPB of the memory (270) can store the modified restored picture to be used as a reference picture in the inter prediction unit (221). The memory (270) can store motion information of a block from which motion information is derived (or encoded) within the current picture and / or motion information of blocks within a picture that has already been restored. The stored motion information can be transferred to the inter prediction unit (221) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (270) can store restored samples of restored blocks within the current picture and transfer them to the intra prediction unit (222).

[0072] FIG. 3 is a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and in which decoding of a video / image signal is performed.

[0073] Referring to FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-prediction unit (332) and an intra-prediction unit (331). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321).

[0074] The entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) described above may be configured by a single hardware component (e.g., a decoding device chipset or processor) depending on the embodiment. In addition, the memory (360) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.

[0075] When a bitstream including video / image information is input, the decoding device (300) can restore the image corresponding to the process in which the video / image information is processed in the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied in the encoding device. Accordingly, the processing unit of decoding may be a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. Then, the restored image signal decoded and output through the decoding device (300) can be reproduced through a reproduction device.

[0076] The decoding device (300) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through the entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The decoding device can decode the picture further based on the information on the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this specification can be decoded through the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit (310) can decode information in a bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values ​​of syntax elements required for image restoration and the quantized values ​​of transform coefficients for residuals. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element in the bitstream, determines a context model using information of the syntax element to be decoded and decoding information of the surrounding and decoding target blocks or information of symbols / bins decoded in the previous step, and predicts the occurrence probability of the bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Information regarding prediction among the information decoded by the entropy decoding unit (310) is provided to the prediction unit (inter prediction unit (332) and intra prediction unit (331)), and residual values ​​on which entropy decoding is performed by the entropy decoding unit (310), i.e., quantized transform coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive a residual signal (residual block, residual samples, residual sample array). In addition, information regarding filtering among the information decoded by the entropy decoding unit (310) can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of a decoding device (300), or the receiving unit may be a component of an entropy decoding unit (310).

[0077] Meanwhile, a decoding device according to the present specification may be called a video / video / picture decoding device, and the decoding device may be divided into an information decoding device (video / video / picture information decoding device) and a sample decoding device (video / video / picture sample decoding device). The information decoding device may include the entropy decoding unit (310), and the sample decoding device may include at least one of the inverse quantization unit (321), the inverse transformation unit (322), the addition unit (340), the filtering unit (350), the memory (360), the inter prediction unit (332), and the intra prediction unit (331).

[0078] The inverse quantization unit (321) can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit (321) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.

[0079] In the inverse transform unit (322), the transform coefficients are inversely transformed to obtain a residual signal (residual block, residual sample array).

[0080] The prediction unit (320) can perform a prediction on the current block and generate a predicted block including prediction samples for the current block. The prediction unit (320) can determine whether intra-prediction or inter-prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit (310), and can determine a specific intra / inter-prediction mode.

[0081] The prediction unit (320) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit (320) can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP) mode. In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described herein. Palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, information about the palette table and palette index may be included and signaled in the video / image information.

[0082] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block, or may be located a certain distance away from the current block, depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The intra prediction unit (331) may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.

[0083] The inter prediction unit (332) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) can construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and information about the prediction can include information indicating an inter prediction mode for the current block.

[0084] The addition unit (340) can generate a restoration signal (restored picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter-prediction unit (332) and / or intra-prediction unit (331)). When there is no residual for the block to be processed, such as when skip mode is applied, the prediction block can be used as the restoration block.

[0085] The addition unit (340) may be referred to as a restoration unit or restoration block generation unit. The generated restoration signal may be used for intra prediction of the next processing target block within the current picture, may be output after filtering as described below, or may be used for inter prediction of the next picture. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.

[0086] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can apply various filtering methods to the restored picture to generate a modified restored picture, and transmit the modified restored picture to the memory (360), specifically, to the DPB of the memory (360). The various filtering methods can include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0087] The (corrected) reconstructed picture stored in the DPB of the memory (360) can be used as a reference picture in the inter prediction unit (332). The memory (360) can store motion information of a block from which motion information is derived (or decoded) in the current picture and / or motion information of blocks in a picture that has already been reconstructed. The stored motion information can be transferred to the inter prediction unit (332) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (360) can store reconstructed samples of reconstructed blocks in the current picture and transfer them to the intra prediction unit (331).

[0088] In this specification, the embodiments described in the filtering unit (260), the inter prediction unit (221), and the intra prediction unit (222) of the encoding device (200) can be applied to the filtering unit (350), the inter prediction unit (332), and the intra prediction unit (331) of the decoding device (300) in the same or corresponding manner, respectively.

[0089] FIG. 4 illustrates a method for generating a bitstream performed in an encoding device (200) according to the present disclosure.

[0090] A video picture to be encoded can be received (S400).

[0091] A received video picture can be encoded to generate a compressed video picture (S410).

[0092] Encoder optimization information (EOI) can be generated (S420).

[0093] Encoder optimization information according to the present disclosure may relate to a received or compressed video picture. The encoder optimization information may define the purpose, properties, and scope (or target) of encoder optimization.

[0094] For example, the encoder optimization information may include an EOI cancellation flag (eoi_cancel_flag). eoi_cancel_flag may relate to the persistence of the encoder optimization information. If the value of eoi_cancel_flag is 1, this may indicate that the persistence of the previously applied encoder optimization information is canceled. For example, if the value of eoi_cancel_flag is 1, this may indicate that the persistence of the encoder optimization information included in the previous PU (prediction unit) in the output order is canceled. If the value of eoi_cancel_flag is 0, this may indicate that the optimization-related information applied during pre-processing or encoding follows.

[0095] The above optimization-related information may include at least one of an EOI persistence flag (eoi_persistence_flag), an EOI identifier for human viewing (eoi_for_human_viewing_idc), an EOI identifier for machine analysis (eoi_for_machine_analysis_idc), or EOI type information (eoi_type). That is, based on the value of eoi_cancel_flag being 0, the above-described optimization-related information may be generated and encoded in the bitstream.

[0096] eoi_persistence_flag may be related to the persistence of optimization information. eoi_persistence_flag may indicate the persistence of the optimization information indicated by eoi_type. If the value of eoi_persistence_flag is 0, this may indicate that the optimization information (identified based on eoi_type) is applied only to the current picture. If the value of eoi_persistence_flag is 1, this may indicate that the optimization information (identified by eoi_type) is applied to the current picture and all subsequent pictures. Here, the subsequent pictures may mean all subsequent pictures of the current layer in output order.

[0097] eoi_for_human_viewing_idc can indicate information for identifying the purpose and target of optimization. A value of 3 for eoi_for_human_viewing_idc may indicate that the purpose of optimization includes human viewing. A value of 2 for eoi_for_human_viewing_idc may indicate that the video is suitable for human viewing but is not specifically optimized for human viewing. A value of 1 for eoi_for_human_viewing_idc may indicate that the video is not suitable for human viewing. A value of 0 for eoi_for_human_viewing_idc may indicate that it is unknown whether the video is suitable for human viewing.

[0098] eoi_for_machine_analysis_idc can indicate information for identifying the purpose and target of optimization. If the value of eoi_for_machine_analysis_idc is 3, this may mean that the purpose of optimization includes machine analysis. If the value of eoi_for_machine_analysis_idc is 2, this may mean that the video is suitable for machine analysis but is not specifically optimized for machine analysis. If the value of eoi_for_machine_analysis_idc is 1, this may mean that the video is not suitable for machine analysis. If the value of eoi_for_machine_analysis_idc is 0, this may mean that it is not known whether the video is suitable for machine analysis.

[0099] eoi_type can indicate the properties of an optimization method. For example, eoi_type can be defined as shown in Table 1 below. eoi_type can identify the optimization properties (or optimization types) defined in Table 1. If the value of eoi_type is 0, this may mean that optimization by the application or user is applied. However, the optimization properties defined in Table 1 are only examples, and new optimization properties may be defined under the structure. Alternatively, eoi_type can identify only some of the optimization properties defined in Table 1. Other optimization properties not defined in Table 1 may be additionally defined in Table 1.

[0100] bitMaskInterpretation0x01Object-based optimization; the pictures for which this SEI message persists have been pre-processed or encoded so that detected objects in the pictures are optimized with respect to other parts of the pictures for the indicated optimization purposes0x02Temporal resampling optimization0x04Spatial resampling optimization0x08Temporal quality optimization0x10Spatial quality optimization; the pictures for which this SEI message persists have been pre-processed or encoded to reduce unnecessary information or improve the quality of necessary information.(e.g to reduce the amount of noise and remove speckles at the picture-level)0x20personal information protection optimization; the pictures for which this SEI message persists have been pre-processed or encoded to protect personal information. (e.g. removal or replacing of personal identifiable information, pseudonymization, anonymization)

[0101] If (eoi_type & bitMask) is not 0, this may indicate that an optimization property with the bitMask value in Table 1 has been applied. If eoi_type is greater than 0 and (eoi_type & bitMask) is 0, this may indicate that an optimization property with the corresponding bitMask value has not been applied. If eoi_type is 0, this may indicate that an optimization determined by the application has been used. For example, optimization properties according to eoi_type may be defined as in Table 2 below.

[0102] ValueInterpretationeoi_type = = 0May be used as determined by the applicationeoi_type > 0 &&( eoi_type & 0x01 ) = = 0No object-based optimization( eoi_type & 0x01 ) != 0With object-based optimizationeoi_type > 0 &&( eoi_type & 0x02 ) = = 0No temporal resampling optimization( eoi_type & 0x02 ) != 0With temporal resampling optimizationeoi_type > 0 &&( eoi_type & 0x04 ) = = 0No spatial resampling optimization( eoi_type & 0x04 ) != 0With spatial resampling optimizationeoi_type > 0 &&( eoi_type & 0x08 ) = = 0No temporal quality optimization( eoi_type & 0x08 ) != 0With temporal quality optimizationeoi_type > 0 &&( eoi_type & 0x10 ) = = 0No spatial quality optimization( eoi_type & 0x10 ) != 0With spatial quality optimization; the pictures for which this SEI message persists have been pre-processed or encoded to reduce unnecessary information or improve the quality of necessary information.(e.g to reduce the amount of noise and remove speckles at the picture-level)eoi_type > 0 &&( eoi_type & 0x20 ) == 0No optimization for personal information protection( eoi_type & 0x20 ) != 0With personal information protection optimization; the pictures for which this SEI message persists have been pre-processed or encoded to protect personal information. (eg removal or replacing of personal identifiable information, pseudonymization, anonymization).

[0103] The encoder optimization information regarding the purpose, properties, and scope of the aforementioned optimization can be applied equally to the embodiments described below. Encoder optimization information can be defined in the SEI message as shown in Table 3 below.

[0104] encoder_optimization_info(payloadSize) {Descriptor eoi_cancel_flagu(1) if(!eoi_cancel_flag){ eoi_persistence_flagu(1) eoi_for_human_viewing_idcu(2) eoi_for_machine_analysis_idcu(2) eoi_typeu(16)}}

[0105] To clarify the purpose of bitstream optimization, constraints on the values ​​of eoi_for_machine_analysis_idc and eoi_for_human_viewing_idc can be defined.

[0106] 1. Constraints to prevent cases where there is no optimization target

[0107] The optimization target defined based on encoder optimization information can be at least one of human viewing and machine analysis. In this case, defining cases where neither human viewing nor machine analysis are optimization targets may be unnecessary and wasteful from a signaling perspective. Therefore, both eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc can be restricted to not have values ​​that indicate unsuitability for the corresponding purpose. In other words, the constraint can be applied that both eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc cannot have values ​​equal to 1.

[0108] Specifically, eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc can be encoded as any integer in the range of 0 to 3. At this time, the combination of values ​​that eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc can have can be limited to 15 cases as shown in Table 4 below. The case where both eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc are 1 may not be allowed.

[0109] Caseeoi_for_human_viewing_idceoi_for_machine_analysis_idc100201302403510612713820921102211231230133114321533

[0110] eoi_for_human_viewing_idc may be encoded before eoi_for_machine_analysis_idc. In this case, the range (or number) of values ​​that can be set for eoi_for_machine_analysis_idc may differ based on the value of eoi_for_human_viewing_idc.

[0111] For example, if the value of eoi_for_human_viewing_idc is encoded as 1, eoi_for_machine_analysis_idc can be encoded as any one of 0, 2, or 3, and can be restricted from being encoded as the value of 1. That is, if the value of eoi_for_human_viewing_idc is encoded as 1, the number of values ​​that can be set for eoi_for_machine_analysis_idc can be 3. On the other hand, if the value of eoi_for_human_viewing_idc is not encoded as 1, eoi_for_machine_analysis_idc can be encoded as any one of 0, 1, 2, or 3. That is, if the value of eoi_for_human_viewing_idc is not encoded as 1, the number of values ​​that can be set for eoi_for_machine_analysis_idc can be 4.

[0112] Alternatively, eoi_for_machine_analysis_idc may be encoded before eoi_for_human_viewing_idc. In this case, the range (or number) of values ​​that can be set for eoi_for_human_viewing_idc may differ based on the value of eoi_for_machine_analysis_idc.

[0113] For example, if the value of eoi_for_machine_analysis_idc is encoded as 1, eoi_for_human_viewing_idc can be encoded as any one of 0, 2, or 3, and can be restricted from being encoded as 1. That is, if the value of eoi_for_machine_analysis_idc is encoded as 1, the number of values ​​that can be set for eoi_for_human_viewing_idc can be 3. On the other hand, if the value of eoi_for_machine_analysis_idc is not encoded as 1, eoi_for_human_viewing_idc can be encoded as any one of 0, 1, 2, or 3. That is, if the value of eoi_for_machine_analysis_idc is not encoded as 1, the number of values ​​that can be set for eoi_for_human_viewing_idc can be 4.

[0114] 2. Constraints to reduce uncertainty in the optimization target

[0115] Considering that the information defined as encoder optimization information indicates whether the information is optimized for human viewing / machine analysis, if the suitability for both human viewing and machine analysis is unknown, it may be unnecessary and wasteful from a signaling perspective. Therefore, both eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc can be restricted to not have values ​​that indicate that the suitability for the corresponding purpose is unknown. In other words, the constraint that both eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc cannot be 0 can be applied.

[0116] Specifically, eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc can be encoded as any integer in the range of 0 to 3. At this time, the combination of values ​​that eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc can have can be limited to 15 cases as shown in Table 5 below. The case where both eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc are 0 may not be allowed.

[0117] Caseeoi_for_human_viewing_idceoi_for_machine_analysis_idc101202303410511612713820921102211231230133114321533

[0118] eoi_for_human_viewing_idc may be encoded before eoi_for_machine_analysis_idc. In this case, the range (or number) of values ​​that can be set for eoi_for_machine_analysis_idc may differ based on the value of eoi_for_human_viewing_idc.

[0119] For example, if the value of eoi_for_human_viewing_idc is encoded as 0, eoi_for_machine_analysis_idc can be encoded as any one of 1, 2, or 3, and can be restricted from being encoded as 0. That is, if the value of eoi_for_human_viewing_idc is encoded as 0, the number of values ​​that can be set for eoi_for_machine_analysis_idc can be 3. On the other hand, if the value of eoi_for_human_viewing_idc is not encoded as 0, eoi_for_machine_analysis_idc can be encoded as any one of 0, 1, 2, or 3. That is, if the value of eoi_for_human_viewing_idc is not encoded as 1, the number of values ​​that can be set for eoi_for_machine_analysis_idc can be 4.

[0120] Alternatively, eoi_for_machine_analysis_idc may be encoded before eoi_for_human_viewing_idc. In this case, the range (or number) of values ​​that can be set for eoi_for_human_viewing_idc may differ based on the value of eoi_for_machine_analysis_idc.

[0121] For example, if the value of eoi_for_machine_analysis_idc is encoded as 0, eoi_for_human_viewing_idc can be encoded as any one of 1, 2, or 3, and can be restricted from being encoded as a value of 0. That is, if the value of eoi_for_machine_analysis_idc is encoded as 0, the number of values ​​that can be set for eoi_for_human_viewing_idc can be 3. On the other hand, if the value of eoi_for_machine_analysis_idc is not encoded as 0, eoi_for_human_viewing_idc can be encoded as any one of 0, 1, 2, or 3. That is, if the value of eoi_for_machine_analysis_idc is not encoded as 0, the number of values ​​that can be set for eoi_for_human_viewing_idc can be 4.

[0122] Constraints 1 and 2 mentioned above may be applied together. That is, the constraint that the values ​​of eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc cannot both be 1, nor can they both be 0 may be applied.

[0123] Specifically, eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc can be encoded as any integer in the range of 0 to 3. At this time, the combination of values ​​that eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc can have can be limited to 14 cases as shown in Table 6 below. Cases where both eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc are 0 and cases where both eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc are 1 may not be allowed.

[0124] Caseeoi_for_human_viewing_idceoi_for_machine_analysis_idc10120230341051261372082192210231130123113321433

[0125] eoi_for_human_viewing_idc may be encoded before eoi_for_machine_analysis_idc. In this case, the range (or number) of values ​​that can be set for eoi_for_machine_analysis_idc may differ based on the value of eoi_for_human_viewing_idc.

[0126] For example, if the value of eoi_for_human_viewing_idc is encoded as 0, eoi_for_machine_analysis_idc can be encoded as any one of 1, 2, or 3, and can be restricted not to be encoded as the value 0. If the value of eoi_for_human_viewing_idc is encoded as 1, eoi_for_machine_analysis_idc can be encoded as any one of 0, 2, or 3, and can be restricted not to be encoded as the value 1. That is, if the value of eoi_for_human_viewing_idc is encoded as 0 or 1, the number of values ​​that can be set for eoi_for_machine_analysis_idc can be 3. On the other hand, if the value of eoi_for_human_viewing_idc is encoded as 2 or 3, eoi_for_machine_analysis_idc can be encoded as any one of 0, 1, 2, or 3. That is, if the value of eoi_for_human_viewing_idc is encoded as 2 or 3, the number of values ​​that can be set for eoi_for_machine_analysis_idc can be 4.

[0127] Alternatively, eoi_for_machine_analysis_idc may be encoded before eoi_for_human_viewing_idc. In this case, the range (or number) of values ​​that can be set for eoi_for_human_viewing_idc may differ based on the value of eoi_for_machine_analysis_idc.

[0128] For example, if the value of eoi_for_machine_analysis_idc is encoded as 0, eoi_for_human_viewing_idc can be encoded as any one of 1, 2, or 3, and can be restricted not to be encoded as the value 0. If the value of eoi_for_machine_analysis_idc is encoded as 1, eoi_for_human_viewing_idc can be encoded as any one of 0, 2, or 3, and can be restricted not to be encoded as the value 1. That is, if the value of eoi_for_machine_analysis_idc is encoded as 0 or 1, the number of values ​​that can be set for eoi_for_human_viewing_idc can be 3. On the other hand, if the value of eoi_for_machine_analysis_idc is encoded as 2 or 3, eoi_for_human_viewing_idc can be encoded as any one of 0, 1, 2, or 3. That is, if the value of eoi_for_machine_analysis_idc is encoded as 2 or 3, the number of values ​​that can be set for eoi_for_human_viewing_idc can be 4.

[0129] 3. Constraints for defining a clear optimization target (1)

[0130] When optimization is performed, the target of optimization must be identifiable. However, if the suitability of human viewing and machine analysis is unclear, it can be difficult to identify the target of optimization and its suitability. Table 7 illustrates cases where the suitability of human viewing and machine analysis is unclear. However, this is merely an example, and some of Cases 1 through 4 may not be included if the suitability of human viewing and machine analysis is unclear.

[0131] Caseeoi_for_machine_analysis_idceoi_for_human_viewing_idc100201310411

[0132] Case 1 may apply when it is unclear whether the data is suitable for human viewing or machine analysis. Case 2 may apply when it is unsuitable for human viewing and unsure whether it is suitable for machine analysis. Case 3 may apply when it is unsuitable for machine analysis and unsure whether it is suitable for human viewing. Case 4 may apply when it is unsuitable for both human viewing and machine analysis.

[0133] In order to use a video for its intended purpose, it may be required to be suitable for at least one of machine analysis and human viewing. That is, a constraint may be applied such that at least one of the values ​​of eoi_for_machine_analysis_idc or eoi_for_human_viewing_idc must be greater than 1.

[0134] Specifically, eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc can be encoded as any integer in the range of 0 to 3. However, the combination of values ​​that eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc can have can be limited to 12 cases as shown in Table 8 below. The minimum value of the sum of eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc can be 2, and the maximum value of the sum of eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc can be 6.

[0135] Caseeoi_for_human_viewing_idceoi_for_machine_analysis_idc102203312413520621722823930103111321233

[0136] eoi_for_human_viewing_idc may be encoded before eoi_for_machine_analysis_idc. In this case, the range (or number) of values ​​that can be set for eoi_for_machine_analysis_idc may differ based on the value of eoi_for_human_viewing_idc.

[0137] For example, if the value of eoi_for_human_viewing_idc is encoded as 0 or 1, eoi_for_machine_analysis_idc can be encoded as either 2 or 3, and can be restricted from being encoded as either 0 or 1. That is, if the value of eoi_for_human_viewing_idc is encoded as 0 or 1, the number of values ​​that can be set for eoi_for_machine_analysis_idc can be 2. On the other hand, if the value of eoi_for_human_viewing_idc is encoded as 2 or 3, eoi_for_machine_analysis_idc can be encoded as either 0, 1, 2, or 3. That is, if the value of eoi_for_human_viewing_idc is encoded as 2 or 3, the number of values ​​that can be set for eoi_for_machine_analysis_idc can be 4.

[0138] Alternatively, eoi_for_machine_analysis_idc may be encoded before eoi_for_human_viewing_idc. In this case, the range (or number) of values ​​that can be set for eoi_for_human_viewing_idc may differ based on the value of eoi_for_machine_analysis_idc.

[0139] For example, if the value of eoi_for_machine_analysis_idc is encoded as 0 or 1, eoi_for_human_viewing_idc can be encoded as either 2 or 3, and can be restricted from being encoded as either 0 or 1. That is, if the value of eoi_for_machine_analysis_idc is encoded as 0 or 1, the number of values ​​that can be set for eoi_for_human_viewing_idc can be 2. On the other hand, if the value of eoi_for_machine_analysis_idc is encoded as 2 or 3, eoi_for_human_viewing_idc can be encoded as either 0, 1, 2, or 3. That is, if the value of eoi_for_machine_analysis_idc is encoded as 2 or 3, the number of values ​​that can be set for eoi_for_human_viewing_idc can be 4.

[0140] 4. Constraints for defining a clear optimization target (2)

[0141] The optimization target identified based on at least one of eoi_for_machine_analysis_idc or eoi_for_human_viewing_idc may be unclear. For example, if the values ​​of both eoi_for_machine_analysis_idc and eoi_for_human_viewing_idc are 2, this may mean that the model is suitable for machine analysis and human viewing, but is not optimized for that purpose. At least one of machine analysis and human viewing needs to be clearly defined as the optimization target. To this end, a constraint may be applied that requires that the value of at least one of eoi_for_machine_analysis_idc and eoi_for_human_viewing_idc be 3. In this case, the combinations of values ​​that eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc can have may be limited to seven cases, as shown in Table 9 below.

[0142] Caseeoi_for_human_viewing_idceoi_for_machine_analysis_idc133232331430523613703

[0143] eoi_for_human_viewing_idc may be encoded before eoi_for_machine_analysis_idc. In this case, the range (or number) of values ​​that can be set for eoi_for_machine_analysis_idc may differ based on the value of eoi_for_human_viewing_idc.

[0144] For example, if the value of eoi_for_human_viewing_idc is encoded as 3, eoi_for_machine_analysis_idc can be encoded as any one of 0, 1, 2, or 3. That is, if the value of eoi_for_human_viewing_idc is encoded as 3, the number of values ​​that can be set for eoi_for_machine_analysis_idc can be 4. On the other hand, if the value of eoi_for_human_viewing_idc is not encoded as 3, eoi_for_machine_analysis_idc can be limited to be encoded as 3. That is, if the value of eoi_for_human_viewing_idc is not encoded as 3, the number of values ​​that can be set for eoi_for_machine_analysis_idc can be 1.

[0145] Alternatively, eoi_for_machine_analysis_idc may be encoded before eoi_for_human_viewing_idc. In this case, the range (or number) of values ​​that can be set for eoi_for_human_viewing_idc may differ based on the value of eoi_for_machine_analysis_idc.

[0146] For example, if the value of eoi_for_machine_analysis_idc is encoded as 3, eoi_for_human_viewing_idc can be encoded as any one of 0, 1, 2, or 3. That is, if the value of eoi_for_machine_analysis_idc is encoded as 3, the number of values ​​that can be set for eoi_for_human_viewing_idc can be 4. On the other hand, if the value of eoi_for_machine_analysis_idc is not encoded as 3, eoi_for_human_viewing_idc can be limited to be encoded as 3. That is, if the value of eoi_for_machine_analysis_idc is not encoded as 3, the number of values ​​that can be set for eoi_for_human_viewing_idc can be 1.

[0147] When the above constraint (2) is applied, there are seven combinations of the values ​​of eoi_for_machine_analysis_idc and eoi_for_human_viewing_idc. In this case, encoding eoi_for_machine_analysis_idc and eoi_for_human_viewing_id each with two bits may be inefficient, since the seven cases can be expressed with three bits. Therefore, eoi_for_machine_analysis_idc and eoi_for_human_viewing_idc can be integrated into one identifier encoded with three bits, and Table 10 shows an example of the one identifier.

[0148] encoder_optimization_info(payloadSize) {Descriptor eoi_cancel_flagu(1) if(!eoi _cancel_flag){ eoi _persistence_flagu(1) eoi _purpose_idcu(3) eoi_typeu(16)}}

[0149] As shown in Table 10, eoi_for_machine_analysis_idc and eoi_for_human_viewing_idc can be defined by replacing them with a single identifier (eoi_purpose_idc). eoi_cancel_flag, eoi_persistence_flag, and eoi_type are as discussed above.

[0150] eoi_purpose_idc can be information for identifying the optimization purpose. Optimization purposes that can be identified based on eoi_purpose_idc include human viewing and machine analysis. Depending on the value of eoi_purpose_idc, the optimization status can be identified in detail as follows: 1) it is unknown whether the optimization is suitable for a specific purpose, 2) the optimization is not suitable for a specific purpose, 3) the optimization is suitable for a specific purpose but not optimized for the specific purpose, and 4) the optimization is suitable for a specific purpose and optimized for the specific purpose. For example, eoi_purpose_idc can be defined as shown in Table 11 below.

[0151] eoi_purpose_idcInterpretation000Suitable for human viewing and machine analysis, optimized for human viewing and machine analysis001Suitable for human viewing and optimized for human viewing. Suitable for machine analysis, but not optimized for machine analysis010Suitable for human viewing, but not optimized for human viewing. Suitable for machine analysis and optimized for machine analysis011Suitable for human viewing and optimized for human viewing. Not suitable for machine analysis100Not suitable for human viewing, suitable for machine analysis and optimized for machine analysis101Suitable for human viewing and optimized for human viewing. Not known whether suitable for machine analysis110Not known whether suitable for human viewing, suitable for machine analysis and optimized for machine analysis

[0152] Depending on the values ​​of eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc or the value of the aforementioned eoi_purpose_idc, there may be restrictions on the values ​​of subsequently defined eoi_type and / or details according to eoi_type.

[0153] Based on eoi_type, the applied optimization properties can be identified. Details regarding the applied optimization properties can be additionally defined in the encoder optimization information. The optimization properties identifiable based on eoi_type may include at least one of object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, privacy protection optimization, or bitdepth optimization.

[0154] Table 12 below is an example of details regarding applied optimization properties according to eoi_type.

[0155] encoder_optimization_info(payloadSize) {Descriptoreoi_cancel_flagu(1)if( !eoi_cancel_flag ) {eoi_persistence_flagu(1)eoi_for_human_viewing_idcu(1)eoi_for_machine_analysis_idcu(1)eoi_typeu(16)if( EoiObjectBasedFlag )eoi_object_based_idcu(16)if( EoiTemporalResamplingFlag ) {eoi_temporal_resampling_type_flagu(1)eoi_num_int_picsue(v)}if( EoiPrivacyProtectionFlag ) {eoi_privacy_protection_type_idcu(4)eoi_privacy_protected_info_typeu(8)}if(EoiBitdepthOptimizationFlag){eoi_bit_depth_optimization_type_flagu(1)eoi_bit_depth_shift_lumau(3)eoi_bit_depth_shift_chromau(3)}}}

[0156] Information related to object-based optimization can be generated and encoded in the bitstream based on a flag (eoiObjectBasedFlag) indicating whether object-based optimization is applied.

[0157] For example, if the value of eoiObjectBasedFlag is 1, this may mean that object-based optimization is applied, and if the value of eoiObjectBasedFlag is 0, this may mean that object-based optimization is not applied. eoiObjectBasedFlag can be derived as shown in the following mathematical expression 1.

[0158] [Mathematical Formula 1]

[0159] eoiObjectBasedFlag = ( ( eoi_type & 0x01 ) > 0 ) ? 1:0

[0160] If the value of eoiObjectBasedFlag is 1, object-based optimization-related information may be generated and encoded in the bitstream. On the other hand, if the value of eoiObjectBasedFlag is 0, object-based optimization-related information may not be generated and may not be encoded in the bitstream.

[0161] Information related to object-based optimization may include an identifier (eoi_object_based_idc) indicating the type of object-based optimization.

[0162] For example, eoi_object_based_idc can be defined as in Table 13 below. Here, if the value of (eoi_object_based_idc & bitMask) is not 0, it may indicate that the type of object-based optimization corresponding to the bitMask value has been applied. If eoi_object_based_idc is greater than 0 and the value of (eoi_object_based_idc & bitMask) is 0, it may indicate that the type of object-based optimization corresponding to the bitMask value has not been applied. If the value of eoi_object_based_idc is 0, it may indicate that the type of object-based optimization defined by the application has been applied. The value of eoi_object_based_idc may be in the range of 0 to 7. The values ​​8 to 65,535 for eoi_object_based_idc are reserved for future use in ITU-T|ISO / IEC and may not be present in the bitstream. If the value of eoi_object_based_idc is in the range of 8 to 65,535, the decoder must ignore the value of eoi_object_based_idc.

[0163] bitMaskInterpretation0x01Areas outside the detected objects have been blurred prior to encoding.0x02Areas outside the detected objects have been encoded with coarser transform-domain quantization than the quantization used for the detected objects.0x04Areas outside the detected objects have been overwritten. For example, an encoding system can overwrite areas outside the detected objects with a constant sample value.

[0164] When the value of eoi_for_human_viewing_idc is 3, the constraint may be that the value of (eoi_object_based_idc & 0x04) cannot be 1. Alternatively, when the value of (eoi_object_based_idc & 0x04) is 1, the constraint may be that the value of eoi_for_human_viewing_idc cannot be 3. Alternatively, when the value of eoi_for_human_viewing_flag is 3, the constraint may be that the value of EoiObjectBasedFlag cannot be 1. For example, the constraint may be that the value of (eoi_type & 0x01) cannot be greater than 1. This is because when the value of EoiObjectBasedFlag is 1, there may be differential quality depending on the region of the image, which may adversely affect human viewing.

[0165] When eoi_for_human_viewing_idc is 3, a constraint may be applied that object-based optimization related information is not generated and encoded in the bitstream. To this end, when eoi_for_human_viewing_idc is 3, the value of EoiObjectBasedFlag may be forced to be derived as 0. On the other hand, when eoi_for_human_viewing_idc is not 3, EoiObjectBasedFlag may be adaptively derived as either 0 or 1, and object-based optimization related information may be generated and encoded in the bitstream based on the value of EoiObjectBasedFlag. That is, when eoi_for_human_viewing_idc is not 3, the constraint described above may not be applied. In this way, whether or not to encode object-based optimization related information may be determined based on the value of eoi_for_human_viewing_idc.

[0166] Information related to temporal resampling optimization can be generated and encoded in the bitstream based on a flag (eoiTemporalResamplingFlag) indicating whether temporal resampling optimization is applied.

[0167] For example, if the value of eoiTemporalResamplingFlag is 1, this may mean that temporal resampling optimization is applied, and if the value of eoiTemporalResamplingFlag is 0, this may mean that temporal resampling optimization is not applied. eoiTemporalResamplingFlag can be derived as shown in the following mathematical expression 2.

[0168] [Equation 2]

[0169] eoiTemporalResamplingFlag = ( ( eoi_type & 0x02 ) > 0 ) ? 1:0

[0170] If the value of eoiTemporalResamplingFlag is 1, temporal resampling optimization related information may be generated and encoded in the bitstream. On the other hand, if the value of eoiTemporalResamplingFlag is 0, temporal resampling optimization related information may not be generated and may not be encoded in the bitstream.

[0171] Information related to spatial resampling optimization can be generated and encoded in the bitstream based on a flag (eoiSpatialResamplingFlag) indicating whether spatial resampling optimization is applied.

[0172] For example, if the value of eoiSpatialResamplingFlag is 1, this may mean that spatial resampling optimization is applied, and if the value of eoiSpatialResamplingFlag is 0, this may mean that spatial resampling optimization is not applied. eoiSpatialResamplingFlag can be derived as shown in the following mathematical expression 3.

[0173] [Equation 3]

[0174] eoiSpatialResamplingFlag = ( ( eoi_type & 0x04 ) > 0 ) ? 1:0

[0175] If the value of eoiSpatialResamplingFlag is 1, spatial resampling optimization related information may be generated and encoded in the bitstream. On the other hand, if the value of eoiSpatialResamplingFlag is 0, spatial resampling optimization related information may not be generated and may not be encoded in the bitstream.

[0176] Information related to temporal quality optimization can be generated based on a flag (eoiTemporalQualityFlag) indicating whether temporal quality optimization is applied and encoded in the bitstream.

[0177] For example, if the value of eoiTemporalQualityFlag is 1, this may mean that temporal quality optimization has been applied, and if the value of eoiTemporalQualityFlag is 0, this may mean that temporal quality optimization has not been applied. eoiTemporalQualityFlag can be derived as shown in the following mathematical expression 4.

[0178] [Equation 4]

[0179] eoiTemporalQualityFlag = ( ( eoi_type & 0x08 ) > 0 ) ? 1:0

[0180] If the value of eoiTemporalQualityFlag is 1, temporal quality optimization-related information may be generated and encoded in the bitstream. On the other hand, if the value of eoiTemporalQualityFlag is 0, temporal quality optimization-related information may not be generated and may not be encoded in the bitstream.

[0181] Information related to spatial quality optimization can be generated based on a flag (eoiSpatialQualityFlag) indicating whether spatial quality optimization is applied and can be encoded in the bitstream.

[0182] For example, if the value of eoiSpatialQualityFlag is 1, this may mean that spatial quality optimization has been applied, and if the value of eoiSpatialQualityFlag is 0, this may mean that spatial quality optimization has not been applied. eoiSpatialQualityFlag can be derived as shown in the following mathematical expression (5).

[0183] [Equation 5]

[0184] eoiSpatialQualityFlag = ( ( eoi_type & 0x10 ) > 0 ) ? 1:0

[0185] If the value of eoiSpatialQualityFlag is 1, spatial quality optimization-related information may be generated and encoded in the bitstream. On the other hand, if the value of eoiSpatialQualityFlag is 0, spatial quality optimization-related information may not be generated and may not be encoded in the bitstream.

[0186] Information related to privacy optimization can be generated based on a flag (eoiPrivacyProtectionFlag) indicating whether privacy optimization is applied and encoded in the bitstream.

[0187] For example, if the value of eoiPrivacyProtectionFlag is 1, this may mean that privacy optimization has been applied, and if the value of eoiPrivacyProtectionFlag is 0, this may mean that privacy optimization has not been applied. eoiPrivacyProtectionFlag can be derived as shown in the following mathematical expression (6).

[0188] [Equation 6]

[0189] eoiPrivacyProtectionFlag = ( ( eoi_type & 0x20 ) > 0 ) ? 1:0

[0190] If the value of eoiPrivacyProtectionFlag is 1, privacy optimization-related information may be generated and encoded in the bitstream. Conversely, if the value of eoiPrivacyProtectionFlag is 0, privacy optimization-related information may not be generated and may not be encoded in the bitstream.

[0191] Information related to privacy optimization may include at least one of an identifier indicating the type of privacy optimization (eoi_privacy_protection_type_idc) or information indicating the type of protected information (eoi_privacy_protected_info_type). For example, eoi_privacy_protection_type_idc may be defined as shown in Table 14 below. eoi_privacy_protected_info_type may be defined as shown in Table 15 below.

[0192] eoi_privacy_protection_type_idcInterpretation0determinated by the application1Blurring; personal information is blurred to make it unidentifiable.2Replacing; personal information is replaced with something different from the original to make it unidentifiable.3Masking; personal information is masked so that it cannot be identified4… Reserved for future use.

[0193] bitMaskInterpretation0x01Information that identifies a person is protected. For example, the face of the person.0x02Information that can identify vehicles is protected. For example, the license plate of the vehicle.0x04Information that can infer locations is protected. For example text or images on signs.

[0194] If the value of eoi_privacy_protected_info_type is greater than 0 and (eoi_privacy_protected_info_type & bitMask) is non-zero, it may indicate that the information corresponding to the bitMask value is protected. If the value of eoi_privacy_protected_info_type is 0, it may indicate that information of an application-defined type is protected. The value of eoi_privacy_protection_info_type may be in the range of 0 to 7. The values ​​8 to 255 for eoi_privacy_protected_info_type are reserved for future use by ITU-T|ISO / IEC and may not be present in the bitstream. If the value of eoi_privacy_protected_info_type is in the range of 8 to 255, the decoder should ignore eoi_privacy_protected_info_type.

[0195] If the value of eoi_for_human_viewing_idc is 3, a constraint may be applied to eoi_type that states that the value of eoiPrivacyProtectionFlag must not be 1. This is because protecting personal information would result in a degradation of human-perceivable information in the video, making it difficult to consider it optimized for human viewing.

[0196] When eoi_for_human_viewing_idc is 3, a constraint may be applied that privacy optimization-related information is not generated and encoded in the bitstream. To this end, when eoi_for_human_viewing_idc is 3, the value of eoiPrivacyProtectionFlag may be forced to be derived as 0. On the other hand, when eoi_for_human_viewing_idc is not 3, eoiPrivacyProtectionFlag may be adaptively derived as either 0 or 1, and privacy optimization-related information may be generated and encoded in the bitstream based on the value of eoiPrivacyProtectionFlag. That is, when eoi_for_human_viewing_idc is not 3, the constraint described above may not be applied. In this way, whether or not privacy optimization-related information is encoded may be determined based on the value of eoi_for_human_viewing_idc.

[0197] Information related to bitdepth optimization can be generated based on a flag (eoiBitdepthOptimizationFlag) indicating whether bitdepth optimization is applied and encoded in the bitstream.

[0198] For example, if the value of eoiBitdepthOptimizationFlag is 1, this may mean that bitdepth optimization is applied, and if the value of eoiBitdepthOptimizationFlag is 0, this may mean that bitdepth optimization is not applied. eoiBitdepthOptimizationFlag can be derived as shown in the following mathematical expression (7).

[0199] [Equation 7]

[0200] eoiBitdepthOptimizationFlag = ( ( eoi_type & 0x40 ) > 0 ) ? 1:0

[0201] If the value of eoiBitdepthOptimizationFlag is 1, bitdepth optimization-related information may be generated and encoded in the bitstream. On the other hand, if the value of eoiBitdepthOptimizationFlag is 0, bitdepth optimization-related information may not be generated and may not be encoded in the bitstream.

[0202] Information related to bit depth optimization may include at least one of information indicating a bit depth optimization type (or a bit shift direction) (eoi_bit_depth_optimization_type_flag), information for identifying a shift size for a luma component (eoi_bit_depth_shift_luma), or information for identifying a shift size for a chroma component (eoi_bit_depth_shift_chroma).

[0203] When the value of eoi_bit_depth_optimization_type_flag is 0, it may indicate that the bit depth optimization type is truncating bit depth, and when the value of eoi_bit_depth_optimization_type_flag is 1, it may indicate that the bit depth optimization type is increasing bit depth. For example, when the value of eoi_bit_depth_optimization_type_flag is 0, a right shift may be applied, and when the value of eoi_bit_depth_optimization_type_flag is 1, a left shift may be applied.

[0204] eoi_bit_depth_shift_luma can indicate the right shift amount for the luma component when truncating bit depth is applied, and can indicate the left shift amount for the luma component when increasing bit depth is applied. For example, the values ​​0, 1, 2, 3, 4, 5, 6, and 7 of eoi_bit_depth_shift_luma can mean that the input sample is shifted by 0, 1, 2, 3, 4, 5, 6, and 7 bits, respectively. If there is no shift, there may be no reason to define the value, in which case the value 0 may be undefined.

[0205] eoi_bit_depth_shift_chroma can indicate the right shift amount for chroma components when truncating bit depth is applied, and can indicate the left shift amount for chroma components when increasing bit depth is applied. For example, the values ​​0, 1, 2, 3, 4, 5, 6, and 7 of eoi_bit_depth_shift_chroma can mean that the input sample is shifted by 0, 1, 2, 3, 4, 5, 6, and 7 bits, respectively. If there is no shift, there may be no reason to define the value, in which case the value 0 may be undefined.

[0206] Based on the value of eoi_for_human_viewing_idc, the range (or number) of values ​​that can be set for eoi_bit_depth_optimization_type_flag may vary. Based on the value of eoi_for_human_viewing_idc, the range (or number) of applicable bit depth optimization types may vary.

[0207] For example, if eoi_for_human_viewing_idc is 3, a constraint that eoi_bit_depth_optimization_type_flag must not be 0 may be applied to eoi_type. This is because excessive truncating bit depth can cause a degradation in the quality of information that humans can perceive through the video, making it difficult to call it optimized for human viewing.

[0208] When eoi_for_human_viewing_idc is 3, a constraint may be applied that eoi_bit_depth_optimization_type_flag is generated with a value of 1 and encoded in the bitstream. That is, when eoi_for_human_viewing_idc is 3, the number of values ​​that can be set for eoi_bit_depth_optimization_type_flag may be 1, and the applicable bit depth optimization type may be limited to increasing bit depth. On the other hand, when eoi_for_human_viewing_idc is not 3, eoi_bit_depth_optimization_type_flag may be generated with a value of either 0 or 1 and encoded in the bitstream. When eoi_for_human_viewing_idc is not 3, the constraint described above may not be applied. That is, if eoi_for_human_viewing_idc is not 3, the number of possible values ​​for eoi_bit_depth_optimization_type_flag can be 2, and the applicable bit depth optimization types can include truncating bit depth and increasing bit depth.

[0209] Alternatively, if the value of eoi_for_human_viewing_idc is 3 and the value of eoi_bit_depth_optimization_type_flag is 0, the values ​​of eoi_bit_depth_shift_luma and eoi_bit_depth_shift_chroma can be constrained to minimize quality degradation due to bit truncation. For example, the values ​​of eoi_bit_depth_shift_luma and eoi_bit_depth_shift_chroma can be constrained so that the size of the optimized bit depth is not less than a certain value (e.g., 8).

[0210] Depending on the bit depth optimization type (or, eoi_bit_depth_optimization_type_flag), the range or number of values ​​that can be set for eoi_for_human_viewing_idc may differ. For example, if truncating bit depth is applied, a constraint may be applied that the value of eoi_for_human_viewing_idc must not be 3. That is, if truncating bit depth is applied, eoi_for_human_viewing_idc can be encoded as any one of the values ​​0, 1, or 2. If truncating bit depth is applied, the number of values ​​that can be set for eoi_for_human_viewing_idc can be 3. If increasing bit depth is applied, no constraint may be applied to the value of eoi_for_human_viewing_idc. That is, if increasing bit depth is applied, eoi_for_human_viewing_idc can be encoded as any one of the values ​​0 to 3. When increasing bit depth is applied, the number of possible values ​​for eoi_for_human_viewing_idc can be 4.

[0211] Depending on whether the size of the optimized bit depth is less than a specific value, the range or number of values ​​that can be set for eoi_for_human_viewing_idc may differ. For example, if the size of the optimized bit depth is less than a specific value (e.g., 8), a constraint may be applied that the value of eoi_for_human_viewing_idc must not be 3. That is, if the size of the optimized bit depth is less than a specific value (e.g., 8), eoi_for_human_viewing_idc may be encoded as any one of 0, 1, or 2. If the size of the optimized bit depth is less than a specific value (e.g., 8), the number of values ​​that can be set for eoi_for_human_viewing_idc may be 3. On the other hand, if the size of the optimized bit depth is greater than or equal to a specific value (e.g., 8), no constraint may be applied to the value of eoi_for_human_viewing_idc. That is, eoi_for_human_viewing_idc can be encoded as any one of the values ​​0 to 3, and the number of values ​​that can be set for eoi_for_human_viewing_idc can be 4.

[0212] A bitstream including the compressed video picture and the encoder optimization information can be generated (S430). The encoder optimization information can be configured in an SEI message of the bitstream. The SEI message can be included in a NAL (network abstraction layer) unit of the bitstream.

[0213] FIG. 5 illustrates a schematic configuration of an encoding device (200) that performs a method for generating a bitstream according to the present disclosure.

[0214] Referring to FIG. 5, the encoding device (200) may include a receiving unit (500), a video compression unit (510), an EOI generation unit (520), and a bitstream generation unit (530).

[0215] The receiving unit (500) can receive one or more video pictures to be encoded.

[0216] The video compression unit (510) can encode one or more received video pictures to generate compressed video pictures.

[0217] The EOI generation unit (520) can generate encoder optimization information. The encoder optimization information can be generated by taking into account the aforementioned constraints, as described with reference to FIG. 4.

[0218] The bitstream generation unit (530) can generate a bitstream including a compressed video picture and encoder optimization information.

[0219] FIG. 6 illustrates a method for restoring a video picture performed in a decoding device (300) according to the present disclosure.

[0220] A bitstream including an encoded video picture can be received (S600).

[0221] The encoded video picture of the bitstream can be restored (S610).

[0222] Video information about an encoded video picture can be extracted from a bitstream. The encoded video picture can be restored based on the extracted video information.

[0223] Additionally, the NAL unit of the bitstream may include an SEI message. The SEI message may define encoder optimization information (EOI). The encoder optimization information may be related to the encoded video picture. The encoder optimization information may be included in the SEI message and signaled through the bitstream.

[0224] The above encoder optimization information may include an EOI cancellation flag (eoi_cancel_flag). Based on eoi_cancel_flag being 0, optimization-related information may be included in the encoder optimization information, as discussed with reference to FIG. 4.

[0225] eoi_for_machine_analysis_idc and eoi_for_human_viewing_idc, signaled through encoder optimization information, are encoded according to the aforementioned constraints and can be decoded into values ​​that comply with the constraints. This is as described with reference to Fig. 4, and a redundant description will be omitted here.

[0226] Based on the value of eoi_for_human_viewing_idc, the range (or number) of values ​​that can be set for eoi_for_machine_analysis_idc may be different. Alternatively, based on the value of eoi_for_machine_analysis_idc, the range (or number) of values ​​that can be set for eoi_for_human_viewing_idc may be different.

[0227] Based on the value of eoi_for_human_viewing_idc, whether or not to decode detailed information about optimization properties according to eoi_type can be determined. Based on the value of eoi_for_human_viewing_idc, the range (or number) of values ​​that can be set for detailed information about optimization properties according to eoi_type can vary. This is as discussed with reference to Fig. 4.

[0228] Depending on the bit depth optimization type (or eoi_bit_depth_optimization_type_flag), the range or number of values ​​that can be set for eoi_for_human_viewing_idc may vary, as illustrated in FIG. 4.

[0229] Depending on whether the optimized bit depth size is less than a certain value, the range or number of values ​​that can be set for eoi_for_human_viewing_idc may vary, as discussed with reference to Figure 4.

[0230] FIG. 7 illustrates a schematic configuration of a decoding device (300) that performs a method for restoring a video picture according to the present disclosure.

[0231] Referring to FIG. 7, the decoding device (300) may include a receiving unit (700), a video information extraction unit (710), and a video restoration unit (720).

[0232] The receiving unit (700) can receive a bitstream including an encoded video picture.

[0233] The video information extraction unit (710) can extract video information about an encoded video picture from a bitstream. In addition, the video information extraction unit (710) can extract an SEI message from a NAL unit of the bitstream. The extracted SEI message can include encoder optimization information.

[0234] The video restoration unit (720) can restore an encoded video picture based on the extracted video information.

[0235] In the embodiments described above, the methods are described based on a flowchart as a series of steps or blocks. However, the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps described above. Furthermore, those skilled in the art will understand that the steps depicted in the flowchart are not exclusive, and other steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments of this document.

[0236] The method according to the embodiments of the present document described above can be implemented in the form of software, and the encoding device and / or decoding device according to the present document can be included in a device that performs image processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.

[0237] When the embodiments in this document are implemented as software, the above-described method can be implemented as a module (process, function, etc.) that performs the above-described function. The module can be stored in memory and executed by a processor. The memory can be internal or external to the processor and can be connected to the processor by various well-known means. The processor can include an application-specific integrated circuit (ASIC), another chipset, logic circuit, and / or data processing device. The memory can include a read-only memory (ROM), a random access memory (RAM), flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be implemented and performed on a processor, a microprocessor, a controller, or a chip. For example, the functional units illustrated in each drawing can be implemented and performed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for implementation (e.g., information on instructions) or an algorithm can be stored on a digital storage medium.

[0238] In addition, the decoding device and encoding device to which the embodiment(s) of the present specification are applied may be included in a multimedia broadcasting transmitting and receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as a video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VoD) service providing device, an OTT (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (argumente reality) device, a video phone video device, a transportation terminal (ex. a vehicle (including an autonomous vehicle) terminal, an airplane terminal, a ship terminal, etc.), and a medical video device, and may be used to process a video signal or a data signal. For example, the OTT (Over the top video) device may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc.

[0239] In addition, the processing method to which the embodiment(s) of the present specification are applied can be produced in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of the present specification can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium can include, for example, a Blu-ray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission via the Internet). In addition, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.

[0240] Additionally, the embodiments of the present disclosure may be implemented as a computer program product by program code, and the program code may be executed on a computer by the embodiments of the present disclosure. The program code may be stored on a computer-readable carrier.

[0241] FIG. 8 illustrates an example of a content streaming system to which embodiments of the present disclosure can be applied.

[0242] Referring to FIG. 8, a content streaming system to which the embodiment(s) of the present specification are applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0243] The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, generates a bitstream, and transmits it to the streaming server. Alternatively, if multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server may be omitted.

[0244] The above bitstream can be generated by an encoding method or a bitstream generation method to which the embodiment(s) of the present specification are applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0245] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, and the streaming server transmits the multimedia data to the user. At this time, the content streaming system may include a separate control server, in which case the control server controls commands / responses between each device within the content streaming system.

[0246] The streaming server can receive content from a media repository and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0247] Examples of the user devices may include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs), digital TVs, desktop computers, digital signage, etc.

[0248] Each server within the above content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.

[0249] The claims set forth in this specification may be combined in various ways. For example, the technical features of the method claims of this specification may be combined and implemented as a device, and the technical features of the device claims of this specification may be combined and implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a device, and the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a method.

Claims

1. A step of receiving a video picture to be encoded; A step of encoding the received video picture to generate a compressed video picture; A step of generating encoder optimization information (EOI); and A step of generating a bitstream including the compressed video picture and the encoder optimization information, The above encoder optimization information is encoded in the NAL (network abstraction layer) unit of the bitstream, A method wherein the encoder optimization information comprises at least one of an EOI identifier for human viewing or an EOI identifier for machine analysis.

2. In paragraph 1, A method wherein both the EOI identifier for the human viewing and the EOI identifier for the machine analysis have a value in the range of 0 to 3.

3. In paragraph 2, A method wherein the EOI identifier for the human viewing and the EOI identifier for the machine analysis are both subject to a constraint that they cannot be 1.

4. In paragraph 3, If the value of the EOI identifier for the above human viewing is 1, this means that the video is not suitable for the above human viewing, A method in which if the value of the EOI identifier for the machine analysis is 1, this means that the video is not suitable for the machine analysis.

5. In paragraph 2, A method in which the number of values ​​that can be set for the other of the EOI identifier for human viewing and the EOI identifier for machine analysis is determined differently based on the value of one of the EOI identifier for human viewing and the EOI identifier for machine analysis.

6. In paragraph 2, A method in which the EOI identifier for the human viewing and the EOI identifier for the machine analysis are both subject to a constraint that they must not be 0.

7. In paragraph 2, A method wherein a constraint is applied that at least one of the EOI identifiers for the human viewing and the EOI identifiers for the machine analysis must have a value of 3.

8. In paragraph 1, The above encoder optimization information includes object-based optimization related information, A method in which whether or not to encode the object-based optimization related information is determined based on the value of the EOI identifier for human viewing.

9. In paragraph 1, The above encoder optimization information includes information related to privacy optimization, A method in which whether to encode the above personal information optimization related information is determined based on the value of the EOI identifier for human viewing.

10. In paragraph 1, The above encoder optimization information includes bit depth optimization related information, The above bit depth optimization related information includes information indicating the type of bit depth optimization, A method wherein a constraint is applied that the value of information indicating the type of bit depth optimization must not be 0 based on the value of the EOI identifier for the human viewing.

11. A step of receiving a bitstream including an encoded video picture; and A step of restoring an encoded video picture included in the above bitstream, The above bitstream includes encoder optimization information (EOI), The above encoder optimization information is obtained from the NAL (network abstraction layer) unit of the bitstream, A method wherein the encoder optimization information comprises at least one of an EOI identifier for human viewing or an EOI identifier for machine analysis.

12. A computer-readable storage medium storing a bitstream generated by the method according to paragraph 1.

13. A step of generating a bitstream; wherein the bitstream is generated based on the steps of: receiving a video picture to be encoded; encoding the received video picture to generate a compressed video picture; generating encoder optimization information (EOI); and generating a bitstream including the compressed video picture and the encoder optimization information; and Including a step of transmitting data including the above bitstream, The above encoder optimization information is encoded in the NAL (network abstraction layer) unit of the bitstream, A method wherein the encoder optimization information comprises at least one of an EOI identifier for human viewing or an EOI identifier for machine analysis.

Citation Information

Patent Citations

  • Niobium precursor compound, composition for forming a niobium-containing film comprising the same, and method for forming a niobium-containing film using the composition

    KR1020240111081A

  • KR20230020428A

  • KR20230074521A

Cited By

  • Systems and methods for signaling encoder optimization privacy information in video coding

    US12641294B2

  • Systems and methods for signaling encoder optimization privacy information in video coding

    US20260107019A1