Method for decoding image information, method for encoding image information, method for bitstream, and computer-readable storage medium storing bitstream

By utilizing encoder optimization information (EOI) SEI messages to determine the source picture for temporal resampling, the method addresses high-resolution video challenges, enhancing coding system reliability and efficiency.

WO2026084541A1PCT designated stage Publication Date: 2026-04-23LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2025-10-20
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

The increasing demand for high-resolution, high-quality video has led to higher transmission and storage costs due to the increase in transmitted information or bits, necessitating high-efficiency video compression technology to minimize confusion and malfunction in coding systems while improving reliability, coding speed, and data transmission efficiency.

Method used

The method involves acquiring and deriving encoder optimization information (EOI) SEI messages to determine whether a current picture is a source for temporal resampling, including optimization type information and picture count information, enabling accurate identification of the original source picture.

Benefits of technology

This approach enhances the reliability and efficiency of the coding system by quickly identifying the original source picture, improving coding speed and efficiency, and reducing data transmission requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025016553_23042026_PF_FP_ABST
    Figure KR2025016553_23042026_PF_FP_ABST
Patent Text Reader

Abstract

A method according to the present disclosure comprises: acquiring image information including an encoder optimization information (EOI) supplemental enhancement information (SEI) message; and deriving information about encoder optimization on the basis of the EOI SEI message, wherein the EOI SEI message includes optimization type information indicating an optimization type including temporal resampling, and on the basis of the optimization type information, whether the current picture included in the same unit as the EOI SEI message is a source picture for the temporal resampling is derived.
Need to check novelty before this filing date? Find Prior Art

Description

A method for decoding image information, a method for encoding image information, a method relating to a bitstream, and a computer-readable storage medium for storing a bitstream

[0001] The present disclosure relates to a method for decoding image information, a method for encoding image information, a method for bitstreams, and a computer-readable storage medium for storing bitstreams.

[0002] Recently, the demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition), has been increasing across various fields. As video data becomes higher in resolution and quality, the relative amount of information or bits transmitted increases compared to conventional video data. This increase in transmitted information or bits leads to higher transmission and storage costs.

[0003] Accordingly, high-efficiency video compression technology is required to effectively transmit, store, and play back high-resolution, high-quality video information.

[0004] The present disclosure aims to suppress, prevent, or minimize confusion and / or malfunction in coding systems.

[0005] The present disclosure aims to improve the reliability of a coding system including an encoding device and a decoding device.

[0006] The present disclosure aims to improve the cocoding speed and coding efficiency of a coding system including an encoding device and a decoding device.

[0007] The present disclosure aims to improve the data transmission efficiency of a coding system including an encoding device and a decoding device.

[0008] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure belongs from the description below.

[0009] According to one aspect of the present disclosure, a method for decoding image information comprises acquiring the image information including an encoder optimization information (EOI) SEI (supplemental enhancement information) message; and deriving information regarding encoder optimization based on the encoder optimization information SEI message, wherein the encoder optimization information SEI message includes optimization type information indicating an optimization type including temporal resampling, and based on the optimization type information, whether a current picture included in the same unit as the encoder optimization information SEI message is a source picture for the temporal resampling is derived.

[0010] According to one aspect of the present disclosure, an apparatus for decoding image information comprises a memory and a processor connected to the memory, wherein the processor acquires the image information including an encoder optimization information (EOI) SEI (supplemental enhancement information) message; and derives information regarding encoder optimization based on the encoder optimization information SEI message, wherein the encoder optimization information SEI message includes optimization type information indicating an optimization type including temporal resampling, and based on the optimization type information, whether a current picture included in the same unit as the encoder optimization information SEI message is a source picture for the temporal resampling is derived.

[0011] In a method or device for decoding the above image information, the encoder optimization information SEI message may further include source picture information regarding the fact that the current picture included in the same unit as the encoder optimization information SEI message is the source picture for the temporal resampling.

[0012] In a method or device for decoding the above-mentioned image information, the encoder optimization information SEI message may further include temporal resampling type information regarding the temporal upsampling and picture count information regarding the number of pictures added or excluded by the temporal resampling.

[0013] In a method or device for decoding the above image information, the source picture information can be obtained based on the temporal resampling type information and the picture count information.

[0014] In a method or device for decoding the above image information, the source picture information can be obtained based on temporal resampling type information indicating the temporal upsampling and picture count information with a value greater than 0.

[0015] In a method or device for decoding the above image information, the temporal resampling type information and the picture count information can be obtained based on optimization type information representing the temporal resampling.

[0016] According to one aspect of the present disclosure, a method for encoding image information comprises: performing encoder optimization; and encoding the image information including an encoder optimization information (EOI) SEI (supplemental enhancement information) message generated based on information regarding the encoder optimization, wherein the encoder optimization information SEI message includes optimization type information indicating an optimization type including temporal resampling, and based on the optimization type information, it is determined whether a current picture included in the same unit as the encoder optimization information SEI message is a source picture for the temporal resampling.

[0017] According to one aspect of the present disclosure, an apparatus for encoding image information comprises a memory and a processor connected to the memory, wherein the processor performs encoder optimization; and encodes the image information including an encoder optimization information (EOI) SEI (supplemental enhancement information) message generated based on information regarding the encoder optimization, wherein the encoder optimization information SEI message includes optimization type information indicating an optimization type including temporal resampling, and based on the optimization type information, it is determined whether a current picture included in the same unit as the encoder optimization information SEI message is a source picture for the temporal resampling.

[0018] In a method or device for encoding the above-mentioned image information, the encoder optimization information SEI message may further include source picture information regarding the fact that the current picture included in the same unit as the encoder optimization information SEI message is the source picture for the temporal resampling.

[0019] In a method or device for encoding the above-mentioned image information, the encoder optimization information SEI message may further include temporal resampling type information regarding the temporal upsampling and picture count information regarding the number of pictures added or excluded by the temporal resampling.

[0020] In a method or device for encoding the above-mentioned image information, the source picture information can be obtained based on the temporal resampling type information and the picture count information.

[0021] In a method or device for encoding the above-mentioned image information, the source picture information can be obtained based on temporal resampling type information indicating the temporal upsampling and picture count information with a value greater than 0.

[0022] In a method or device for encoding the above-mentioned image information, the temporal resampling type information and the picture count information can be obtained based on optimization type information representing the temporal resampling.

[0023] A method for a bitstream according to one aspect of the present disclosure comprises: performing encoder optimization; generating a bitstream for said image information including an encoder optimization information (EOI) SEI (supplemental enhancement information) message generated based on said encoder optimization information; and transmitting data for said bitstream, wherein the encoder optimization information SEI message includes optimization type information indicating an optimization type including temporal resampling, and based on said optimization type information, it is determined whether a current picture included in the same unit as the encoder optimization information SEI message is a source picture for said temporal resampling.

[0024] According to one aspect of the present disclosure, an apparatus for a bitstream comprises: at least one processor that performs encoder optimization and generates a bitstream for said image information including an encoder optimization information (EOI) SEI (supplemental enhancement information) message generated based on information regarding said encoder optimization; and a transmission unit that transmits data for said bitstream, wherein the encoder optimization information SEI message includes optimization type information indicating an optimization type including temporal resampling, and based on said optimization type information, it is determined whether a current picture included in the same unit as the encoder optimization information SEI message is a source picture for said temporal resampling.

[0025] According to one aspect of the present disclosure, a computer-readable storage medium for storing a bitstream, wherein the storage medium stores a bitstream for the image information including an encoder optimization information (EOI) SEI (supplemental enhancement information) message generated based on information regarding encoder optimization, wherein the encoder optimization information SEI message includes optimization type information indicating an optimization type including temporal resampling, and based on the optimization type information, it is determined whether a current picture included in the same unit as the encoder optimization information SEI message is a source picture for the temporal resampling.

[0026] The features briefly summarized above regarding the present disclosure are merely exemplary aspects of the detailed description of the present disclosure that follows and do not limit the scope of the present disclosure.

[0027] According to the present disclosure, information regarding whether the current picture is the original source picture can be provided using source picture information.

[0028] According to the present disclosure, by enabling the coding system to provide information information that is more similar to the image information of the original picture, the reliability of image information transmission by the coding system can be improved.

[0029] According to the present disclosure, by enabling the coding system to identify the original source picture more quickly, the coding speed and coding efficiency of the coding system can be improved.

[0030] According to the present disclosure, the data transmission efficiency of a coding system can be improved by enabling the coding system to identify the original source picture using less data.

[0031] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs from the description below.

[0032] FIG. 1 is a schematic diagram of a VCM system to which embodiments of the present disclosure can be applied.

[0033] FIG. 2 is a schematic diagram showing a VCM pipeline structure to which embodiments of the present disclosure can be applied.

[0034] FIG. 3 is a schematic diagram of an image / video encoder to which embodiments of the present disclosure can be applied.

[0035] FIG. 4 is a schematic diagram of an image / video decoder to which embodiments of the present disclosure can be applied.

[0036] FIG. 5 is a flowchart schematically illustrating a feature / feature map encoding procedure to which embodiments of the present disclosure can be applied.

[0037] FIG. 6 is a flowchart schematically illustrating a feature / feature map decoding procedure to which embodiments of the present disclosure can be applied.

[0038] FIG. 7 is a diagram illustrating a method for decoding image information according to one embodiment of the present disclosure.

[0039] FIG. 8 is a diagram illustrating a method for encoding image information according to one embodiment of the present disclosure.

[0040] FIG. 9 is a drawing showing an example of a content streaming system to which embodiments of the present disclosure can be applied.

[0041] FIG. 10 is a drawing showing another example of a content streaming system to which embodiments of the present disclosure may be applied.

[0042] Hereinafter, embodiments of the present disclosure are described in detail with reference to the attached drawings so that those skilled in the art can easily implement them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein.

[0043] In describing the embodiments of the present disclosure, detailed descriptions of known configurations or functions are omitted if it is determined that such descriptions could obscure the essence of the present disclosure. Additionally, parts of the drawings unrelated to the description of the present disclosure have been omitted, and similar parts are denoted by similar reference numerals.

[0044] In the present disclosure, when a component is described as being "connected," "combined," or "joined" with another component, this may include not only a direct connection but also an indirect connection in which another component exists in between. Furthermore, when a component is described as "comprising" or "having" another component, this means that, unless specifically stated otherwise, it does not exclude the other component but may include an additional component.

[0045] In the present disclosure, terms such as first, second, etc. are used solely for the purpose of distinguishing one component from another and do not limit the order or importance of the components unless specifically stated otherwise. Accordingly, within the scope of the present disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and likewise, a second component in one embodiment may be referred to as a first component in another embodiment.

[0046] In this disclosure, distinct components are intended to clearly describe their respective features and do not imply that the components are separate. That is, multiple components may be integrated to form a single hardware or software unit, or a single component may be distributed to form multiple hardware or software units. Accordingly, such integrated or distributed embodiments are included within the scope of this disclosure, unless otherwise noted.

[0047] In the present disclosure, the components described in various embodiments do not necessarily mean essential components, and some may be optional components. Accordingly, embodiments consisting of a subset of the components described in one embodiment are also included within the scope of the present disclosure. Furthermore, embodiments including other components in addition to the components described in various embodiments are also included within the scope of the present disclosure.

[0048] The present disclosure relates to the encoding and decoding of images, and the terms used in the present disclosure may have the ordinary meanings commonly used in the technical field to which the present disclosure belongs, unless newly defined in the present disclosure.

[0049] The present disclosure may be applied to methods disclosed in the VVC (Versatile Video Coding) standard and / or the VCM (Video Coding for Machines) standard. Additionally, the present disclosure may be applied to methods disclosed in the EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard) or next-generation video / video coding standards (e.g., H.267 or H.268, etc.).

[0050] The present disclosure presents various embodiments regarding video / image coding, and unless otherwise noted, said embodiments may be performed in combination with one another. In the present disclosure, "video" may refer to a set of a series of images over time. "Image" may be information generated by artificial intelligence (AI). Input information used by AI in the process of performing a series of tasks, information generated during the information processing process, and output information may be used as images. "Picture" generally refers to a unit representing a single image at a specific time, and a slice / tile is a coding unit that constitutes a part of the picture in coding. A single picture may be composed of one or more slices / tiles. Additionally, a slice / tile may include one or more coding tree units (CTUs). The said CTU may be divided into one or more CUs. A tile is a rectangular area existing within a specific tile row and a specific tile column within a picture, and may be composed of multiple CTUs. A tile column can be defined as a rectangular area of ​​CTUs, has a height equal to the height of the picture, and may have a width specified by a syntax element signaled from a bitstream portion such as a picture parameter set. A tile row can be defined as a rectangular area of ​​CTUs, has a width equal to the width of the picture, and may have a height specified by a syntax element signaled from a bitstream portion such as a picture parameter set. A tile scan is a predetermined sequential ordering method of CTUs that divides a picture. Here, CTUs may be sequentially ordered according to a CTU raster scan within a tile, and tiles within a picture may be sequentially ordered according to the raster scan order of the picture's tiles.A slice may contain an integer number of complete tiles or an integer number of consecutive rows of complete CTUs within a single tile of a picture. A slice may be contained exclusively in a single NAL unit. A picture may consist of one or more tile groups. A tile group may contain one or more tiles. A brick may represent a rectangular area of ​​rows of CTUs within a tile in a picture. A tile may contain one or more bricks. A brick may represent a rectangular area of ​​rows of CTUs within a tile. A tile may be divided into multiple bricks, and each brick may contain one or more rows of CTUs belonging to the tile. A tile that is not divided into multiple bricks may also be treated as a brick.

[0051] In the present disclosure, "pixel" or "pel" may refer to the smallest unit constituting a picture (or image). Additionally, "sample" may be used as a term corresponding to pixel. A sample may generally represent a pixel or a pixel value, may represent only the pixel / pixel value of the luminance component, or may represent only the pixel / pixel value of the chroma component.

[0052] In one embodiment, particularly when applied to VCM, the pixel / pixel value may represent the independent information of each component or the pixel / pixel value of a component generated through combination, synthesis, or analysis, when there is a picture composed of a set of components having different characteristics and meanings. For example, in an RGB input, only the pixel / pixel value of R may be represented, only the pixel / pixel value of G may be represented, or only the pixel / pixel value of B may be represented. For example, only the pixel / pixel value of the Luma component synthesized using the R, G, and B components may be represented. For example, only the pixel / pixel value of an image or information extracted from the R, G, and B components through analysis may be represented.

[0053] In this disclosure, "unit" may represent a basic unit of image processing. A unit may include at least one of a specific area of ​​a picture and information related to that area. A unit may include one luminance block and two chroma (e.g., Cb, Cr) blocks. Depending on the case, the term "unit" may be used interchangeably with terms such as "sample array," "block," or "area." In general, an MxN block may include samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows. In one embodiment, particularly when applied to VCM, a unit may represent a basic unit containing information for performing a specific task.

[0054] In the present disclosure, "current block" may mean one of "current coding block," "current coding unit," "block to be encoded," "block to be decoded," or "block to be processed." When prediction is performed, "current block" may mean "current prediction block" or "block to be predicted." When transformation (inverse transformation) / quantization (inverse quantization) is performed, "current block" may mean "current transformation block" or "block to be transformed." When filtering is performed, "current block" may mean "block to be filtered."

[0055] Additionally, in the present disclosure, "current block" may mean "the chroma block of the current block" unless there is an explicit description of "chroma block." "The chroma block of the current block" may be expressed by including an explicit description of "chroma block," such as "chroma block" or "current chroma block."

[0056] In the present disclosure, " / " and "," may be interpreted as "and / or." For example, "A / B" and "A, B" may be interpreted as "A and / or B." Additionally, "A / B / C" and "A, B, C" may mean "at least one of A, B and / or C."

[0057] In the present disclosure, "or" may be interpreted as "and / or". For example, "A or B" may mean 1) "A" only, 2) "B" only, or 3) "A and B". Alternatively, in the present disclosure, "or" may mean "additionally or alternatively".

[0058] The present disclosure relates to VCM (Video / image coding for machines).

[0059] VCM refers to a compression technology that encodes and decodes parts of source images or video, or information acquired from source images or video, for the purpose of machine vision. In VCM, the targets for encoding and decoding can be referred to as features. Features can refer to information extracted from source images or video based on task objectives, requirements, the surrounding environment, etc. Features may have information forms different from those of the source images or video; accordingly, the compression method and representation format of the features may also differ from those of the video source.

[0060] VCMs can be applied to various fields. For example, in surveillance systems that recognize and track objects or people, VCMs can be used to store or transmit object recognition information. Furthermore, in intelligent transportation or smart traffic systems, VCMs can be used to transmit vehicle location information collected from GPS, sensing information collected from LIDAR, radar, etc., and various vehicle control information to other vehicles or infrastructure. Additionally, in the field of smart cities, VCMs can be used to perform individual tasks for interconnected sensor nodes or devices.

[0061] The present disclosure provides various embodiments relating to feature / feature map coding. Unless otherwise specifically stated, the embodiments of the present disclosure may each be implemented individually or in combination of two or more.

[0062] FIG. 1 is a schematic diagram of a VCM system to which embodiments of the present disclosure can be applied.

[0063] Referring to FIG. 1, the VCM system may include an encoding device (10) and a decoding device (20).

[0064] The encoding device (10) can generate a bitstream by compressing / encoding features / feature maps extracted from source images / videos and transmit the generated bitstream to a decoding device (20) via a storage medium or network. The encoding device (10) may also be referred to as a feature encoding device. In a VCM system, features / feature maps may be generated in each hidden layer of a neural network. The size and number of channels of the generated feature map may vary depending on the type of neural network or the location of the hidden layer. In the present disclosure, the feature map may be referred to as a feature set.

[0065] The encoding device (10) may include a feature acquisition unit (11), an encoding unit (12), and a transmission unit (13).

[0066] The feature acquisition unit (11) can acquire features / feature maps for a source image / video. According to an embodiment, the feature acquisition unit (11) can acquire features / feature maps from an external device, such as a feature extraction network. In this case, the feature acquisition unit (11) performs a feature receiving interface function. Alternatively, the feature acquisition unit (11) may acquire features / feature maps by running a neural network (e.g., CNN, DNN, etc.) with the source image / video as input. In this case, the feature acquisition unit (11) performs a feature extraction network function.

[0067] According to an embodiment, the encoding device (10) may further include a source image generation unit (not shown) for acquiring a source image / video. The source image generation unit may be implemented as an image sensor, a camera module, etc., and may acquire the source image / video through a process of capturing, synthesizing, or generating the image / video. In this case, the generated source image / video may be transmitted to a feature extraction network and used as input data for extracting features / feature maps.

[0068] The encoding unit (12) can encode the feature / feature map acquired by the feature acquisition unit (11). The encoding unit (12) can perform a series of procedures, such as prediction, transformation, and quantization, to increase encoding efficiency. The encoded data (encoded feature / feature map information) can be output in the form of a bitstream. The bitstream containing the encoded feature / feature map information may be referred to as a VCM bitstream.

[0069] The transmission unit (13) can transmit feature / feature map information or data output in the form of a bitstream to a decoding device (20) via a digital storage medium or network in the form of a file or streaming. Here, the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit (13) may include elements for creating a media file having a predetermined file format or elements for transmitting data via a broadcasting / communication network.

[0070] The decoding device (20) can obtain feature / feature map information from the encoding device (10) and restore the feature / feature map based on the obtained information.

[0071] The decoding device (20) may include a receiving unit (21) and a decoding unit (22).

[0072] The receiving unit (21) receives a bitstream from the encoding device (10), obtains feature / feature map information from the received bitstream, and transmits it to the decoding unit (22).

[0073] The decoding unit (22) can decode the feature / feature map based on the acquired feature / feature map information. To increase decoding efficiency, the decoding unit (22) can perform a series of procedures such as inverse quantization, inverse transform, and prediction corresponding to the operation of the encoding unit (14).

[0074] According to an embodiment, the decoding device (20) may further include a task analysis / rendering unit (23).

[0075] The task analysis / rendering unit (23) can perform task analysis based on the decoded feature / feature map. Additionally, the task analysis / rendering unit (23) can render the decoded feature / feature map into a form suitable for task execution. Various machine (oriented) tasks can be performed based on the task analysis results and the rendered feature / feature map.

[0076] In summary, the VCM system can encode / decode features extracted from source images / videos according to user and / or machine requests, task objectives, and surrounding environments, and perform various machine (oriented) tasks based on the decoded features. The VCM system may be implemented by extending / redesigning a video / image coding system and can perform various encoding / decoding methods defined in the VCM standard.

[0077] FIG. 2 is a schematic diagram showing a VCM pipeline structure to which embodiments of the present disclosure can be applied.

[0078] Referring to FIG. 2, the VCM pipeline (200) may include a first pipeline (210) for encoding / decoding images / videos and a second pipeline (220) for encoding / decoding features / feature maps. In the present disclosure, the first pipeline (210) may be referred to as a video codec pipeline, and the second pipeline (220) may be referred to as a feature codec pipeline.

[0079] The first pipeline (210) may include a first stage (video encoder) (211) that encodes an input video and a second stage (video decoder) (212) that decodes the encoded video to generate a restored video. The restored video may be used for human viewing, i.e., for human vision.

[0080] The second pipeline (220) may include a third stage (feature extraction network) (221) for extracting features / feature maps from an input image / video, a fourth stage (VCM encoder) (222) for encoding the extracted features / feature maps, and a fifth stage (VCM decoder) (223) for decoding the encoded features / feature maps to generate restored features / feature maps. The restored features / feature maps may be used for machine (vision) tasks. Here, a machine (vision) task may refer to a task in which images / videos are consumed by a machine. Machine (vision) tasks may be applied to service scenarios such as surveillance, intelligent transportation, smart cities, intelligent industries, intelligent content, etc. According to an embodiment, the restored features / feature maps may also be used for human vision.

[0081] According to an embodiment, the features / feature map encoded in the fourth stage (222) may be transferred to the first stage (221) and used to encode the image / video. In this case, an additional bitstream may be generated based on the encoded features / feature map, and the generated additional bitstream may be transferred to the second stage (222) and used to decode the image / video.

[0082] According to an embodiment, the feature / feature map decoded at the fifth stage (223) can be transferred to the second stage (222) and used to decode an image / video.

[0083] FIG. 2 illustrates a case where the VCM pipeline (200) includes a first pipeline (210) and a second pipeline (220), but this is merely exemplary and the embodiments of the present disclosure are not limited thereto. For example, the VCM pipeline (200) may include only the second pipeline (220), or the second pipeline (220) may be extended to a plurality of feature codec pipelines.

[0084] Meanwhile, in the first pipeline (210), the first stage (211) may be performed by an image / video encoder, and the second stage (212) may be performed by an image / video decoder. Additionally, in the second pipeline (220), the third stage (221) may be performed by a VCM encoder (or a feature / feature map encoder), and the fourth stage (222) may be performed by a VCM decoder (or a feature / feature map decoder). The encoder / decoder structure will be described in detail below.

[0085] FIG. 3 is a schematic diagram of an image / video encoder to which embodiments of the present disclosure can be applied.

[0086] Referring to FIG. 3, the image / video encoder (300) may include an image partitioner (310), a predictor (320), a residual processor (330), an entropy encoder (340), an adder (350), a filter (360), and a memory (370). The predictor (320) may include an inter-predictor (321) and an intra-predictor (322). The residual processor (330) may include a transformer (332), a quantizer (333), a dequantizer (334), and an inverse transformer (335). The residual processing unit (330) may further include a subtractor (331). The addition unit (350) may be referred to as a reconstructor or a reconstructed block generator. The above-described image segmentation unit (310), prediction unit (320), residual processing unit (330), entropy encoding unit (340), addition unit (350), and filtering unit (360) may be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to the embodiment. Additionally, the memory (370) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The above-described hardware component may further include the memory (370) as an internal / external component.

[0087] The image segmentation unit (310) can divide an input image (or picture, picture) input to the image / video encoder (300) into one or more processing units. For example, a processing unit may be referred to as a coding unit (CU). A coding unit may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, a single coding unit may be divided into multiple coding units of a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first and the binary-tree structure and / or ternary structure may be applied later. Alternatively, the binary-tree structure may be applied first. An image / video coding procedure according to the present disclosure may be performed based on a final coding unit that is no longer subdivided. In this case, the maximum coding unit may be used directly as the final coding unit based on coding efficiency according to image characteristics, or, if necessary, the coding unit may be recursively subdivided into lower-depth coding units so that a coding unit of the optimal size is used as the final coding unit. Here, the term "coding procedure" may include procedures such as prediction, transformation, and restoration described below. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may each be subdivided or partitioned from the aforementioned final coding unit.The prediction unit may be a unit of sample prediction, and the transformation unit may be a unit that derives transformation coefficients and / or a unit that derives a residual signal from transformation coefficients.

[0088] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the context. In general, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chroma component. The term "sample" may be used as a counterpart to "pixel" or "pel."

[0089] The image / video encoder (300) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (predicted block, prediction sample array) output from an inter prediction unit (321) or an intra prediction unit (322) from an input image signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (332). In this case, as illustrated, the unit that subtracts the prediction signal (predicted block, prediction sample array) from the input image signal (original block, original sample array) within the image / video encoder (300) may be referred to as a subtraction unit (331). The prediction unit can perform a prediction on a block to be processed (hereinafter referred to as the current block) and generate a predicted block containing prediction samples for the current block. The prediction unit can determine whether intra prediction is applied or inter prediction is applied at the current block or CU level. The prediction unit can generate various information regarding the prediction, such as prediction mode information, and transmit it to the entropy encoding unit (340). The information regarding the prediction can be encoded in the entropy encoding unit (340) and output in the form of a bitstream.

[0090] The intra prediction unit (322) can predict the current block by referencing samples within the current picture. At this time, the referenced samples may be located near the current block or away from it, depending on the prediction mode. In intra prediction, the prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a Planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. The intra prediction unit (322) may determine the prediction mode applied to the current block by using the prediction mode applied to the surrounding blocks.

[0091] The inter prediction unit (321) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference picture indices. Motion information may further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture containing the reference blocks and the reference picture containing the temporal neighboring blocks may be the same or different. Temporal surrounding blocks may be referred to as collocated reference blocks, collocated CUs (colCU), etc., and a reference picture containing temporal surrounding blocks may be referred to as a collocated picture (colPic). For example, the inter prediction unit (321) may construct a list of motion information candidates based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter prediction unit (321) may use the motion information of surrounding blocks as the motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0092] The prediction unit (320) can generate a prediction signal based on various prediction methods. For example, the prediction unit may apply intra prediction or inter prediction for a single block, and may also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, for example, screen content coding (SCC). IBC basically performs prediction within the current picture, but it can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may utilize at least one of the inter prediction techniques described in this disclosure. The palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, sample values ​​within the picture can be signaled based on information regarding the palette table and palette index.

[0093] The prediction signal generated by the prediction unit (320) can be used to generate a restoration signal or to generate a residual signal. The transformation unit (332) can generate transform coefficients by applying a transformation technique to the residual signal. For example, the transformation technique may include at least one of the Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loeve Transform (KLT), Graph-Based Transform (GBT), or Conditionally Non-linear Transform (CNT). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transformation obtained based on generating a prediction signal using all previously reconstructed pixels. Additionally, the transformation process may be applied to a pixel block of the same size in a square, or to a block of variable size that is not square.

[0094] The quantization unit (333) quantizes the transformation coefficients and transmits them to the entropy encoding unit (340), and the entropy encoding unit (340) can encode the quantized signal (information regarding the quantized transformation coefficients) and output it as a bitstream. The information regarding the quantized transformation coefficients may be called residual information. The quantization unit (333) can rearrange the block-shaped quantized transformation coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information regarding the quantized transformation coefficients based on the one-dimensional vector-shaped quantized transformation coefficients. The entropy encoding unit (340) can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit (340) may encode information necessary for image / video restoration (e.g., values ​​of syntax elements) together or separately, in addition to the quantized transform coefficients. The encoded information (e.g., encoded image / video information) may be transmitted or stored in the form of a bitstream in units of network abstraction layer (NAL) units. The image / video information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the image / video information may further include general constraint information. Furthermore, the image / video information may further include methods for generating and using the encoded information, purposes, etc. In this disclosure, information and / or syntax elements transmitted / signaled from the image / video encoder to the image / video decoder may be included in the image / video information.Video information may be encoded through the encoding procedure described above and included in a bitstream. The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) that transmits a signal output from the entropy encoding unit (340) and / or a storage unit (not shown) that stores it may be configured as internal / external elements of the video encoder (300), or the transmission unit may be included in the entropy encoding unit (340).

[0095] The quantized transform coefficients output from the quantization unit (333) can be used to generate a prediction signal. For example, a residual signal (residual block or residual samples) can be restored by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (334) and the inverse transformation unit (335). The adder (350) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter-prediction unit (321) or the intra-prediction unit (322). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The adder (350) may be called a reconstructed unit or a reconstructed block generation unit. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, and can also be used for inter prediction of the next picture after filtering as described below.

[0096] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.

[0097] The filtering unit (360) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (360) can generate a modified restored picture by applying various filtering methods to the restored picture, and can store the modified restored picture in memory (370), specifically in the DPB of memory (370). Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (360) can generate various information regarding filtering and transmit it to the entropy encoding unit (340). The information regarding filtering can be encoded in the entropy encoding unit (340) and output in the form of a bitstream.

[0098] The modified restored picture transmitted to memory (370) can be used as a reference picture in the inter-prediction unit (321). This allows for avoiding prediction mismatches in the encoder and decoder units and improving encoding efficiency.

[0099] The DPB of the memory (370) can store the modified restored picture to be used as a reference picture in the inter-prediction unit (321). The memory (370) can store motion information of blocks from which motion information is derived (or encoded) within the current picture and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (321) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (370) can store restoration samples of restored blocks within the current picture and can transmit the stored restoration samples to the intra-prediction unit (322).

[0100] Meanwhile, the VCM encoder (or feature / feature map encoder) may have a structure identical or similar to the image / video encoder (300) described with reference to FIG. 3 in that it performs a series of procedures such as prediction, transformation, and quantization to encode features / feature maps. However, the VCM encoder differs from the image / video encoder (300) in that it encodes features / feature maps, and accordingly, the names of each unit (or component) (e.g., image segmentation unit (310), etc.) and the specific operation details may differ from those of the image / video encoder (300). The specific operation details of the VCM encoder will be described in detail later.

[0101] FIG. 4 is a schematic diagram of an image / video decoder to which embodiments of the present disclosure can be applied.

[0102] Referring to FIG. 4, the image / video decoder (400) may include an entropy decoder (410), a residual processor (420), a predictor (430), an adder (440), a filter (450), and a memory (460). The predictor (430) may include an inter-predictor (431) and an intra-predictor (432). The residual processor (420) may include a dequantizer (421) and an inverse transformer (422). The aforementioned entropy decoding unit (410), residual processing unit (420), prediction unit (430), addition unit (440), and filtering unit (450) may be configured by a single hardware component (e.g., a decoder chipset or a processor) according to an embodiment. Additionally, the memory (460) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (460) as an internal / external component.

[0103] When a bitstream containing video / image information is input, the video / video decoder (400) can restore the video / video in correspondence with the process in which the video / video information is processed in the video / video encoder (300) of FIG. 3. For example, the video / video decoder (400) can derive units / blocks based on block division information obtained from the bitstream. The video / video decoder (400) can perform decoding using a processing unit applied in the video / video encoder. Accordingly, the processing unit for decoding may be, for example, a coding unit, and the coding unit may be divided according to a quad tree structure, a binary tree structure, and / or a binary tree structure from a coding tree unit or a maximum coding unit. One or more conversion units may be derived from the coding unit. And, the restored video signal decoded and output through the video / video decoder (400) can be played back through a playback device.

[0104] The image / video decoder (400) can receive a signal output from the encoder (300) of FIG. 3 in the form of a bitstream, and the received signal can be decoded through an entropy decoding unit (410). For example, the entropy decoding unit (410) can parse the bitstream to derive information (e.g., image / video information) necessary for image restoration (or picture restoration). The image / video information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the image / video information may further include general constraint information. Furthermore, the image / video information may include a method for generating decoded information, a method for using it, a purpose, etc. The image / video decoder (400) may further decode the picture based on information regarding parameter sets and / or general constraint information. Information and / or syntax elements that are signaled / received can be obtained from the bitstream by decoding through a decoding procedure. For example, the entropy decoding unit (410) can decode information within the bitstream based on coding methods such as exponential chord coding, CAVLC, or CABAC, and output values ​​of syntax elements required for image restoration and quantized values ​​of transform coefficients regarding residuals. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, determine a context model using information on the syntax element to be decoded and decoding information of surrounding and decoding target blocks or information on symbols / bins decoded in the previous step, predict the probability of occurrence of the bin according to the determined context model, and perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin after determining the context model. Among the information decoded in the entropy decoding unit (410), information regarding prediction is provided to the prediction unit (inter prediction unit (432) and intra prediction unit (431)), and the residual value for which entropy decoding was performed in the entropy decoding unit (410), i.e., quantized transformation coefficients and related parameter information, can be input to the residual processing unit (420). The residual processing unit (420) can derive residual signals (residual blocks, residual samples, residual sample array). In addition, among the information decoded in the entropy decoding unit (410), information regarding filtering can be provided to the filtering unit (450). Meanwhile, a receiver (not shown) that receives a signal output from an image / video encoder may be further configured as an internal / external element of an image / video decoder (400), or the receiver may be a component of an entropy decoding unit (410). Meanwhile, the image / video decoder according to the present disclosure may be called an image / video decoding device, and the image / video decoder may be divided into an information decoder (image / video information decoder) and a sample decoder (image / video sample decoder). In this case, the information decoder may include an entropy decoding unit (410), and the sample decoder may include at least one of an inverse quantization unit (321), an inverse transform unit (322), an adder (440), a filtering unit (450), a memory (460), an inter prediction unit (432), and an intra prediction unit (431).

[0105] In the inverse quantization unit (421), the quantized transform coefficients can be inversely quantized to output transform coefficients. The inverse quantization unit (421) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed in the image / video encoder. The inverse quantization unit (321) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.

[0106] In the inverse conversion unit (422), the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).

[0107] The prediction unit (430) can perform a prediction for the current block and generate a predicted block containing prediction samples for the current block. Based on the information regarding the prediction output from the entropy decoding unit (410), the prediction unit can determine whether an intra prediction or an inter prediction is applied to the current block and can determine a specific intra / inter prediction mode.

[0108] The prediction unit (420) can generate a prediction signal based on various prediction methods. For example, the prediction unit may apply intra prediction or inter prediction for a single block, and may also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, for example, screen content coding (SCC). IBC basically performs prediction within the current picture, but it can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. The palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and palette index may be included in the image / video information and signaled.

[0109] The intra prediction unit (431) can predict the current block by referring to samples within the current picture. The referenced samples may be located next to the current block or away from it, depending on the prediction mode. In intra prediction, the prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit (431) may determine the prediction mode applied to the current block by using the prediction mode applied to the surrounding blocks.

[0110] The inter prediction unit (432) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. Motion information may include a motion vector and a reference picture index. Motion information may further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (432) can construct a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-prediction can be performed based on various prediction modes, and information regarding the prediction may include information indicating the mode of inter-prediction for the current block.

[0111] The adder (440) can generate a restoration signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the prediction signal (predicted block, predicted sample array) output from the prediction unit (including the inter prediction unit (432) and / or the intra prediction unit (431)). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the restoration block.

[0112] The addition unit (440) may be called a restoration unit or a restoration block generation unit. The generated restoration signal may be used for intra-predicting the next block to be processed within the current picture, may be output after filtering as described below, or may be used for inter-predicting the next picture.

[0113] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.

[0114] The filtering unit (450) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (450) can generate a modified restored picture by applying various filtering methods to the restored picture, and can transmit the modified restored picture to memory (460), specifically to the DPB of memory (460). Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0115] The (modified) restored picture stored in the DPB of the memory (460) can be used as a reference picture in the inter-prediction unit (432). The memory (460) can store motion information of blocks from which motion information within the current picture has been derived (or decoded) and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (432) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (460) can store restoration samples of blocks restored within the current picture and transmit them to the intra-prediction unit (431).

[0116] Meanwhile, the VCM decoder (or feature / feature map decoder) may have a structure identical or similar to the image / video decoder (400) described above with reference to FIG. 4 in that it performs a series of procedures such as prediction, inverse transformation, and inverse quantization to decode features / feature maps. However, the VCM decoder differs from the image / video decoder (400) in that it targets features / feature maps for decoding, and accordingly, the names of each unit (or component) (e.g., DPB, etc.) and the specific operation details may differ from those of the image / video decoder (400). The operation of the VCM decoder may correspond to the operation of the VCM encoder, and the specific operation details will be described in detail later.

[0117] FIG. 5 is a flowchart schematically illustrating a feature / feature map encoding procedure to which embodiments of the present disclosure can be applied.

[0118] Referring to FIG. 5, the feature / feature map encoding procedure may include a prediction procedure (S510), a residual processing procedure (S520), and an information encoding procedure (S530).

[0119] The prediction procedure (S510) can be performed by the prediction unit (320) described above with reference to FIG. 3.

[0120] Specifically, the intra prediction unit (322) can predict the current block (i.e., the set of feature elements currently being encoded) by referencing feature elements within the current feature / feature map. Intra prediction can be performed based on the spatial similarity of the feature elements constituting the feature / feature map. For example, feature elements included in the same Region of Interest (RoI) within an image / video can be presumed to have similar data distribution characteristics. Accordingly, the intra prediction unit (322) can predict the current block by referencing the previously restored feature elements within the region of interest containing the current block. At this time, the referenced feature elements may be located adjacent to the current block or spaced apart from the current block depending on the prediction mode. Intra prediction modes for feature / feature map encoding may include a plurality of non-directional prediction modes and a plurality of directional prediction modes. Non-directional prediction modes may include prediction modes corresponding, for example, to the DC mode and planner mode of a video encoding procedure. Additionally, directional modes may include prediction modes corresponding, for example, to 33 directional modes or 65 directional modes of a video encoding procedure. However, this is merely an example, and the types and number of intra prediction modes may be set or changed in various ways depending on the embodiment.

[0121] The inter prediction unit (321) can predict the current block based on a reference block (i.e., a set of referenced feature elements) specified by motion information on the reference feature / feature map. Inter prediction can be performed based on the temporal similarity of the feature elements constituting the feature / feature map. For example, temporally consecutive features may have similar data distribution characteristics. Therefore, the inter prediction unit (321) can predict the current block by referring to the restored feature elements of features temporally adjacent to the current feature. At this time, motion information for specifying the referenced feature elements may include motion vectors and reference feature / feature map indices. The motion information may further include information regarding the direction of inter prediction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current feature / feature map and temporal neighboring blocks existing within the reference feature / feature map. A reference feature / feature map containing a reference block and a reference feature / feature map containing temporal surrounding blocks may be the same or different. Temporal surrounding blocks may be referred to as collocated reference blocks, and a reference feature / feature map containing temporal surrounding blocks may be referred to as a collocated feature / feature map. The inter-prediction unit (321) may construct a list of motion information candidates based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference feature / feature map index of the current block.Inter prediction can be performed based on various prediction modes, for example, in the case of skip mode and merge mode, the inter prediction unit (321) can use the motion information of surrounding blocks as the motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted. In the case of motion vector prediction (MVP) mode, the motion vector of surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference. In addition to the intra prediction and inter prediction described above, the prediction unit (320) can generate a prediction signal based on various prediction methods.

[0122] The prediction signal generated by the prediction unit (320) can be used to generate a residual signal (residual block, residual feature elements) (S520). The residual processing procedure (S520) can be performed by the residual processing unit (330) described above with reference to FIG. 3. Then, (quantized) transformation coefficients can be generated through a transformation and / or quantization procedure for the residual signal, and the entropy encoding unit (340) can encode information regarding the (quantized) transformation coefficients as residual information within the bitstream (S530). In addition, the entropy encoding unit (340) can encode information necessary for feature / feature map restoration, such as prediction information (e.g., prediction mode information, motion information, etc.), in addition to the residual information within the bitstream.

[0123] Meanwhile, the feature / feature map encoding procedure may further include a procedure (S530) for encoding information for feature / feature map restoration (e.g., prediction information, resident information, partitioning information, etc.) and outputting it in the form of a bitstream, as well as a procedure for generating a restored feature / feature map for the current feature / feature map and a procedure (optional) for applying in-loop filtering to the restored feature / feature map.

[0124] The VCM encoder can derive (modified) residual feature(s) from quantized transform coefficient(s) through inverse quantization and inverse transform, and can generate a reconstructed feature / feature map based on the predicted feature(s) and (modified) residual feature(s) which are the outputs of step S510. The reconstructed feature / feature map thus generated may be identical to the reconstructed feature / feature map generated by the VCM decoder. If an in-loop filtering procedure is performed on the reconstructed feature / feature map, a modified reconstructed feature / feature map may be generated through the in-loop filtering procedure on the reconstructed feature / feature map. The modified reconstructed feature / feature map is stored in a decoded feature buffer (DFB) or memory and can be used as a reference feature / feature map in a subsequent prediction procedure of the feature / feature map. Additionally, information (parameters) related to (in-loop) filtering may be encoded and output in the form of a bitstream. Through the in-loop filtering procedure, noise that may occur during feature / feature map coding can be removed, and the performance of feature / feature map-based tasks can be improved. Furthermore, by performing the in-loop filtering procedure at both the encoder and decoder stages, the consistency of prediction results can be guaranteed, the reliability of feature / feature map coding can be enhanced, and the amount of data transmitted for feature / feature map coding can be reduced.

[0125] FIG. 6 is a flowchart schematically illustrating a feature / feature map decoding procedure to which embodiments of the present disclosure can be applied.

[0126] Referring to FIG. 6, the feature / feature map decoding procedure may include an image / video information acquisition procedure (S610), a feature / feature map restoration procedure (S620–S640), and an in-loop filtering procedure for the restored feature / feature map (S650). The feature / feature map restoration procedure may be performed based on prediction signals and residual signals obtained through the inter / intra prediction (S620) and residual processing (S630) and the inverse quantization and inverse transformation process for the quantized transformation coefficients described in the present disclosure. A modified restored feature / feature map may be generated through the in-loop filtering procedure for the restored feature / feature map, and the modified restored feature / feature map may be output as a decoded feature / feature map. The decoded feature / feature map can be stored in a decoding feature buffer (DFB) or memory and used as a reference feature / feature map in the inter-prediction procedure during subsequent decoding of the feature / feature map. In some cases, the aforementioned in-loop filtering procedure may be omitted. In this case, the restored feature / feature map can be output as is as the decoded feature / feature map, or stored in a decoding feature buffer (DFB) or memory and used as a reference feature / feature map in the inter-prediction procedure during subsequent decoding of the feature / feature map.

[0127] The SEI message related to the present disclosure will be described below.

[0128] Table 1 shows an example of encoder optimization information (EOI) SEI (Supplemental Enhancement Information) message syntax according to one embodiment.

[0129] [Table 1]

[0130]

[0131] An example of encoder optimization information SEI message semantics according to one embodiment is described.

[0132] Encoder optimization information SEI messages are used to indicate whether the video is optimized for human viewing or machine analysis, and what type of optimization was applied during the preprocessing or encoding process.

[0133] If eoi_cancel_flag is 1, it indicates that the persistence of the encoder optimization information SEI message included in the previous PU in the output order is canceled. If eoi_cancel_flag is 0, it indicates that optimization information applied during the preprocessing or encoding stage follows.

[0134] eoi_persistence_flag indicates the persistence of the optimization information provided in this SEI message. If eoi_persistence_flag is 0, it indicates that the optimization information is applied only to the current picture. If eoi_persistence_flag is 1, it indicates that the optimization information is applied to the current picture and all pictures following the current layer in output order, and persists until one or more of the following conditions are met:

[0135] - Until new CLVS of the current layer start.

[0136] - Until the bitstream ends.

[0137] - Until the picture of the current layer associated with the encoder optimization information SEI message is output following the current picture in the output order.

[0138] If eoi_for_human_viewing_idc is 3, it indicates that human viewing is included in the applied optimization objectives. If eoi_for_human_viewing_idc is 2, it indicates that the video is suitable for human viewing but has not been specifically optimized. If eoi_for_human_viewing_idc is 1, it indicates that the video is unsuitable for human viewing. If eoi_for_human_viewing_idc is 0, it indicates that it is unclear whether the video is suitable for human viewing.

[0139] If eoi_for_machine_analysis_idc is 3, it indicates that machine analysis is included in the applied optimization objectives. If eoi_for_machine_analysis_idc is 2, it indicates that the image is suitable for machine analysis but has not been specifically optimized. If eoi_for_machine_analysis_idc is 1, it indicates that the image is not suitable for machine analysis. If eoi_for_machine_analysis_idc is 0, it indicates that it is unknown whether the image is suitable for machine analysis.

[0140] As a bitstream conformance requirement, the values ​​of eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc must not be 1 at the same time.

[0141] eoi_type indicates the type of optimization method specified in Table 2, and if ( eoi_type & bitMask ) is different from 0, it means that the optimization type with the bitMask value in Table 2 has been applied. If eoi_type is greater than 0 and (eoi_type & bitMask) is 0, it indicates that the optimization type with the bitMask value has not been applied. If eoi_type is 0, it indicates that the optimization determined by the application has been used.

[0142] [Table 2]

[0143]

[0144] The variables EoiObjectBasedFlag, EoiTemporalResamplingFlag, EoiSpatialResamplingFlag, EoiTemporalQualityFlag, EoiSpatialQualityFlag, and EoiPrivacyProtectionFlag each specify whether eoi_type includes the types of object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, and privacy protection optimization, and are derived as follows:

[0145] [Formula 1]

[0146]

[0147] For reference, for example, if a specific top temporal sublayer is encoded with coarse quantization that makes quality variations unpleasant for human viewers but does not degrade machine performance, eoi_for_human_viewing_flag and eoi_for_machine_analaysis_flag can be set to 0 and 1, respectively, and eoi_type can be set to a value that makes EoiTemporalQualityFlag 1.

[0148] When eoi_persistence_flag is 0, EoiTemporalResamplingFlag must be 0 and EoiTemporalQualityFlag must be 0 as bitstream conformity requirements.

[0149] If eoi_object_based_idc exists, it indicates the object-based optimization type specified in Table 3, and if (eoi_object_based_idc & bitMask) is not 0, it means that the object-based optimization type associated with the bitMask value in Table 3 has been applied. If eoi_object_based_idc is greater than 0 and (eoi_object_based_idc & bitMask) is 0, it indicates that the object-based optimization type associated with the bitMask value has not been applied. If eoi_object_based_idc is 0, it indicates that an application-defined type of object-based optimization has been applied. The eoi_object_based_idc value must be in the range from 0 to 7 (inclusive) in a bitstream conforming to the present disclosure. Values ​​of eoi_object_based_idc from 8 to 65,535 (inclusive) are reserved for future use and must not exist in a bitstream conforming to the present disclosure. If the value of eoi_object_based_idc is from 8 to 65,535 (including that value), decoders conforming to this version of the specification must ignore eoi_object_based_idc.

[0150] [Table 3]

[0151]

[0152] If eoi_temporal_resampling_type_flag is 0, it indicates that temporal resampling optimization is a subsampling operation. If eoi_temporal_resampling_type_flag is 1, it indicates that temporal resampling optimization is an upsampling operation.

[0153] If eoi_num_int_pics is greater than 0, it indicates that within the duration of this SEI message, the number of pictures excluded between each pair of coded pictures by the encoding system in the output order (when eoi_temporal_resampling_type_flag is 0) or added between each pair of source pictures for encoding (when eoi_temporal_resampling_type_flag is 1) is constant. When eoi_temporal_resampling_type_flag is 0 and eoi_num_int_pics is greater than 0, eoi_num_int_pics indicates the number of pictures excluded between pairs of coded pictures by the encoding system in the output order. When eoi_temporal_resampling_type_flag is 1 and eoi_num_int_pics is greater than 0, eoi_num_int_pics indicates the number of pictures added between each pair of source pictures by the encoding system for encoding.

[0154] If eoi_num_int_pics is 0, it indicates that the number of pictures excluded between each coded picture pair in the output order by the encoding system (when eoi_temporal_resampling_type_flag is 0) or added between source picture pairs for encoding (when eoi_temporal_resampling_type_flag is 1) during the duration of this SEI message is unknown or fluctuating.

[0155] The value of eoi_num_int_pics must be in the range from 0 to 63 (inclusive).

[0156] If eoi_orig_pic_dimensions_flag is 1, it indicates that the eoi_orig_pic_width and eoi_orig_pic_height syntax elements exist. If eoi_orig_pic_dimensions_flag is 0, it indicates that eoi_orig_pic_width and eoi_orig_pic_height do not exist.

[0157] If eoi_orig_pic_width and eoi_orig_pic_height exist, they represent the width and height of the original source picture in luma samples, respectively.

[0158] If eoi_spatial_resampling_type_flag is 0, it indicates that spatial resampling optimization is a subsampling operation. If eoi_spatial_resampling_type_flag is 1, it indicates that spatial resampling optimization is an upsampling operation.

[0159] If eoi_privacy_protection_method_idc exists, it indicates the method / algorithm used to apply privacy protection optimization as specified in Table 4.

[0160] [Table 4]

[0161]

[0162] If eoi_privacy_info_type exists, and eoi_privacy_info_type is greater than 0 and ( eoi_privacy_info_type & bitMask ) is not equal to 0, it indicates the type of protected information specified in Table 5. This indicates that the information type with the bitMask value in Table 5 is protected. If eoi_privacy_info_type is 0, it indicates that the application-defined information type is protected. In a bitstream conforming to the present disclosure, the eoi_privacy_info_type value must have a range from 0 to 7 (inclusive). Values ​​of eoi_privacy_info_type from 8 to 255 are reserved for future use and must not exist in a bitstream conforming to the present disclosure. If the eoi_privacy_info_type value is from 8 to 255 (inclusive), a decoder conforming to this version of the specification must ignore eoi_privacy_info_type.

[0163] [Table 5]

[0164]

[0165] Conventional encoder optimization information (EOI) SEI message designs include signal information indicating that temporal upsampling has been applied. However, EOI SEI messages do not include signal information indicating the location of the source picture within the bitstream. Consequently, it is very difficult, if not impossible, for a decoder / application to distinguish between the source picture and the picture generated during the temporal upsampling process. This distinction may be useful or necessary when the decoder / application needs to delete / remove a picture from the bitstream, and it is a natural course of action to choose to delete the generated picture instead of the source picture.

[0166] In one embodiment, the following items may be applied individually or in combination.

[0167] If the EOI SEI message includes temporal upsampling optimization as one of the optimization types, it indicates a constraint that the first picture in the decoding order within CLVS must be the original source picture.

[0168] If the EOI SEI message includes temporal upsampling optimization as one of the optimization types, it indicates a constraint that the first picture in the output order within CLVS must be the original source picture.

[0169] Specify a syntax element so that if the value is 1, the first picture in the decoding order of the CLVS is the original source picture, and otherwise (i.e., if the value is 0), it means there is no such indication (i.e., it is not known whether the first picture in the decoding order of the CLVS is the source picture or an added picture).

[0170] Specifies a syntax element, where if the value is 1, it indicates that the picture existing in the same access unit as the SEI message is the original source picture, and otherwise (i.e., if 0), it means that there is no such indication (i.e., it is not known whether the first picture in the decoding order in CLVS is the source picture or the added picture).

[0171] Adds a syntax element representing the maximum TemporalId value of the temporal sublayer to which the original source picture belongs.

[0172] Adds a syntax element indicating whether the original source picture belongs to each temporal sub-layer.

[0173] If the EOI SEI message includes temporal upsampling optimization as one of the optimization types, a constraint is specified that the access unit containing the source picture for the said temporal upsampling optimization must exist.

[0174] An example of encoder optimization information SEI message semantics according to one embodiment is described. Semantics different from the encoder optimization information SEI message semantics described above are described, and descriptions of semantics identical to the encoder optimization information SEI message semantics described above are replaced with descriptions of the encoder optimization information SEI message semantics described above. Specifically, among the encoder optimization information SEI message semantics described above, semantics regarding the original source picture of temporal upsampling are described, and descriptions of other semantics are replaced with descriptions of the encoder optimization information SEI message semantics described above.

[0175] If eoi_num_int_pics is greater than 0, it indicates that within the duration of this SEI message, the number of pictures excluded between each pair of coded pictures by the encoding system in the output order (when eoi_temporal_resampling_type_flag is 0) or added between each pair of source pictures for encoding (when eoi_temporal_resampling_type_flag is 1) is constant. When eoi_temporal_resampling_type_flag is 0 and eoi_num_int_pics is greater than 0, eoi_num_int_pics indicates the number of pictures excluded between pairs of coded pictures by the encoding system in the output order. When eoi_temporal_resampling_type_flag is 1 and eoi_num_int_pics is greater than 0, eoi_num_int_pics indicates the number of pictures added between each pair of source pictures by the encoding system for encoding.

[0176] If eoi_num_int_pics is 0, it indicates that the number of pictures excluded between each coded picture pair in the output order by the encoding system (when eoi_temporal_resampling_type_flag is 0) or added between source picture pairs for encoding (when eoi_temporal_resampling_type_flag is 1) during the duration of this SEI message is unknown or fluctuating.

[0177] The value of eoi_num_int_pics must be in the range from 0 to 63 (inclusive).

[0178] If EoiTemporalResamplingFlag is 1, eoi_temporal_resampling_type_flag is 1, and eoi_num_int_pics is greater than 0, the bitstream conformance requirement is that the first picture in the decoding order within CLVS must be the source picture.

[0179] An example of encoder optimization information SEI message semantics according to one embodiment is described. Semantics different from the encoder optimization information SEI message semantics described above are described, and descriptions of semantics identical to the encoder optimization information SEI message semantics described above are replaced with descriptions of the encoder optimization information SEI message semantics described above. Specifically, among the encoder optimization information SEI message semantics described above, semantics regarding the original source picture of temporal upsampling are described, and descriptions of other semantics are replaced with descriptions of the encoder optimization information SEI message semantics described above.

[0180] If eoi_num_int_pics is greater than 0, it indicates that within the duration of this SEI message, the number of pictures excluded between each pair of coded pictures by the encoding system in the output order (when eoi_temporal_resampling_type_flag is 0) or added between each pair of source pictures for encoding (when eoi_temporal_resampling_type_flag is 1) is constant. When eoi_temporal_resampling_type_flag is 0 and eoi_num_int_pics is greater than 0, eoi_num_int_pics indicates the number of pictures excluded between pairs of coded pictures by the encoding system in the output order. When eoi_temporal_resampling_type_flag is 1 and eoi_num_int_pics is greater than 0, eoi_num_int_pics indicates the number of pictures added between each pair of source pictures by the encoding system for encoding.

[0181] If eoi_num_int_pics is 0, it indicates that the number of pictures excluded between each coded picture pair in the output order by the encoding system (when eoi_temporal_resampling_type_flag is 0) or added between source picture pairs for encoding (when eoi_temporal_resampling_type_flag is 1) during the duration of this SEI message is unknown or fluctuating.

[0182] The value of eoi_num_int_pics must be in the range from 0 to 63 (inclusive).

[0183] If EoiTemporalResamplingFlag is 1, eoi_temporal_resampling_type_flag is 1, and eoi_num_int_pics is greater than 0, the bitstream conformity requirement is that the first picture in the output order within CLVS must be the source picture.

[0184] An example of encoder optimization information SEI message semantics according to one embodiment is described. Semantics different from the encoder optimization information SEI message semantics described above are described, and descriptions of semantics identical to the encoder optimization information SEI message semantics described above are replaced with descriptions of the encoder optimization information SEI message semantics described above. Specifically, among the encoder optimization information SEI message semantics described above, semantics regarding the original source picture of temporal upsampling are described, and descriptions of other semantics are replaced with descriptions of the encoder optimization information SEI message semantics described above.

[0185] [Table 6]

[0186]

[0187] If eoi_clvs_start_pic_is_src_pic_flag is 1, it indicates that the first picture in the decoding order in CLVS is the source picture. If eoi_clvs_start_pic_is_src_pic_flag is 0, it does not provide such an indication.

[0188] An example of encoder optimization information SEI message semantics according to one embodiment is described. Semantics different from the encoder optimization information SEI message semantics described above are described, and descriptions of semantics identical to the encoder optimization information SEI message semantics described above are replaced with descriptions of the encoder optimization information SEI message semantics described above. Specifically, among the encoder optimization information SEI message semantics described above, semantics regarding the original source picture of temporal upsampling are described, and descriptions of other semantics are replaced with descriptions of the encoder optimization information SEI message semantics described above.

[0189] [Table 7]

[0190]

[0191] If eoi_src_pic_flag is 1, it indicates that the picture within the same access unit containing the EOI SEI message is the source picture. If eoi_src_pic_flag is 0, it does not provide such an indication.

[0192] An example of encoder optimization information SEI message semantics according to one embodiment is described. Semantics different from the encoder optimization information SEI message semantics described above are described, and descriptions of semantics identical to the encoder optimization information SEI message semantics described above are replaced with descriptions of the encoder optimization information SEI message semantics described above. Specifically, among the encoder optimization information SEI message semantics described above, semantics regarding the original source picture of temporal upsampling are described, and descriptions of other semantics are replaced with descriptions of the encoder optimization information SEI message semantics described above.

[0193] [Table 8]

[0194]

[0195] If eoi_max_tid_src_pic_present_flag is 1, it indicates that the syntax element eoi_max_tid_src_pic exists in the SEI message. If eoi_max_tid_src_pic_present_flag is 0, it indicates that the syntax element eoi_max_tid_src_pic does not exist in the SEI message. When eoi_max_tid_src_pic_present_flag is 1, there is a bitstream conformance requirement that the source picture and the upsampled picture must not belong to the same time sublayer.

[0196] eoi_max_tid_src_pic, if present, represents the maximum TemporalId value of the temporal sublayer to which the original source picture belongs.

[0197] An example of encoder optimization information SEI message semantics according to one embodiment is described. Semantics different from the encoder optimization information SEI message semantics described above are described, and descriptions of semantics identical to the encoder optimization information SEI message semantics described above are replaced with descriptions of the encoder optimization information SEI message semantics described above. Specifically, among the encoder optimization information SEI message semantics described above, semantics regarding the original source picture of temporal upsampling are described, and descriptions of other semantics are replaced with descriptions of the encoder optimization information SEI message semantics described above.

[0198] [Table 9]

[0199]

[0200] If eoi_sublayer_upsampling_info_present_flag is 1, it indicates that the syntax element eoi_max_sublayer_minus1 and / or eoi_sublayer_upsampled_flag[ i ] is present in the SEI message. If eoi_max_tid_src_pic_present_flag is 0, it indicates that the syntax elements eoi_max_sublayer_minus1 and eoi_sublayer_upsampled_flag[ i ] are not present in the SEI message. When eoi_sublayer_upsampling_info_present_flag is 1, there is a bitstream conformance requirement that the source picture and the upsampled picture must not belong to the same time sublayer.

[0201] The value of eoi_max_sublayers_minus1 plus 1 represents the maximum number of temporal sublayers for eoi_sublayer_upsampled_flag[ i ] to be signaled. If s_max_sublayers_minus1 does not exist, it is considered to be equal to TemporalId. The value of eoi_max_sublayers_minus1 must be equal to or greater than the TemporalId of the EOI SEI message.

[0202] The variable eoiMinTemporalSublayer is set as follows:

[0203] - When eoi_persistence_flag is 1, eoiMinTemporalSublayer is equal to 0.

[0204] - Otherwise, eoiMinTemporalSublayer is equal to eoi_max_sublayers_minus1.

[0205] When eoi_sublayer_upsampled_picture_flag[ i ] exists, if 1 it indicates that the decoded output picture belonging to the i-th temporal sublayer has been upsampled and does not match the original source picture that was not modified. If eoi_sublayer_upsampled_picture_flag[ i ] is 0, it does not provide such an indication. If it does not exist, the value of eoi_sublayer_upsampled_picture_flag[ i ] is considered to be 0.

[0206] For reference, if the TemporalId of an EOI SEI message is greater than 0 and the EOI SEI message persists for one or more pictures with a lower TemporalId, the encoder may repeat the information of the EOI SEI message by including that information in one or more EOI SEI messages with a lower TemporalId to prevent information loss when the pictures of the temporal sublayer are lost or removed.

[0207] An example of encoder optimization information SEI message semantics according to one embodiment is described. Semantics different from the encoder optimization information SEI message semantics described above are described, and descriptions of semantics identical to the encoder optimization information SEI message semantics described above are replaced with descriptions of the encoder optimization information SEI message semantics described above. Specifically, among the encoder optimization information SEI message semantics described above, semantics regarding the original source picture of temporal upsampling are described, and descriptions of other semantics are replaced with descriptions of the encoder optimization information SEI message semantics described above.

[0208] If eoi_num_int_pics is greater than 0, it indicates that within the duration of this SEI message, the number of pictures excluded between each pair of coded pictures by the encoding system in the output order (when eoi_temporal_resampling_type_flag is 0) or added between each pair of source pictures for encoding (when eoi_temporal_resampling_type_flag is 1) is constant. When eoi_temporal_resampling_type_flag is 0 and eoi_num_int_pics is greater than 0, eoi_num_int_pics indicates the number of pictures excluded between pairs of coded pictures by the encoding system in the output order. When eoi_temporal_resampling_type_flag is 1 and eoi_num_int_pics is greater than 0, eoi_num_int_pics indicates the number of pictures added between each pair of source pictures by the encoding system for encoding.

[0209] If eoi_num_int_pics is 0, it indicates that the number of pictures excluded between each coded picture pair in the output order by the encoding system (when eoi_temporal_resampling_type_flag is 0) or added between source picture pairs for encoding (when eoi_temporal_resampling_type_flag is 1) during the duration of this SEI message is unknown or fluctuating.

[0210] The value of eoi_num_int_pics must be in the range from 0 to 63 (inclusive).

[0211] When EoiTemporalResamplingFlag is 1 and eoi_temporal_resampling_type_flag is 1, there is a constraint that access units containing EOI SEI messages must include a source picture.

[0212] For reference, eoi_temporal_int_pics can be used to identify pictures generated from temporal upsampling optimization. The (eoi_temporal_int_pics)th picture from the current picture (i.e., the picture associated with the EOI SEI message) is a picture generated from temporal upsampling optimization.

[0213] The terms or names described below (e.g., names of syntax elements or variables, etc.) are merely examples, and the technical features of the present disclosure are not limited to the terms, etc. described below. For example, the image information described below may include various information according to the embodiments described in the present disclosure and may include information described in at least one of the tables described above.

[0214] The operations described below do not constitute an essential component of one embodiment, and at least some of the operations described below may be omitted. Furthermore, the operations described below do not constitute a sufficient component of one embodiment, and previously described operations may be added. Moreover, unless they contradict previously described operations, the operations described below form one embodiment integrally with previously described operations and do not form a separate embodiment distinct from previously described operations.

[0215] FIG. 7 is a diagram illustrating a method for decoding image information according to one embodiment of the present disclosure.

[0216] Terms or names (e.g., names of syntax elements or names of variables, etc.) are merely examples, and the technical features of the present disclosure are not limited to the terms, etc. described in FIG. 7. For example, the image information described in FIG. 7 may include various information according to the embodiments described in the present disclosure and may include information described in at least one of the tables described above.

[0217] The decoding method (S700) may include operations described below. The operations described below do not constitute an essential component of the decoding method according to one embodiment, and at least some of the operations described below may be omitted. Furthermore, the operations described below do not constitute a sufficient component of the decoding method according to one embodiment, and the previously described operations may be added. Moreover, unless the operations described below contradict the previously described operations, they form an embodiment integrally with the previously described operations and do not form a separate embodiment distinct from the previously described operations.

[0218] The decoding method (S700) can be executed by a decoding device including a memory and a processor electrically connected to the memory, for example, by a processor.

[0219] The decoding device can acquire image information (S710).

[0220] For example, a processor of a decoding device may acquire image information. The image information may include SEI (supplemental enhancement information) messages. SEI messages may convey specific types of information that assist in processes related to the decoding, display, or other purposes of the image information. Here, SEI messages may not be necessary for the decoding process to determine the sample values ​​of the decoded picture.

[0221] For example, the SEI message may include Encoder optimization information (EOI) SEI (Supplemental Enhancement Information) messages.

[0222] Encoder optimization information SEI messages may include information regarding SEI message handling and information regarding encoder optimization. For example, an encoder optimization information SEI message may provide encoder optimization information indicating whether the video is optimized for human viewing or machine analysis, and what type of optimization was applied during the preprocessing or encoding process. For example, encoder optimization may include optimization for human viewing, optimization for machine analysis, object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, or privacy optimization.

[0223] Encoder optimization information SEI messages can have various names. For example, an encoder optimization information SEI message may be referred to as an EOI SEI message, an EOI SEI related message, EOI SEI information, EOI SEI related information, encoder optimization information message, EOI message, EOI related message, EOI information, EOI related information, etc.

[0224] The encoder optimization information SEI message can take various forms. For example, the encoder optimization information SEI message may be a syntax element or a syntax structure containing one or more syntax elements. Additionally, the encoder optimization information SEI message may be a raw byte sequence payload (RBSP) containing one or more syntax elements or one or more syntax structures. For example, the encoder optimization information SEI message may be represented as encoder_optimization_info(payloadSize), but is not limited thereto.

[0225] The encoder optimization information SEI message may include cancellation information of the encoder optimization information SEI message, persistence information of the encoder optimization information SEI message, human viewing optimization information, machine analysis optimization information, optimization type information, object-based optimization type information, temporal resampling type information, picture count information, source picture information, information regarding spatial resampling type, privacy protection method type information and / or privacy type information.

[0226] Human viewing optimization information can indicate whether the optimization purpose involves human viewing. In other words, human viewing optimization information can indicate whether the optimized image is optimized for human viewing, suitable, unsuitable, or unknown.

[0227] Human viewing optimization information may take various forms and may be expressed by various names. For example, human viewing optimization information may be a syntax element or a syntax structure containing one or more syntax elements. For example, human viewing optimization information that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits or variable-length bits. Human viewing optimization information that is a syntax element may be expressed as eoi_for_human_viewing_idc or eoi_for_human_viewing_flag, but is not limited thereto.

[0228] Machine analysis optimization information can indicate whether the optimization objective involves machine analysis. In other words, machine analysis optimization information can indicate whether the optimized image is optimized for machine analysis, is suitable, unsuitable, or unknown.

[0229] Machine analysis optimization information may take various forms and may be represented by various names. For example, machine analysis optimization information may be a syntax element or a syntax structure containing one or more syntax elements. For example, machine analysis optimization information that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits or variable-length bits. Machine analysis optimization information that is a syntax element may be represented as eoi_for_machine_analysis_idc or eoi_for_machine_analysis_flag, but is not limited thereto.

[0230] Optimization type information indicates the type of optimization applied. Encoder optimization may include object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, and / or privacy optimization. One or more of object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, and privacy optimization may be applied, and optimization type information may indicate the type of all applied optimizations.

[0231] Optimization type information may take various forms and may be represented by various names. For example, optimization type information may be a syntax element or a syntax structure containing one or more syntax elements. For example, optimization type information that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits (e.g., 16 bits) or a variable-length bit. Optimization type information that is a syntax element may be represented as eoi_type, eoi_type_flag, or eoi_type_idc, but is not limited thereto.

[0232] Based on optimization type information, multiple optimization type variables can be derived to indicate whether object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, and privacy protection optimization are applied, respectively. For example, the multiple optimization type variables may include, but are not limited to, EoiObjectBasedFlag, EoiTemporalResamplingFlag, EoiSpatialResamplingFlag, EoiTemporalQualityFlag, EoiSpatialQualityFlag, and EoiPrivacyProtectionFlag.

[0233] Object-based optimization type information is obtained when the optimization type indicated by the optimization type information includes object-based optimization, and may indicate the type of object-based optimization including blurring, quantization, and / or overwriting.

[0234] Object-based optimization type information may take various forms and may be represented by various names. For example, object-based optimization type information may be a syntax element or a syntax structure containing one or more syntax elements. For example, object-based optimization type information that is a syntax element may include multiple flags consisting of one bit or an indicator consisting of variable-length bits. Optimization type information that is a syntax element may be represented as eoi_object_based_flag or eoi_object_based_idc, but is not limited thereto.

[0235] Temporal resampling type information is obtained when the optimization type indicated by the optimization type information includes temporal resampling, and may indicate the type of temporal resampling. Temporal resampling may include temporal upsampling and temporal subsimplification. Temporal resampling type information may indicate whether the applied temporal resampling is temporal upsampling or temporal subsimplification. For example, temporal resampling type information with a value of 0 indicates that the applied temporal resampling is temporal subsampling. Temporal resampling type information with a value of 1 indicates that the applied temporal resampling is temporal upsampling. However, this is not limited thereto, and what is indicated by temporal resampling type information with a value of 1 may be interchangeable with what is indicated by temporal resampling type information with a value of 0.

[0236] Temporal resampling type information may take various forms and may be represented by various names. For example, temporal resampling type information may be a syntax element or a syntax structure containing one or more syntax elements. For example, temporal resampling type information that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits or variable-length bits. Temporal resampling type information that is a syntax element may be represented as eoi_temporal_resampling_type_flag or eoi_temporal_resampling_type_idc, but is not limited thereto.

[0237] Picture count information is obtained when the optimization type indicated by the optimization type information includes temporal resampling, and indicates whether there are pictures added between a pair of temporally adjacent original source pictures or excluded between a pair of temporally adjacent decoded pictures by temporal resampling, and may also indicate the number of pictures added between a pair of temporally adjacent original source pictures by temporal upsampling or the number of pictures excluded between a pair of temporally adjacent decoded pictures by temporal subsampling. For example, picture count information with a value equal to 0 may indicate that there are no pictures added or excluded by temporal resampling. Additionally, picture count information with a value greater than 0 may indicate the number of pictures added by temporal upsampling or the number of pictures excluded by temporal subsampling.

[0238] Picture count information may take various forms and may be represented by various names. For example, picture count information may be a syntax element or a syntax structure containing one or more syntax elements. For example, picture count information that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits or a variable length bit. Picture count information that is a syntax element may be represented as eoi_num_int_pics, etc., but is not limited thereto.

[0239] Source picture information is obtained when the optimization type indicated by the optimization type information includes temporal resampling, and can indicate whether the current picture is the original source picture.

[0240] Source picture information can provide information regarding whether the current picture is the original source picture or a picture added by temporal upsampling when a picture is added by temporal upsampling.

[0241] When a picture is added through temporal upsampling, information regarding whether the current picture is the original source picture or a picture added through temporal upsampling can be provided in various ways.

[0242] For example, source picture information may indicate whether a picture contained in the same picture unit or access unit as the encoder optimization information SEI message is the original source picture. For example, source picture information with a value of 1 may indicate that a picture contained in the same picture unit or access unit as the encoder optimization information SEI message is the original source picture. Additionally, source picture information with a value of 0 may indicate that no information regarding the original source picture is provided.

[0243] Original source pictures can be identified based on source picture information. For example, if one picture is added between a pair of temporally adjacent original source pictures by temporal upsampling, the picture included in the same picture unit or access unit as the encoder optimization information SEI message is the original source picture, the picture following the original source picture is the added picture, and the picture following the added picture is the original source picture. For example, if two pictures are added between a pair of temporally adjacent original source pictures by temporal upsampling, the picture included in the same picture unit or access unit as the encoder optimization information SEI message is the original source picture, the two consecutive pictures following the original source picture are the added pictures, and the picture following the added pictures is the original source picture.

[0244] As another example, source picture information can indicate whether the first picture in the decoding order within the CLVS is the source picture. Alternatively, the first picture in the decoding order within the CLVS may be restricted to being the source picture.

[0245] As another example, source picture information can indicate whether the first picture in the output order within the CLVS is the source picture. Alternatively, the first picture in the output order within the CLVS may be restricted to being the source picture.

[0246] As another example, source picture information can represent the maximum TemporalId value of the temporal sublayer to which the original source picture belongs.

[0247] As another example, source picture information may represent a temporal sublayer containing a picture added by upsampling of the encoding device.

[0248] As another example, source picture information can indicate the number of pictures up to the picture added by the upsampling of the encoding device.

[0249] Since source picture information provides information regarding whether the current picture is the original source picture when a picture is added by temporal upsampling, it is efficient to acquire source picture information only when a picture is added by temporal upsampling. For example, source picture information can be acquired if the optimization type indicated by the optimization type information includes temporal resampling and the picture count information indicates that there are pictures added or excluded by temporal resampling.

[0250] Source picture information may take various forms and may be represented by various names. For example, source picture information may be a syntax element or a syntax structure containing one or more syntax elements. For example, source picture information that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits or a variable length bit. Source picture information that is a syntax element may be represented as eoi_clvs_start_pic_is_src_pic_flag, eoi_src_pic_flag, eoi_max_tid_src_pic, eoi_sublayer_upsampled_picture_flag[ i ] or eoi_temporal_int_pics, but is not limited thereto.

[0251] Information regarding the spatial resampling type is obtained when the optimization type indicated by the optimization type information includes spatial resampling, and can indicate the type of spatial resampling. Spatial resampling includes resampling in the horizontal direction (width direction) and resampling in the vertical direction (height direction) in terms of space, and may include spatial upsampling and spatial subsampling in terms of sampling. In other words, spatial resampling may include upsampling / subsampling in the horizontal direction (width direction) and upsampling / subsampling in the vertical direction (height direction).

[0252] Information regarding spatial resampling types can be provided in various ways. For example, information regarding horizontal (width direction) upsampling / subsampling and vertical (height direction) upsampling / subsampling can be provided by providing information regarding the width and height of the original source picture. Additionally, resampling information indicating horizontal (width direction) upsampling / subsampling and resampling information indicating vertical (height direction) upsampling / subsampling can be provided directly.

[0253] For example, information regarding the spatial resampling type may include information providing the spatial resampling type (e.g., eoi_orig_pic_dimensions_flag), information regarding the width and height of the original source picture (e.g., eoi_orig_pic_width and eoi_orig_pic_height), and / or information regarding the spatial resampling type (e.g., eoi_spatial_resampling_type_flag).

[0254] Information regarding spatial resampling type may take various forms and may be expressed by various names. For example, information regarding spatial resampling type may be a syntax element or a syntax structure containing one or more syntax elements. For example, information regarding spatial resampling type that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits or variable-length bits. Information regarding spatial resampling type that is a syntax element may include syntax elements expressed as eoi_orig_pic_dimensions_flag, eoi_orig_pic_width, eoi_orig_pic_height, and eoi_spatial_resampling_type_flag, but is not limited thereto.

[0255] Information on the types of privacy protection methods may indicate the methods / algorithms used to apply privacy protection optimization. The methods / algorithms used to apply privacy protection optimization may include delegation, blurring, substitution, masking, etc., in the application.

[0256] Information on the type of privacy protection method may take various forms and may be expressed by various names. For example, information on the type of privacy protection method may be a syntax element or a syntax structure containing one or more syntax elements. For example, information on the type of privacy protection method that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits or variable-length bits. Information on the type of privacy protection method that is a syntax element may be expressed as eoi_privacy_protection_method_flag or eoi_privacy_protection_method_idc, but is not limited thereto.

[0257] The personal information type information may indicate the types of personal information protected by encoding optimization. The types of personal information protected by encoding optimization may include an individual's face, vehicle license plates, location information, etc.

[0258] Personal information type information may take various forms and may be expressed by various names. For example, personal information type information may be a syntax element or a syntax structure containing one or more syntax elements. For example, personal information type information that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits or variable-length bits. Personal information type information that is a syntax element may be expressed as eoi_privacy_info_type, eoi_privacy_info_type_flag, or eoi_privacy_info_type_idc, but is not limited thereto.

[0259] The decoding device can derive encoder optimization information (S720).

[0260] For example, the processor of a decoding device can process encoder optimization information SEI messages. Based on the processing of the encoder optimization information SEI messages, the decoding device can derive information regarding encoder optimization.

[0261] The encoder optimization information SEI message may include information regarding temporal resampling. Additionally, if the temporal resampling represents temporal upsampling, the encoder optimization information SEI message may further include source picture information to identify the original source picture from the picture added by the temporal upsampling.

[0262] As such, identifying the original source picture can provide various technical benefits. The original source picture may contain image information that is more similar to the original picture when compared to the picture added by the upsampling of the encoding device. Therefore, if subsampling is required by the decoding device, the decoding device can remove the picture added by the upsampling of the encoding device and retain the original source picture. By doing so, the decoding device can provide the user with image information that is more similar to the original picture.

[0263] In this way, providing information regarding whether the current picture is the original source picture using source picture information can improve the reliability of image information transmission by the coding system by enabling the coding system to provide information that is more similar to the image information of the original picture.

[0264] Providing information on whether the current picture is the original source picture using source picture information can improve the coding speed and coding efficiency of the coding system by enabling the coding system to identify the original source picture more quickly.

[0265] Providing information on whether the current picture is the original source picture using source picture information can improve the data transmission efficiency of the coding system by enabling the coding system to identify the original source picture using less data.

[0266] FIG. 8 is a diagram illustrating a method for encoding image information according to one embodiment of the present disclosure.

[0267] The terms or names described in FIG. 8 (e.g., names of syntax elements or variables, etc.) are merely examples, and the technical features of the present disclosure are not limited to the terms, etc. described in FIG. 8. For example, the image information described in FIG. 8 may include various information according to the embodiments described in the present disclosure and may include information described in at least one of the tables described above.

[0268] The encoding method (S800) may include operations described below. The operations described below do not constitute an essential component of the decoding method according to one embodiment, and at least some of the operations described below may be omitted. Furthermore, the operations described below do not constitute a sufficient component of the encoding method according to one embodiment, and the previously described operations may be added. Moreover, unless the operations described below contradict the previously described operations, they form an embodiment integrally with the previously described operations and do not form a separate embodiment distinct from the previously described operations.

[0269] The encoding method (S800) may be executed by an encoding device including a memory and a processor electrically connected to the memory, for example, by a processor.

[0270] The encoding device can perform encoder optimization (S810).

[0271] For example, the processor of the encoding device can perform encoder optimization. For example, encoder optimization may include optimization for human viewing, optimization for machine analysis, object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, or privacy optimization.

[0272] The encoding device can encode video information (S820).

[0273] For example, the processor of the encoding device can generate information about encoder optimization based on the performed encoder optimization, generate an encoder optimization information (EOI) SEI (supplemental enhancement information) message based on the generated information about encoder optimization, and encode image information including the encoder optimization information SEI message.

[0274] The encoder optimization information SEI message may include information regarding SEI message handling and information regarding encoder optimization. For example, the encoder optimization information SEI message may provide encoder optimization information indicating whether the video is optimized for human viewing or machine analysis, and what type of optimization was applied during the preprocessing or encoding process.

[0275] Encoder optimization information SEI messages can have various names. For example, an encoder optimization information SEI message may be referred to as an EOI SEI message, an EOI SEI related message, EOI SEI information, EOI SEI related information, encoder optimization information message, EOI message, EOI related message, EOI information, EOI related information, etc.

[0276] The encoder optimization information SEI message can take various forms. For example, the encoder optimization information SEI message may be a syntax element or a syntax structure containing one or more syntax elements. Additionally, the encoder optimization information SEI message may be a raw byte sequence payload (RBSP) containing one or more syntax elements or one or more syntax structures. For example, the encoder optimization information SEI message may be represented as encoder_optimization_info(payloadSize), but is not limited thereto.

[0277] The encoder optimization information SEI message may include cancellation information of the encoder optimization information SEI message, persistence information of the encoder optimization information SEI message, human viewing optimization information, machine analysis optimization information, optimization type information, object-based optimization type information, temporal resampling type information, picture count information, source picture information, information regarding spatial resampling type, privacy protection method type information and / or privacy type information.

[0278] The human viewing optimization information, machine analysis optimization information, optimization type information, object-based optimization type information, temporal resampling type information, picture count information, source picture information, information regarding spatial resampling type, privacy protection method type information and / or privacy type information may be the same as the human viewing optimization information, machine analysis optimization information, optimization type information, object-based optimization type information, temporal resampling type information, picture count information, source picture information, information regarding spatial resampling type, privacy protection method type information and / or privacy type information described above in operation 710 of FIG. 7.

[0279] The descriptions of human viewing optimization information, machine analysis optimization information, optimization type information, object-based optimization type information, temporal resampling type information, picture count information, source picture information, information regarding spatial resampling type, privacy protection method type information and / or privacy type information are replaced with the descriptions of human viewing optimization information, machine analysis optimization information, optimization type information, object-based optimization type information, temporal resampling type information, picture count information, source picture information, information regarding spatial resampling type, privacy protection method type information and / or privacy type information described above in operation 710 of FIG. 7.

[0280] In this way, the encoding device can encode video information including encoder optimization information SEI messages.

[0281] Here, the encoder optimization information SEI message may include information regarding temporal resampling. Additionally, if the temporal resampling represents temporal upsampling, the encoder optimization information SEI message may further include source picture information for identifying the original source picture from the picture added by the temporal upsampling.

[0282] As such, identifying the original source picture can provide various technical benefits. The original source picture may contain image information that is more similar to the original picture when compared to the picture added by the upsampling of the encoding device. Therefore, if subsampling is required by the decoding device, the decoding device can remove the picture added by the upsampling of the encoding device and retain the original source picture. By doing so, the decoding device can provide the user with image information that is more similar to the original picture.

[0283] In this way, providing information regarding whether the current picture is the original source picture using source picture information can improve the reliability of image information transmission by the coding system by enabling the coding system to provide information that is more similar to the image information of the original picture.

[0284] Providing information on whether the current picture is the original source picture using source picture information can improve the coding speed and coding efficiency of the coding system by enabling the coding system to identify the original source picture more quickly.

[0285] Providing information on whether the current picture is the original source picture using source picture information can improve the data transmission efficiency of the coding system by enabling the coding system to identify the original source picture using less data.

[0286] A bitstream is generated based on video information encoded according to the encoding method (S800) described above, and the bitstream can be stored on a computer-readable storage medium.

[0287] In addition, a bitstream is generated based on video information encoded according to the encoding method (S800) described above, and the bitstream can be transmitted through a transmission unit and / or a transmission medium.

[0288] In this disclosure, as an example, the names of the syntax elements described above are all arbitrarily designated for clarity of explanation and are not intended to limit the names of the syntax elements. Additionally, each syntax element may be referred to as information. Furthermore, syntax elements may be obtained from a bitstream, but may also be derived from other syntax elements, and this may also be included in the embodiments of this disclosure.

[0289] In addition, the bitstream generated by the video encoding method may be stored on a non-transient computer-readable recording medium.

[0290] In addition, as another example, a bitstream generated by a video encoding method may be transmitted to another device (e.g., a video decoding device, etc.). In this case, the method of transmitting the bitstream may include a process of transmitting the bitstream.

[0291] The exemplary methods of the present disclosure are described as a series of operations for clarity of description, but this is not intended to limit the order in which the steps are performed, and if necessary, each step may be performed simultaneously or in a different order. To implement the method according to the present disclosure, additional steps may be included in addition to the steps exemplified, steps excluding some steps and including the remaining steps, or steps excluding some steps and including additional steps.

[0292] In the present disclosure, an image encoding device or an image decoding device performing a predetermined operation (step) may perform an operation (step) to check the conditions or circumstances for performing the said operation (step). For example, if it is stated that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or the image decoding device may perform an operation to check whether the said predetermined condition is satisfied, and then perform the said predetermined operation.

[0293] The various embodiments of the present disclosure are not intended to list all possible combinations but to describe representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combination of two or more.

[0294] The embodiments described in this disclosure may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each figure may be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for implementation may be stored in a digital storage medium.

[0295] In addition, the decoder (decoding device) and encoder (encoding device) to which the embodiment(s) of the present disclosure are applied may be included in multimedia broadcasting transmission and reception devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video conversation devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, Video on Demand (VoD) service providers, OTT video (Over the top video) devices, internet streaming service providers, 3D video devices, VR (virtual reality) devices, AR (argument reality) devices, video phone video devices, transportation terminals (e.g., vehicle (including autonomous vehicle) terminals, robot terminals, airplane terminals, ship terminals, etc.), and medical video devices, and may be used to process video signals or data signals. For example, OTT (Over the top video) devices may include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, DVRs (Digital Video Recorders), etc.

[0296] Additionally, the processing method to which the embodiment(s) of the present disclosure are applied may be produced in the form of a program that is executed by a computer and may be stored on a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of this document may also be stored on a computer-readable recording medium. A computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. A computer-readable recording medium may include, for example, a Blu-ray disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. Additionally, a computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission over the Internet). Additionally, a bitstream generated by an encoding method may be stored on a computer-readable recording medium or transmitted via a wired or wireless communication network.

[0297] Additionally, the embodiment(s) of the present disclosure may be implemented as a computer program product by program code, and the program code may be executed on a computer by the embodiment(s) of the present disclosure. The program code may be stored on a carrier that is readable by a computer.

[0298] FIG. 9 is a drawing showing an example of a content streaming system to which embodiments of the present disclosure can be applied.

[0299] Referring to FIG. 9, a content streaming system to which an embodiment of the present disclosure is applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0300] The encoding server compresses content input from multimedia input devices, such as smartphones, cameras, and camcorders, into digital data to generate a bitstream, and transmits it to the streaming server. As another example, if multimedia input devices, such as smartphones, cameras, and camcorders, generate the bitstream directly, the encoding server can be omitted.

[0301] A bitstream may be generated by a video encoding method and / or video encoding device to which an embodiment of the present disclosure is applied, and a streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0302] A streaming server transmits multimedia data to a user device based on user requests made through a web server, and the web server can act as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server forwards the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which can perform the role of controlling commands and responses between devices within the content streaming system.

[0303] A streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a seamless streaming service, the streaming server may store the bitstream for a certain period of time.

[0304] Examples of user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs (head-mounted displays)), digital TVs, desktop computers, digital signage, etc.

[0305] Each server within the content streaming system can be operated as a distributed server, and in this case, data received from each server can be processed in a distributed manner.

[0306] FIG. 10 is a drawing showing another example of a content streaming system to which embodiments of the present disclosure may be applied.

[0307] Referring to FIG. 10, in an embodiment such as VCM, a task may be performed at a user terminal or at an external device (e.g., a streaming server, an analysis server, etc.) depending on the performance of the device, the user's request, the characteristics of the task to be performed, etc. In this way, in order to transmit information necessary for performing the task to an external device, the user terminal may generate a bitstream containing information necessary for performing the task (e.g., information such as the task, the neural network, and / or the purpose) directly or through an encoding server.

[0308] The analysis server can decode encoded information received from the user terminal (or from the encoding server) and then execute the task requested by the user terminal. The analysis server can transmit the results obtained through task execution back to the user terminal or to other associated service servers (e.g., web servers). For example, the analysis server can transmit the results obtained from performing a task to identify a fire to a fire-related server. The analysis server may include a separate control server, in which case the control server can play a role in controlling commands and responses between each device and server associated with the analysis server. Additionally, the analysis server may request desired information from the web server based on information regarding tasks that the user device intends to perform and tasks that it can perform. When the analysis server requests a desired service from the web server, the web server forwards it to the analysis server, and the analysis server can transmit the corresponding data to the user terminal. In this case, the control server of the content streaming system can perform the role of controlling commands and responses between each device within the streaming system.

[0309] An embodiment according to the present disclosure can be used to encode / decode video.

Claims

1. Regarding the method of decoding video information, Acquiring the above image information including encoder optimization information (EOI) and SEI (supplemental enhancement information) messages; Including deriving information regarding encoder optimization based on the above encoder optimization information SEI message, The above encoder optimization information SEI message includes optimization type information indicating an optimization type that includes temporal resampling, and A method for determining whether the current picture included in the same unit as the encoder optimization information SEI message is the source picture for the temporal resampling, based on the optimization type information above.

2. In Paragraph 1, A method in which the above encoder optimization information SEI message further includes source picture information regarding the current picture included in the same unit as the above encoder optimization information SEI message being the source picture for the above temporal resampling.

3. In Paragraph 2, A method in which the above encoder optimization information SEI message further includes temporal resampling type information regarding the above temporal upsampling and picture count information regarding the number of pictures added or excluded by the above temporal resampling.

4. In Paragraph 3, A method for obtaining the above source picture information based on the above temporal resampling type information and the above picture count information.

5. In Paragraph 4, A method for obtaining the above source picture information based on temporal resampling type information indicating the above temporal upsampling and picture count information with a value greater than 0.

6. In Paragraph 3, A method for obtaining the above temporal resampling type information and the above picture count information based on optimization type information representing the above temporal resampling.

7. Regarding the method of encoding video information, Perform encoder optimization; The method includes encoding the image information including an Encoder Optimization Information (EOI) SEI (supplemental enhancement information) message generated based on the information regarding the encoder optimization above, wherein The above encoder optimization information SEI message includes optimization type information indicating an optimization type that includes temporal resampling, and A method for determining, based on the above optimization type information, whether the current picture included in the same unit as the above encoder optimization information SEI message is the source picture for the above temporal resampling.

8. In Paragraph 7, A method in which the above encoder optimization information SEI message further includes source picture information regarding the current picture included in the same unit as the above encoder optimization information SEI message being the source picture for the above temporal resampling.

9. In Paragraph 8, A method in which the above encoder optimization information SEI message further includes temporal resampling type information regarding the above temporal upsampling and picture count information regarding the number of pictures added or excluded by the above temporal resampling.

10. In Paragraph 9, A method for obtaining the above source picture information based on the above temporal resampling type information and the above picture count information.

11. In Paragraph 10, A method for obtaining the above source picture information based on temporal resampling type information indicating the above temporal upsampling and picture count information with a value greater than 0.

12. In Paragraph 9, A method for obtaining the above temporal resampling type information and the above picture count information based on optimization type information representing the above temporal resampling.

13. Regarding methods concerning bitstreams, Perform encoder optimization; Generating a bitstream for the image information including an Encoder optimization information (EOI) SEI (supplemental enhancement information) message generated based on the information regarding the encoder optimization; Includes transmitting data for the above bitstream, The above encoder optimization information SEI message includes optimization type information indicating an optimization type that includes temporal resampling, and A method for determining, based on the above optimization type information, whether the current picture included in the same unit as the above encoder optimization information SEI message is the source picture for the above temporal resampling.

14. In a computer-readable storage medium for storing a bitstream, The above storage medium stores a bitstream for the image information including an Encoder Optimization Information (EOI) SEI (supplemental enhancement information) message generated based on information regarding encoder optimization, and The above encoder optimization information SEI message includes optimization type information indicating an optimization type that includes temporal resampling, and A storage medium in which, based on the optimization type information above, it is determined whether the current picture included in the same unit as the encoder optimization information SEI message is the source picture for the temporal resampling.