Image decoding method, image encoding method, computer-readable storage medium, and data transmission method for image
By managing multiple EOI SEI messages with distinct identifiers and persistence information, the method enhances video compression efficiency and reduces transmission costs for high-resolution video data.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2025-10-17
- Publication Date
- 2026-04-23
AI Technical Summary
The increasing demand for high-resolution, high-quality video leads to higher transmission and storage costs due to the increase in transmitted information or bits, necessitating high-efficiency video compression technology to manage and process large amounts of data effectively.
The method involves identifying and managing multiple Encoder Optimization Information (EOI) supplemental enhancement information (SEI) messages with distinct EOI identifiers and persistence information to improve data transmission and coding efficiency by preventing confusion and malfunctions in the coding system.
This approach allows for efficient processing of multiple EOI SEI messages with different identifiers and persistence, preventing repeated transmission of the same information and improving processing stability and clarity in video decoding.
Smart Images

Figure KR2025016505_23042026_PF_FP_ABST
Abstract
Description
Video decoding method, video encoding method, computer-readable storage medium and method for transmitting data for video
[0001] The present disclosure relates to a method for decoding / encoding image information, a computer-readable storage medium for storing image information, and a method for transmitting image information.
[0002] Recently, the demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition), has been increasing across various fields. As video data becomes higher in resolution and quality, the relative amount of information or bits transmitted increases compared to conventional video data. This increase in transmitted information or bits leads to higher transmission and storage costs.
[0003] Accordingly, high-efficiency video compression technology is required to effectively transmit, store, and play back high-resolution, high-quality video information.
[0004] The present disclosure aims to prevent confusion and malfunction in a coding system by identifying encoder optimization information SEI messages through EOI identifiers, thereby independently defining the scope of application and persistence of each encoder optimization information.
[0005] The present disclosure aims to improve the reliability of a coding system by enabling a decoding device to efficiently determine the association with a previous SEI message of the same identifier by referring to an EOI identifier.
[0006] The present disclosure aims to improve data transmission efficiency and coding efficiency by managing multiple encoder optimization information SEI messages having different EOI identifiers and persistence information in parallel.
[0007] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure belongs from the description below.
[0008] According to one aspect of the present disclosure, a method for decoding image information comprises the steps of obtaining an Encoder Optimization Information (EOI) supplemental enhancement information (SEI) message associated with the optimization of a picture from a bitstream, and obtaining the Encoder Optimization Information based on the EOI SEI message, wherein the EOI SEI message comprises an EOI identifier for identifying the EOI SEI message, EOI cancellation information associated with the persistence of the EOI SEI message, and EOI persistence information.
[0009] According to one aspect of the present disclosure, an apparatus for decoding image information obtains an Encoder Optimization Information (SEI) supplemental enhancement information message associated with the optimization of a picture from a bitstream, obtains the Encoder Optimization Information based on the EOI SEI message, and the EOI SEI message comprises an EOI identifier for identifying the EOI SEI message, EOI cancellation information associated with the persistence of the EOI SEI message, and EOI persistence information.
[0010] A method or device for decoding the above image information may be characterized by the simultaneous existence of two or more EOI SEI messages corresponding to different EOI identifiers.
[0011] In a method or device for decoding the above image information, when the value of the EOI cancellation information is 1, the persistence of an EOI SEI message having the same EOI identifier as the EOI SEI message of the current picture among the EOI SEI messages for the previous picture of the current picture may be terminated.
[0012] In a method or device for decoding the above image information, when the value of the EOI cancellation information is 0, the encoder optimization information according to the EOI SEI message may be applied to the current picture and the subsequent picture.
[0013] In a method or device for decoding the above image information, when the value of the EOI persistence information is 0, the encoder optimization information according to the EOI SEI message may be characterized as being applied only to the current picture.
[0014] In a method or device for decoding the above image information, when the value of the EOI persistence information is 1, the encoder optimization information according to the EOI SEI message may be characterized as being applied to the current picture and subsequent picture until a predefined termination condition is satisfied.
[0015] In a method or device for decoding the above image information, the termination condition may be characterized as being at least one of the following: when a new encoded layered image sequence (CLVS) is started within the same unit, when a bitstream is terminated, or when another EOI SEI message having the same EOI identifier as the current picture is included within the same unit.
[0016] In a method or device for decoding the above image information, the value of the EOI identifier may be characterized as being in the range of 0 to 63.
[0017] In a method or device for decoding the above image information, the EOI SEI message may be characterized by further including EOI type information indicating an encoder optimization type corresponding to a predefined bitMask value.
[0018] In a method or device for decoding the above image information, it may be characterized in that whether to apply an encoder optimization type corresponding to the bitmask value to the current picture is determined based on the EOI type information and the bitmask value.
[0019] In a method or device for decoding the above image information, if the value of the EOI type information is greater than 0 and the bit value corresponding to the bitmask value in the value of the EOI type information is 0, the encoder optimization type corresponding to the bitmask value may be unspecified.
[0020] According to one aspect of the present disclosure, a method for encoding image information comprises the steps of performing encoder optimization and encoding image information including an Encoder Optimization Information (EOI) SEI (supplemental enhancement information) message generated based on information regarding the encoder optimization, wherein the EOI SEI message includes an EOI identifier for identifying the EOI SEI message, EOI cancellation information associated with the persistence of the EOI SEI message, and EOI persistence information.
[0021] According to one aspect of the present disclosure, an apparatus for encoding image information performs encoder optimization and encodes image information including an Encoder Optimization Information (EOI) SEI (supplemental enhancement information) message generated based on information regarding the encoder optimization, wherein the EOI SEI message includes an EOI identifier for identifying the EOI SEI message, EOI cancellation information associated with the persistence of the EOI SEI message, and EOI persistence information.
[0022] In a method or device for encoding the above image information, the image information may be characterized by simultaneously including two or more EOI SEI messages according to different EOI identifiers.
[0023] In a method or device for encoding the above image information, the EOI cancellation information may be characterized by being generated based on whether the persistence of an EOI SEI message having the same EOI identifier as the EOI SEI message of the current picture is terminated among the EOI SEI messages for the previous picture of the current picture.
[0024] In a method or device for encoding the above-mentioned image information, the EOI persistence information may be characterized by being generated based on whether encoder optimization information according to the EOI SEI message is applied to the current picture and subsequent picture until a predefined termination condition is satisfied.
[0025] In a method or device for encoding the above image information, the value of the EOI identifier may be characterized as being in the range of 0 or more and 63 or less.
[0026] In a method or device for encoding the above-mentioned image information, the EOI SEI message may further include EOI type information indicating an encoder optimization type corresponding to a predefined bitMask value.
[0027] In a method or device for encoding the above image information, the bit value corresponding to the bitmask value in the EOI type information or the value of the EOI type information may be characterized by being generated based on whether the encoder optimization type corresponding to the bitmask value is unspecified in the current picture.
[0028] According to one aspect of the present disclosure, a bitstream generated by a video encoding method is stored in a computer-readable storage medium. The storage medium stores a bitstream for video information including an Encoder Optimization Information SEI (supplemental enhancement information) message generated based on information regarding encoder optimization, and the EOI SEI message is characterized by including an EOI identifier for identifying the EOI SEI message, EOI cancellation information associated with the persistence of the EOI SEI message, and EOI persistence information.
[0029] According to one aspect of the present disclosure, a method for transmitting data for an image comprises the steps of performing encoder optimization and obtaining image information including an Encoder Optimization Information (EOI) SEI (supplemental enhancement information) message generated based on information regarding the encoder optimization, and transmitting the data including the image information, wherein the EOI SEI message includes an EOI identifier for identifying the EOI SEI message, EOI cancellation information associated with the persistence of the EOI SEI message, and EOI persistence information.
[0030] The features briefly summarized above regarding the present disclosure are merely exemplary aspects of the detailed description of the present disclosure that follows and do not limit the scope of the present disclosure.
[0031] According to the present disclosure, by transmitting Encoder Optimization Information (SEI) supplemental enhancement information messages having respective identifiers through EOI identifiers, optimization information having different persistences can be managed independently and in parallel.
[0032] According to the present disclosure, since a plurality of encoder optimization information SEI messages have different EOI identifiers and persistence information, the decoder can efficiently process by selectively extracting only the information of the encoder optimization SEI message corresponding to a specific EOI identifier.
[0033] According to the present disclosure, by distinguishing encoder optimization information SEI messages by EOI identifiers, the problem of having to repeatedly retransmit the same information can be resolved.
[0034] According to the present disclosure, by clearly distinguishing between the unapplied state and the unspecified state between encoder optimization types used in parallel, mutual interference between multiple encoder optimization types can be prevented, and the clarity of information transmission to the decoder and processing stability can be improved.
[0035] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs from the description below.
[0036] FIG. 1 is a schematic diagram of a VCM system to which embodiments of the present disclosure can be applied.
[0037] FIG. 2 is a schematic diagram showing a VCM pipeline structure to which embodiments of the present disclosure can be applied.
[0038] FIG. 3 is a schematic diagram of an image / video encoder to which embodiments of the present disclosure can be applied.
[0039] FIG. 4 is a schematic diagram of an image / video decoder to which embodiments of the present disclosure can be applied.
[0040] FIG. 5 is a flowchart schematically illustrating a feature / feature map encoding procedure to which embodiments of the present disclosure can be applied.
[0041] FIG. 6 is a flowchart schematically illustrating a feature / feature map decoding procedure to which embodiments of the present disclosure can be applied.
[0042] FIG. 7 is a diagram illustrating a method for decoding image information according to one embodiment of the present disclosure.
[0043] FIG. 8 is a diagram illustrating a method for encoding image information according to one embodiment of the present disclosure.
[0044] FIG. 9 is a drawing showing an example of a content streaming system to which embodiments of the present disclosure can be applied.
[0045] FIG. 10 is a drawing showing another example of a content streaming system to which embodiments of the present disclosure may be applied.
[0046] Hereinafter, embodiments of the present disclosure are described in detail with reference to the attached drawings so that those skilled in the art can easily implement them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein.
[0047] In describing the embodiments of the present disclosure, detailed descriptions of known configurations or functions are omitted if it is determined that such descriptions could obscure the essence of the present disclosure. Additionally, parts of the drawings unrelated to the description of the present disclosure have been omitted, and similar parts are denoted by similar reference numerals.
[0048] In the present disclosure, when a component is described as being "connected," "combined," or "joined" with another component, this may include not only a direct connection but also an indirect connection in which another component exists in between. Furthermore, when a component is described as "comprising" or "having" another component, this means that, unless specifically stated otherwise, it does not exclude the other component but may include an additional component.
[0049] In the present disclosure, terms such as first, second, etc. are used solely for the purpose of distinguishing one component from another and do not limit the order or importance of the components unless specifically stated otherwise. Accordingly, within the scope of the present disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and likewise, a second component in one embodiment may be referred to as a first component in another embodiment.
[0050] In this disclosure, distinct components are intended to clearly describe their respective features and do not imply that the components are separate. That is, multiple components may be integrated to form a single hardware or software unit, or a single component may be distributed to form multiple hardware or software units. Accordingly, such integrated or distributed embodiments are included within the scope of this disclosure, unless otherwise noted.
[0051] In the present disclosure, the components described in various embodiments do not necessarily mean essential components, and some may be optional components. Accordingly, embodiments consisting of a subset of the components described in one embodiment are also included within the scope of the present disclosure. Furthermore, embodiments including other components in addition to the components described in various embodiments are also included within the scope of the present disclosure.
[0052] The present disclosure relates to the encoding and decoding of images, and the terms used in the present disclosure may have the ordinary meanings commonly used in the technical field to which the present disclosure belongs, unless newly defined in the present disclosure.
[0053] The present disclosure may be applied to methods disclosed in the VVC (Versatile Video Coding) standard and / or the VCM (Video Coding for Machines) standard. Additionally, the present disclosure may be applied to methods disclosed in the EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard) or next-generation video / video coding standards (e.g., H.267 or H.268, etc.).
[0054] The present disclosure presents various embodiments regarding video / image coding, and unless otherwise noted, said embodiments may be performed in combination with one another. In the present disclosure, "video" may refer to a set of a series of images over time. "Image" may be information generated by artificial intelligence (AI). Input information used by AI in the process of performing a series of tasks, information generated during the information processing process, and output information may be used as images. "Picture" generally refers to a unit representing a single image at a specific time, and a slice / tile is a coding unit that constitutes a part of the picture in coding. A single picture may be composed of one or more slices / tiles. Additionally, a slice / tile may include one or more coding tree units (CTUs). The said CTU may be divided into one or more CUs. A tile is a rectangular area existing within a specific tile row and a specific tile column within a picture, and may be composed of multiple CTUs. A tile column can be defined as a rectangular area of CTUs, has a height equal to the height of the picture, and may have a width specified by a syntax element signaled from a bitstream portion such as a picture parameter set. A tile row can be defined as a rectangular area of CTUs, has a width equal to the width of the picture, and may have a height specified by a syntax element signaled from a bitstream portion such as a picture parameter set. A tile scan is a predetermined sequential ordering method of CTUs that divides a picture. Here, CTUs may be sequentially ordered according to a CTU raster scan within a tile, and tiles within a picture may be sequentially ordered according to the raster scan order of the picture's tiles.A slice may contain an integer number of complete tiles or an integer number of consecutive rows of complete CTUs within a single tile of a picture. A slice may be contained exclusively in a single NAL unit. A picture may consist of one or more tile groups. A tile group may contain one or more tiles. A brick may represent a rectangular area of rows of CTUs within a tile in a picture. A tile may contain one or more bricks. A brick may represent a rectangular area of rows of CTUs within a tile. A tile may be divided into multiple bricks, and each brick may contain one or more rows of CTUs belonging to the tile. A tile that is not divided into multiple bricks may also be treated as a brick.
[0055] In the present disclosure, "pixel" or "pel" may refer to the smallest unit constituting a picture (or image). Additionally, "sample" may be used as a term corresponding to pixel. A sample may generally represent a pixel or a pixel value, may represent only the pixel / pixel value of the luminance component, or may represent only the pixel / pixel value of the chroma component.
[0056] In one embodiment, particularly when applied to VCM, the pixel / pixel value may represent the independent information of each component or the pixel / pixel value of a component generated through combination, synthesis, or analysis, when there is a picture composed of a set of components having different characteristics and meanings. For example, in an RGB input, only the pixel / pixel value of R may be represented, only the pixel / pixel value of G may be represented, or only the pixel / pixel value of B may be represented. For example, only the pixel / pixel value of the Luma component synthesized using the R, G, and B components may be represented. For example, only the pixel / pixel value of an image or information extracted from the R, G, and B components through analysis may be represented.
[0057] In this disclosure, "unit" may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to that area. A unit may include one luminance block and two chroma (e.g., Cb, Cr) blocks. Depending on the case, the term "unit" may be used interchangeably with terms such as "sample array," "block," or "area." In general, an MxN block may include samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows. In one embodiment, particularly when applied to VCM, a unit may represent a basic unit containing information for performing a specific task.
[0058] In the present disclosure, "current block" may mean one of "current coding block," "current coding unit," "block to be encoded," "block to be decoded," or "block to be processed." When prediction is performed, "current block" may mean "current prediction block" or "block to be predicted." When transformation (inverse transformation) / quantization (inverse quantization) is performed, "current block" may mean "current transformation block" or "block to be transformed." When filtering is performed, "current block" may mean "block to be filtered."
[0059] Additionally, in the present disclosure, "current block" may mean "the chroma block of the current block" unless there is an explicit description of "chroma block." "The chroma block of the current block" may be expressed by including an explicit description of "chroma block," such as "chroma block" or "current chroma block."
[0060] In the present disclosure, " / " and "," may be interpreted as "and / or." For example, "A / B" and "A, B" may be interpreted as "A and / or B." Additionally, "A / B / C" and "A, B, C" may mean "at least one of A, B and / or C."
[0061] In the present disclosure, "or" may be interpreted as "and / or". For example, "A or B" may mean 1) "A" only, 2) "B" only, or 3) "A and B". Alternatively, in the present disclosure, "or" may mean "additionally or alternatively".
[0062] The present disclosure relates to VCM (Video / image coding for machines).
[0063] VCM refers to a compression technology that encodes and decodes parts of source images or video, or information acquired from source images or video, for the purpose of machine vision. In VCM, the targets for encoding and decoding can be referred to as features. Features can refer to information extracted from source images or video based on task objectives, requirements, the surrounding environment, etc. Features may have information forms different from those of the source images or video; accordingly, the compression method and representation format of the features may also differ from those of the video source.
[0064] VCMs can be applied to various fields. For example, in surveillance systems that recognize and track objects or people, VCMs can be used to store or transmit object recognition information. Furthermore, in intelligent transportation or smart traffic systems, VCMs can be used to transmit vehicle location information collected from GPS, sensing information collected from LIDAR, radar, etc., and various vehicle control information to other vehicles or infrastructure. Additionally, in the field of smart cities, VCMs can be used to perform individual tasks for interconnected sensor nodes or devices.
[0065] The present disclosure provides various embodiments relating to feature / feature map coding. Unless otherwise specifically stated, the embodiments of the present disclosure may each be implemented individually or in combination of two or more.
[0066] FIG. 1 is a schematic diagram of a VCM system to which embodiments of the present disclosure can be applied.
[0067] Referring to FIG. 1, the VCM system may include an encoding device (10) and a decoding device (20).
[0068] The encoding device (10) can generate a bitstream by compressing / encoding features / feature maps extracted from source images / videos and transmit the generated bitstream to a decoding device (20) via a storage medium or network. The encoding device (10) may also be referred to as a feature encoding device. In a VCM system, features / feature maps may be generated in each hidden layer of a neural network. The size and number of channels of the generated feature map may vary depending on the type of neural network or the location of the hidden layer. In the present disclosure, the feature map may be referred to as a feature set.
[0069] The encoding device (10) may include a feature acquisition unit (11), an encoding unit (12), and a transmission unit (13).
[0070] The feature acquisition unit (11) can acquire features / feature maps for a source image / video. According to an embodiment, the feature acquisition unit (11) can acquire features / feature maps from an external device, such as a feature extraction network. In this case, the feature acquisition unit (11) performs a feature receiving interface function. Alternatively, the feature acquisition unit (11) may acquire features / feature maps by running a neural network (e.g., CNN, DNN, etc.) with the source image / video as input. In this case, the feature acquisition unit (11) performs a feature extraction network function.
[0071] According to an embodiment, the encoding device (10) may further include a source image generation unit (not shown) for acquiring a source image / video. The source image generation unit may be implemented as an image sensor, a camera module, etc., and may acquire the source image / video through a process of capturing, synthesizing, or generating the image / video. In this case, the generated source image / video may be transmitted to a feature extraction network and used as input data for extracting features / feature maps.
[0072] The encoding unit (12) can encode the feature / feature map acquired by the feature acquisition unit (11). The encoding unit (12) can perform a series of procedures, such as prediction, transformation, and quantization, to increase encoding efficiency. The encoded data (encoded feature / feature map information) can be output in the form of a bitstream. The bitstream containing the encoded feature / feature map information may be referred to as a VCM bitstream.
[0073] The transmission unit (13) can transmit feature / feature map information or data output in the form of a bitstream to a decoding device (20) via a digital storage medium or network in the form of a file or streaming. Here, the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit (13) may include elements for creating a media file having a predetermined file format or elements for transmitting data via a broadcasting / communication network.
[0074] The decoding device (20) can obtain feature / feature map information from the encoding device (10) and restore the feature / feature map based on the obtained information.
[0075] The decoding device (20) may include a receiving unit (21) and a decoding unit (22).
[0076] The receiving unit (21) receives a bitstream from the encoding device (10), obtains feature / feature map information from the received bitstream, and transmits it to the decoding unit (22).
[0077] The decoding unit (22) can decode the feature / feature map based on the acquired feature / feature map information. To increase decoding efficiency, the decoding unit (22) can perform a series of procedures such as inverse quantization, inverse transform, and prediction corresponding to the operation of the encoding unit (14).
[0078] According to an embodiment, the decoding device (20) may further include a task analysis / rendering unit (23).
[0079] The task analysis / rendering unit (23) can perform task analysis based on the decoded feature / feature map. Additionally, the task analysis / rendering unit (23) can render the decoded feature / feature map into a form suitable for task execution. Various machine (oriented) tasks can be performed based on the task analysis results and the rendered feature / feature map.
[0080] In summary, the VCM system can encode / decode features extracted from source images / videos according to user and / or machine requests, task objectives, and surrounding environments, and perform various machine (oriented) tasks based on the decoded features. The VCM system may be implemented by extending / redesigning a video / image coding system and can perform various encoding / decoding methods defined in the VCM standard.
[0081] FIG. 2 is a schematic diagram showing a VCM pipeline structure to which embodiments of the present disclosure can be applied.
[0082] Referring to FIG. 2, the VCM pipeline (200) may include a first pipeline (210) for encoding / decoding images / videos and a second pipeline (220) for encoding / decoding features / feature maps. In the present disclosure, the first pipeline (210) may be referred to as a video codec pipeline, and the second pipeline (220) may be referred to as a feature codec pipeline.
[0083] The first pipeline (210) may include a first stage (video encoder) (211) that encodes an input video and a second stage (video decoder) (212) that decodes the encoded video to generate a restored video. The restored video may be used for human viewing, i.e., for human vision.
[0084] The second pipeline (220) may include a third stage (feature extraction network) (221) for extracting features / feature maps from an input image / video, a fourth stage (VCM encoder) (222) for encoding the extracted features / feature maps, and a fifth stage (VCM decoder) (223) for decoding the encoded features / feature maps to generate restored features / feature maps. The restored features / feature maps may be used for machine (vision) tasks. Here, a machine (vision) task may refer to a task in which images / videos are consumed by a machine. Machine (vision) tasks may be applied to service scenarios such as surveillance, intelligent transportation, smart cities, intelligent industries, intelligent content, etc. According to an embodiment, the restored features / feature maps may also be used for human vision.
[0085] According to an embodiment, the features / feature map encoded in the fourth stage (222) may be transferred to the first stage (221) and used to encode the image / video. In this case, an additional bitstream may be generated based on the encoded features / feature map, and the generated additional bitstream may be transferred to the second stage (222) and used to decode the image / video.
[0086] According to an embodiment, the feature / feature map decoded at the fifth stage (223) can be transferred to the second stage (222) and used to decode an image / video.
[0087] FIG. 2 illustrates a case where the VCM pipeline (200) includes a first pipeline (210) and a second pipeline (220), but this is merely exemplary and the embodiments of the present disclosure are not limited thereto. For example, the VCM pipeline (200) may include only the second pipeline (220), or the second pipeline (220) may be extended to a plurality of feature codec pipelines.
[0088] Meanwhile, in the first pipeline (210), the first stage (211) may be performed by an image / video encoder, and the second stage (212) may be performed by an image / video decoder. Additionally, in the second pipeline (220), the third stage (221) may be performed by a VCM encoder (or a feature / feature map encoder), and the fourth stage (222) may be performed by a VCM decoder (or a feature / feature map decoder). The encoder / decoder structure will be described in detail below.
[0089] FIG. 3 is a schematic diagram of an image / video encoder to which embodiments of the present disclosure can be applied.
[0090] Referring to FIG. 3, the image / video encoder (300) may include an image partitioner (310), a predictor (320), a residual processor (330), an entropy encoder (340), an adder (350), a filter (360), and a memory (370). The predictor (320) may include an inter-predictor (321) and an intra-predictor (322). The residual processor (330) may include a transformer (332), a quantizer (333), a dequantizer (334), and an inverse transformer (335). The residual processing unit (330) may further include a subtractor (331). The addition unit (350) may be referred to as a reconstructor or a reconstructed block generator. The above-described image segmentation unit (310), prediction unit (320), residual processing unit (330), entropy encoding unit (340), addition unit (350), and filtering unit (360) may be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to the embodiment. Additionally, the memory (370) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The above-described hardware component may further include the memory (370) as an internal / external component.
[0091] The image segmentation unit (310) can divide an input image (or picture, picture) input to the image / video encoder (300) into one or more processing units. For example, a processing unit may be referred to as a coding unit (CU). A coding unit may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, a single coding unit may be divided into multiple coding units of a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first and the binary-tree structure and / or ternary structure may be applied later. Alternatively, the binary-tree structure may be applied first. An image / video coding procedure according to the present disclosure may be performed based on a final coding unit that is no longer subdivided. In this case, the maximum coding unit may be used directly as the final coding unit based on coding efficiency according to image characteristics, or, if necessary, the coding unit may be recursively subdivided into lower-depth coding units so that a coding unit of the optimal size is used as the final coding unit. Here, the term "coding procedure" may include procedures such as prediction, transformation, and restoration described below. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may each be subdivided or partitioned from the aforementioned final coding unit.The prediction unit may be a unit of sample prediction, and the transformation unit may be a unit that derives transformation coefficients and / or a unit that derives a residual signal from transformation coefficients.
[0092] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the context. In general, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chroma component. The term "sample" may be used as a counterpart to "pixel" or "pel."
[0093] The image / video encoder (300) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (predicted block, prediction sample array) output from an inter prediction unit (321) or an intra prediction unit (322) from an input image signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (332). In this case, as illustrated, the unit that subtracts the prediction signal (predicted block, prediction sample array) from the input image signal (original block, original sample array) within the image / video encoder (300) may be referred to as a subtraction unit (331). The prediction unit can perform a prediction on a block to be processed (hereinafter referred to as the current block) and generate a predicted block containing prediction samples for the current block. The prediction unit can determine whether intra prediction is applied or inter prediction is applied at the current block or CU level. The prediction unit can generate various information regarding the prediction, such as prediction mode information, and transmit it to the entropy encoding unit (340). The information regarding the prediction can be encoded in the entropy encoding unit (340) and output in the form of a bitstream.
[0094] The intra prediction unit (322) can predict the current block by referencing samples within the current picture. At this time, the referenced samples may be located near the current block or away from it, depending on the prediction mode. In intra prediction, the prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a Planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. The intra prediction unit (322) may determine the prediction mode applied to the current block by using the prediction mode applied to the surrounding blocks.
[0095] The inter prediction unit (321) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference picture indices. Motion information may further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture containing the reference blocks and the reference picture containing the temporal neighboring blocks may be the same or different. Temporal surrounding blocks may be referred to as collocated reference blocks, collocated CUs (colCU), etc., and a reference picture containing temporal surrounding blocks may be referred to as a collocated picture (colPic). For example, the inter prediction unit (321) may construct a list of motion information candidates based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter prediction unit (321) may use the motion information of surrounding blocks as the motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0096] The prediction unit (320) can generate a prediction signal based on various prediction methods. For example, the prediction unit may apply intra prediction or inter prediction for a single block, and may also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, for example, screen content coding (SCC). IBC basically performs prediction within the current picture, but it can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may utilize at least one of the inter prediction techniques described in this disclosure. The palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, sample values within the picture can be signaled based on information regarding the palette table and palette index.
[0097] The prediction signal generated by the prediction unit (320) can be used to generate a restoration signal or to generate a residual signal. The transformation unit (332) can generate transform coefficients by applying a transformation technique to the residual signal. For example, the transformation technique may include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transformation obtained based on generating a prediction signal using all previously reconstructed pixels. Additionally, the transformation process may be applied to a pixel block of the same size in a square, or to a block of variable size that is not square.
[0098] The quantization unit (333) quantizes the transformation coefficients and transmits them to the entropy encoding unit (340), and the entropy encoding unit (340) can encode the quantized signal (information regarding the quantized transformation coefficients) and output it as a bitstream. The information regarding the quantized transformation coefficients may be called residual information. The quantization unit (333) can rearrange the block-shaped quantized transformation coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information regarding the quantized transformation coefficients based on the one-dimensional vector-shaped quantized transformation coefficients. The entropy encoding unit (340) can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit (340) may encode information necessary for image / video restoration (e.g., values of syntax elements) together or separately, in addition to the quantized transform coefficients. The encoded information (e.g., encoded image / video information) may be transmitted or stored in the form of a bitstream in units of network abstraction layer (NAL) units. The image / video information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the image / video information may further include general constraint information. Furthermore, the image / video information may further include methods for generating and using the encoded information, purposes, etc. In this disclosure, information and / or syntax elements transmitted / signaled from the image / video encoder to the image / video decoder may be included in the image / video information.Video information may be encoded through the encoding procedure described above and included in a bitstream. The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) that transmits a signal output from the entropy encoding unit (340) and / or a storage unit (not shown) that stores it may be configured as internal / external elements of the video encoder (300), or the transmission unit may be included in the entropy encoding unit (340).
[0099] The quantized transform coefficients output from the quantization unit (333) can be used to generate a prediction signal. For example, a residual signal (residual block or residual samples) can be restored by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (334) and the inverse transformation unit (335). The adder (350) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter-prediction unit (321) or the intra-prediction unit (322). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The adder (350) may be called a reconstructed unit or a reconstructed block generation unit. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, and can also be used for inter prediction of the next picture after filtering as described below.
[0100] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.
[0101] The filtering unit (360) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (360) can generate a modified restored picture by applying various filtering methods to the restored picture, and can store the modified restored picture in memory (370), specifically in the DPB of memory (370). Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (360) can generate various information regarding filtering and transmit it to the entropy encoding unit (340). The information regarding filtering can be encoded in the entropy encoding unit (340) and output in the form of a bitstream.
[0102] The modified restored picture transmitted to memory (370) can be used as a reference picture in the inter-prediction unit (321). This allows prediction mismatches in the encoder and decoder units to be avoided and encoding efficiency to be improved.
[0103] The DPB of the memory (370) can store the modified restored picture to be used as a reference picture in the inter-prediction unit (321). The memory (370) can store motion information of blocks from which motion information is derived (or encoded) within the current picture and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (321) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (370) can store restoration samples of restored blocks within the current picture and can transmit the stored restoration samples to the intra-prediction unit (322).
[0104] Meanwhile, the VCM encoder (or feature / feature map encoder) may have a structure identical or similar to the image / video encoder (300) described with reference to FIG. 3 in that it performs a series of procedures such as prediction, transformation, and quantization to encode features / feature maps. However, the VCM encoder differs from the image / video encoder (300) in that it encodes features / feature maps, and accordingly, the names of each unit (or component) (e.g., image segmentation unit (310), etc.) and the specific operation details may differ from those of the image / video encoder (300). The specific operation details of the VCM encoder will be described in detail later.
[0105] FIG. 4 is a schematic diagram of an image / video decoder to which embodiments of the present disclosure can be applied.
[0106] Referring to FIG. 4, the image / video decoder (400) may include an entropy decoder (410), a residual processor (420), a predictor (430), an adder (440), a filter (450), and a memory (460). The predictor (430) may include an inter-predictor (431) and an intra-predictor (432). The residual processor (420) may include a dequantizer (421) and an inverse transformer (422). The aforementioned entropy decoding unit (410), residual processing unit (420), prediction unit (430), addition unit (440), and filtering unit (450) may be configured by a single hardware component (e.g., a decoder chipset or a processor) according to an embodiment. Additionally, the memory (460) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (460) as an internal / external component.
[0107] When a bitstream containing video / image information is input, the video / video decoder (400) can restore the video / video in correspondence with the process in which the video / video information is processed in the video / video encoder (300) of FIG. 3. For example, the video / video decoder (400) can derive units / blocks based on block division information obtained from the bitstream. The video / video decoder (400) can perform decoding using a processing unit applied in the video / video encoder. Accordingly, the processing unit for decoding may be, for example, a coding unit, and the coding unit may be divided according to a quad tree structure, a binary tree structure, and / or a binary tree structure from a coding tree unit or a maximum coding unit. One or more conversion units may be derived from the coding unit. And, the restored video signal decoded and output through the video / video decoder (400) can be played back through a playback device.
[0108] The image / video decoder (400) can receive a signal output from the encoder (300) of FIG. 3 in the form of a bitstream, and the received signal can be decoded through an entropy decoding unit (410). For example, the entropy decoding unit (410) can parse the bitstream to derive information (e.g., image / video information) necessary for image restoration (or picture restoration). The image / video information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the image / video information may further include general constraint information. Furthermore, the image / video information may include a method for generating decoded information, a method for using it, a purpose, etc. The image / video decoder (400) may further decode the picture based on information regarding parameter sets and / or general constraint information. Information and / or syntax elements that are signaled / received can be obtained from the bitstream by decoding through a decoding procedure. For example, the entropy decoding unit (410) can decode information within the bitstream based on coding methods such as exponential chord coding, CAVLC, or CABAC, and output values of syntax elements required for image restoration and quantized values of transform coefficients regarding residuals. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, determine a context model using information on the syntax element to be decoded and decoding information of surrounding and decoding target blocks or information on symbols / bins decoded in the previous step, predict the probability of occurrence of the bin according to the determined context model, and perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin after determining the context model. Among the information decoded in the entropy decoding unit (410), information regarding prediction is provided to the prediction unit (inter prediction unit (432) and intra prediction unit (431)), and the residual value for which entropy decoding was performed in the entropy decoding unit (410), i.e., quantized transformation coefficients and related parameter information, can be input to the residual processing unit (420). The residual processing unit (420) can derive residual signals (residual blocks, residual samples, residual sample array). In addition, among the information decoded in the entropy decoding unit (410), information regarding filtering can be provided to the filtering unit (450). Meanwhile, a receiver (not shown) that receives a signal output from an image / video encoder may be further configured as an internal / external element of an image / video decoder (400), or the receiver may be a component of an entropy decoding unit (410). Meanwhile, the image / video decoder according to the present disclosure may be called an image / video decoding device, and the image / video decoder may be divided into an information decoder (image / video information decoder) and a sample decoder (image / video sample decoder). In this case, the information decoder may include an entropy decoding unit (410), and the sample decoder may include at least one of an inverse quantization unit (321), an inverse transform unit (322), an adder (440), a filtering unit (450), a memory (460), an inter prediction unit (432), and an intra prediction unit (431).
[0109] In the inverse quantization unit (421), the quantized transform coefficients can be inversely quantized to output transform coefficients. The inverse quantization unit (421) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed in the image / video encoder. The inverse quantization unit (321) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.
[0110] In the inverse conversion unit (422), the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).
[0111] The prediction unit (430) can perform a prediction for the current block and generate a predicted block containing prediction samples for the current block. Based on the information regarding the prediction output from the entropy decoding unit (410), the prediction unit can determine whether an intra prediction or an inter prediction is applied to the current block and can determine a specific intra / inter prediction mode.
[0112] The prediction unit (420) can generate a prediction signal based on various prediction methods. For example, the prediction unit may apply intra prediction or inter prediction for a single block, and may also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, for example, screen content coding (SCC). IBC basically performs prediction within the current picture, but it can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. The palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and palette index may be included in the image / video information and signaled.
[0113] The intra prediction unit (431) can predict the current block by referring to samples within the current picture. The referenced samples may be located next to the current block or away from it, depending on the prediction mode. In intra prediction, the prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit (431) may determine the prediction mode applied to the current block by using the prediction mode applied to the surrounding blocks.
[0114] The inter prediction unit (432) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. Motion information may include a motion vector and a reference picture index. Motion information may further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (432) can construct a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-prediction can be performed based on various prediction modes, and information regarding the prediction may include information indicating the mode of inter-prediction for the current block.
[0115] The adder (440) can generate a restoration signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the prediction signal (predicted block, predicted sample array) output from the prediction unit (including the inter prediction unit (432) and / or the intra prediction unit (431)). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the restoration block.
[0116] The addition unit (440) may be called a restoration unit or a restoration block generation unit. The generated restoration signal may be used for intra-predicting the next block to be processed within the current picture, may be output after filtering as described below, or may be used for inter-predicting the next picture.
[0117] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0118] The filtering unit (450) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (450) can generate a modified restored picture by applying various filtering methods to the restored picture, and can transmit the modified restored picture to memory (460), specifically to the DPB of memory (460). Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0119] The (modified) restored picture stored in the DPB of the memory (460) can be used as a reference picture in the inter-prediction unit (432). The memory (460) can store motion information of blocks from which motion information within the current picture has been derived (or decoded) and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (432) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (460) can store restoration samples of blocks restored within the current picture and transmit them to the intra-prediction unit (431).
[0120] Meanwhile, the VCM decoder (or feature / feature map decoder) may have a structure identical or similar to the image / video decoder (400) described above with reference to FIG. 4 in that it performs a series of procedures such as prediction, inverse transformation, and inverse quantization to decode features / feature maps. However, the VCM decoder differs from the image / video decoder (400) in that it targets features / feature maps for decoding, and accordingly, the names of each unit (or component) (e.g., DPB, etc.) and the specific operation details may differ from those of the image / video decoder (400). The operation of the VCM decoder may correspond to the operation of the VCM encoder, and the specific operation details will be described in detail later.
[0121] FIG. 5 is a flowchart schematically illustrating a feature / feature map encoding procedure to which embodiments of the present disclosure can be applied.
[0122] Referring to FIG. 5, the feature / feature map encoding procedure may include a prediction procedure (S510), a residual processing procedure (S520), and an information encoding procedure (S530).
[0123] The prediction procedure (S510) can be performed by the prediction unit (320) described above with reference to FIG. 3.
[0124] Specifically, the intra prediction unit (322) can predict the current block (i.e., the set of feature elements currently being encoded) by referencing feature elements within the current feature / feature map. Intra prediction can be performed based on the spatial similarity of the feature elements constituting the feature / feature map. For example, feature elements included in the same Region of Interest (RoI) within an image / video can be presumed to have similar data distribution characteristics. Accordingly, the intra prediction unit (322) can predict the current block by referencing the previously restored feature elements within the region of interest containing the current block. At this time, the referenced feature elements may be located adjacent to the current block or spaced apart from the current block depending on the prediction mode. Intra prediction modes for feature / feature map encoding may include a plurality of non-directional prediction modes and a plurality of directional prediction modes. Non-directional prediction modes may include prediction modes corresponding, for example, to the DC mode and planner mode of a video encoding procedure. Additionally, directional modes may include prediction modes corresponding, for example, to 33 directional modes or 65 directional modes of a video encoding procedure. However, this is merely an example, and the types and number of intra prediction modes may be set or changed in various ways depending on the embodiment.
[0125] The inter prediction unit (321) can predict the current block based on a reference block (i.e., a set of referenced feature elements) specified by motion information on the reference feature / feature map. Inter prediction can be performed based on the temporal similarity of the feature elements constituting the feature / feature map. For example, temporally consecutive features may have similar data distribution characteristics. Therefore, the inter prediction unit (321) can predict the current block by referring to the restored feature elements of features temporally adjacent to the current feature. At this time, motion information for specifying the referenced feature elements may include motion vectors and reference feature / feature map indices. The motion information may further include information regarding the direction of inter prediction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current feature / feature map and temporal neighboring blocks existing within the reference feature / feature map. A reference feature / feature map containing a reference block and a reference feature / feature map containing temporal surrounding blocks may be the same or different. Temporal surrounding blocks may be referred to as collocated reference blocks, and a reference feature / feature map containing temporal surrounding blocks may be referred to as a collocated feature / feature map. The inter-prediction unit (321) may construct a list of motion information candidates based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference feature / feature map index of the current block.Inter prediction can be performed based on various prediction modes, for example, in the case of skip mode and merge mode, the inter prediction unit (321) can use the motion information of surrounding blocks as the motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted. In the case of motion vector prediction (MVP) mode, the motion vector of surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference. In addition to the intra prediction and inter prediction described above, the prediction unit (320) can generate a prediction signal based on various prediction methods.
[0126] The prediction signal generated by the prediction unit (320) can be used to generate a residual signal (residual block, residual feature elements) (S520). The residual processing procedure (S520) can be performed by the residual processing unit (330) described above with reference to FIG. 3. Then, (quantized) transformation coefficients can be generated through a transformation and / or quantization procedure for the residual signal, and the entropy encoding unit (340) can encode information regarding the (quantized) transformation coefficients as residual information within the bitstream (S530). In addition, the entropy encoding unit (340) can encode information necessary for feature / feature map restoration, such as prediction information (e.g., prediction mode information, motion information, etc.), in addition to the residual information within the bitstream.
[0127] Meanwhile, the feature / feature map encoding procedure may further include a procedure (S530) for encoding information for feature / feature map restoration (e.g., prediction information, resident information, partitioning information, etc.) and outputting it in the form of a bitstream, as well as a procedure for generating a restored feature / feature map for the current feature / feature map and a procedure (optional) for applying in-loop filtering to the restored feature / feature map.
[0128] The VCM encoder can derive (modified) residual feature(s) from quantized transform coefficient(s) through inverse quantization and inverse transform, and can generate a reconstructed feature / feature map based on the predicted feature(s) and (modified) residual feature(s) which are the outputs of step S510. The reconstructed feature / feature map thus generated may be identical to the reconstructed feature / feature map generated by the VCM decoder. If an in-loop filtering procedure is performed on the reconstructed feature / feature map, a modified reconstructed feature / feature map may be generated through the in-loop filtering procedure on the reconstructed feature / feature map. The modified reconstructed feature / feature map is stored in a decoded feature buffer (DFB) or memory and can be used as a reference feature / feature map in a subsequent prediction procedure of the feature / feature map. Additionally, information (parameters) related to (in-loop) filtering may be encoded and output in the form of a bitstream. Through the in-loop filtering procedure, noise that may occur during feature / feature map coding can be removed, and the performance of feature / feature map-based tasks can be improved. Furthermore, by performing the in-loop filtering procedure at both the encoder and decoder stages, the consistency of prediction results can be guaranteed, the reliability of feature / feature map coding can be enhanced, and the amount of data transmitted for feature / feature map coding can be reduced.
[0129] FIG. 6 is a flowchart schematically illustrating a feature / feature map decoding procedure to which embodiments of the present disclosure can be applied.
[0130] Referring to FIG. 6, the feature / feature map decoding procedure may include an image / video information acquisition procedure (S610), a feature / feature map restoration procedure (S620–S640), and an in-loop filtering procedure for the restored feature / feature map (S650). The feature / feature map restoration procedure may be performed based on prediction signals and residual signals obtained through the inter / intra prediction (S620) and residual processing (S630) and the inverse quantization and inverse transformation process for the quantized transformation coefficients described in the present disclosure. A modified restored feature / feature map may be generated through the in-loop filtering procedure for the restored feature / feature map, and the modified restored feature / feature map may be output as a decoded feature / feature map. The decoded feature / feature map can be stored in a decoding feature buffer (DFB) or memory and used as a reference feature / feature map in the inter-prediction procedure during subsequent decoding of the feature / feature map. In some cases, the aforementioned in-loop filtering procedure may be omitted. In this case, the restored feature / feature map can be output as is as the decoded feature / feature map, or stored in a decoding feature buffer (DFB) or memory and used as a reference feature / feature map in the inter-prediction procedure during subsequent decoding of the feature / feature map.
[0131] The SEI message related to the present disclosure will be described below.
[0132] Table 1 shows an example of an encoder optimization information (EOI) SEI message syntax according to one embodiment.
[0133] [Table 1]
[0134]
[0135] An example of encoder optimization information SEI message semantics according to one embodiment is described.
[0136] Encoder Optimization Information SEI messages are used to indicate whether the video is optimized for human viewing or machine analysis, and what type of optimization was applied during the preprocessing or encoding process.
[0137] If eoi_cancel_flag is 1, it indicates that the persistence of the encoder optimization information SEI message included in the previous PU in the output order is canceled. If eoi_cancel_flag is 0, it indicates that optimization information applied during the preprocessing or encoding stage follows.
[0138] eoi_persistence_flag indicates the persistence of the optimization information provided in this SEI message. If eoi_persistence_flag is 0, it indicates that the optimization information is applied only to the current picture. If eoi_persistence_flag is 1, it indicates that the optimization information is applied to the current picture and all pictures following the current layer in output order, and persists until one or more of the following conditions are met.
[0139] - Until new CLVS of the current layer start.
[0140] - Until the bitstream ends.
[0141] - Until the picture of the current layer associated with the encoder optimization information SEI message is output following the current picture in the output order.
[0142] If eoi_for_human_viewing_idc is 3, it indicates that human viewing is included in the applied optimization objectives. If eoi_for_human_viewing_idc is 2, it indicates that the video is suitable for human viewing but has not been specifically optimized. If eoi_for_human_viewing_idc is 1, it indicates that the video is unsuitable for human viewing. If eoi_for_human_viewing_idc is 0, it indicates that it is unclear whether the video is suitable for human viewing.
[0143] If eoi_for_machine_analysis_idc is 3, it indicates that machine analysis is included in the applied optimization objectives. If eoi_for_machine_analysis_idc is 2, it indicates that the image is suitable for machine analysis but has not been specifically optimized. If eoi_for_machine_analysis_idc is 1, it indicates that the image is not suitable for machine analysis. If eoi_for_machine_analysis_idc is 0, it indicates that it is unknown whether the image is suitable for machine analysis.
[0144] As a bitstream conformance requirement, the values of eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc must not be 1 at the same time.
[0145] eoi_type indicates the type of optimization method specified in Table 2. If (eoi_type & bitMask) is not equal to 0, it means that the optimization type corresponding to the bitMask value listed in Table 2 has been applied. If eoi_type is greater than 0 and (eoi_type & bitMask) is equal to 0, it means that the optimization type corresponding to the bitMask value has not been applied. If eoi_type is equal to 0, it indicates that the optimization determined by the application has been used.
[0146] Table 2 below defines which optimization type each bit (bitMask) configuration of the eoi_type value represents.
[0147] [Table 2]
[0148]
[0149] The variables EoiObjectBasedFlag, EoiTemporalResamplingFlag, EoiSpatialResamplingFlag, EoiTemporalQualityFlag, EoiSpatialQualityFlag, and EoiPrivacyProtectionFlag each specify whether eoi_type includes object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, and privacy protection optimization, and are derived according to Table 3 below.
[0150] [Table 3]
[0151]
[0152] For reference, for example, if a specific top temporal sublayer is encoded with coarse quantization that makes quality variations unpleasant for human viewers but does not degrade machine performance, eoi_for_human_viewing_flag and eoi_for_machine_analaysis_flag can be set to 0 and 1, respectively, and eoi_type can be set to a value that makes EoiTemporalQualityFlag 1.
[0153] When eoi_persistence_flag is 0, EoiTemporalResamplingFlag must be 0 and EoiTemporalQualityFlag must be 0 as bitstream conformity requirements.
[0154] If eoi_object_based_idc exists, it indicates an object-based optimization type as specified in Table 4. If (eoi_object_based_idc & bitMask) is not 0, it means that an object-based optimization type associated with the bitMask value in Table 3 has been applied. If eoi_object_based_idc is greater than 0 and (eoi_object_based_idc & bitMask) is 0, it means that an object-based optimization type associated with the bitMask value has not been applied. If eoi_object_based_idc is 0, it means that an object-based optimization type of the application-defined type has been applied. The eoi_object_based_idc value must be in the range of 0 to 7 (inclusive) in a bitstream conforming to the present disclosure. Values of eoi_object_based_idc from 8 to 65,535 (inclusive) are reserved for future use and must not exist in a bitstream conforming to the present disclosure. If the value of eoi_object_based_idc is from 8 to 65,535 (including that value), decoders conforming to this version of the specification must ignore eoi_object_based_idc.
[0155] Table 4 below defines which object-based optimization type each bit (bitMask) of the eoi_object_based_idc value represents.
[0156] [Table 4]
[0157]
[0158] If eoi_temporal_resampling_type_flag is 0, it indicates that temporal resampling optimization is a subsampling operation. If eoi_temporal_resampling_type_flag is 1, it indicates that temporal resampling optimization is an upsampling operation.
[0159] If eoi_num_int_pics is greater than 0, it indicates that within the persistence of this SEI message, the number of pictures excluded by the encoding system between each pair of encoded pictures in the output order (when eoi_temporal_resampling_type_flag is 0), or the number of pictures added between source pairs of pictures for encoding (when eoi_temporal_resampling_type_flag is 1), is constant. If eoi_temporal_resampling_type_flag is 0 and eoi_num_int_pics is greater than 0, eoi_num_int_pics specifies the number of pictures excluded by the encoding system between each pair of encoded pictures in the output order. If eoi_temporal_resampling_type_flag is 1 and eoi_num_int_pics is greater than 0, eoi_num_int_pics indicates the number of pictures added by the encoding system between source pairs of pictures for encoding.
[0160] If eoi_num_int_pics is 0, it indicates that within the persistence of the SEI message, the number of pictures excluded by the encoding system between each pair of encoded pictures in the output order (when eoi_temporal_resampling_type_flag is 0), or the number of pictures added between pairs of source pictures for encoding (when eoi_temporal_resampling_type_flag is 1) is unknown or varies.
[0161] The value of eoi_num_int_pics must be in the range of 0 to 63 (inclusive).
[0162] If eoi_orig_pic_dimensions_flag is 1, it indicates that the eoi_orig_pic_width and eoi_orig_pic_height syntax elements exist. If eoi_orig_pic_dimensions_flag is 0, it specifies that the eoi_orig_pic_width and eoi_orig_pic_height syntax elements do not exist.
[0163] If eoi_orig_pic_width and eoi_orig_pic_height exist, they represent the width and height of the original source picture in luma samples, respectively.
[0164] If eoi_spatial_resampling_type_flag is 0, it indicates that spatial resampling optimization is a subsampling operation. If eoi_spatial_resampling_type_flag is 1, it indicates that spatial resampling optimization is an upsampling operation.
[0165] If eoi_privacy_protection_method_idc exists, it indicates the method or algorithm used for privacy protection optimization as specified in Table 5.
[0166] Table 5 below defines which privacy protection optimization method or algorithm each bit (bitMask) of the eoi_privacy_protection_type_idc value represents.
[0167] [Table 5]
[0168]
[0169] If eoi_privacy_info_type exists, it indicates the type of protected information as specified in Table 6. If eoi_privacy_info_type is greater than 0 and (eoi_privacy_info_type & bitMask) is not equal to 0, it indicates that the type of information corresponding to the bitMask value listed in Table 6 is protected. If eoi_privacy_info_type is equal to 0, it indicates that the type of information defined by the application is protected. In a bitstream conforming to the present disclosure, the value of eoi_privacy_info_type must be in the range of 0 to 7. If the value of eoi_privacy_info_type is in the range of 8 to 255, such value is reserved for future use by ITU-T | ISO / IEC and must not exist in a bitstream conforming to the present disclosure. If the value of eoi_privacy_info_type is in the range of 8 to 255, a decoder conforming to the present disclosure must ignore eoi_privacy_info_type.
[0170] [Table 6]
[0171]
[0172] The current Encoder Optimization Information (EOI) message design includes signal information regarding various types of encoder optimization methods applied before or during the encoding process. Additionally, only one EOI SEI message can be active at any given time, which implies that all optimizations must have the same persistence. However, this persistence may vary depending on the different types of optimization. For example, it may be desirable to apply spatial upsampling to all source pictures and image quality enhancement to only some of the pictures.
[0173] The present disclosure proposes enabling multiple EOI SEI messages to represent optimization types having different persistence ranges within a single Access Unit or Current Layer.
[0174] One embodiment provides a solution to the problem described above. Each item may be applied individually or in combination.
[0175] 1. Allows multiple EOI SEI messages to be enabled simultaneously.
[0176] 2. Add a syntax element representing an identifier to identify the Encoder Optimization Information SEI message as follows.
[0177] - Identify EOI SEI messages using eoi_id.
[0178] - When the value of eoi_cancel_flag is equal to 1, previous EOI SEI messages with the same eoi_id in this EOI SEI message are canceled.
[0179] 3. Modify the semantics of eoi_type as follows.
[0180] - If eoi_type is greater than 0 and ( eoi_type & bitMask ) is equal to 0, the optimization type corresponding to the bitMask value is considered unspecified.
[0181] One embodiment relates to items 1 and 2 described above. One embodiment is based on VSEI and VVC specifications.
[0182] According to one embodiment, the Encoder Optimization Information SEI message syntax is as shown in Table 7 below.
[0183] [Table 7]
[0184]
[0185] Encoder Optimization Information SEI messages are used to indicate whether the video is optimized for human viewing or machine analysis, and what type of optimization was applied during the preprocessing or encoding process.
[0186] eoi_id specifies the identifier of the EOI SEI message. The value of eoi_id must be within the range of 0 to 63.
[0187] If the value of eoi_cancel_flag is 1, it specifies that the persistence of all previous encoder optimization information SEI messages with the same eoi_id in the output order is canceled. If the value of eoi_cancel_flag is 0, it indicates that information regarding optimizations applied during preprocessing or encoding is provided subsequently.
[0188] eoi_persistence_flag specifies the persistence of optimization information provided in SEI messages. If the value of eoi_persistence_flag is 0, the optimization information is applied only to the current picture. If the value of eoi_persistence_flag is 1, the optimization information is applied to all subsequent pictures in the same hierarchy as the current picture and persists until one or more of the following conditions become true.
[0189] - When a new CLVS (Coded Layer-wise Video Sequence) of the same layer starts
[0190] - When the bitstream ends
[0191] - When a picture containing an encoder optimization information SEI message with the same eoi_id value is output according to the output order following the current picture in the same layer
[0192] One embodiment relates to Item 3 described above. One embodiment is based on VSEI and VVC specifications.
[0193] eoi_type indicates the type of optimization method as specified in Table 2 above. If the value of (eoi_type & bitMask) is not 0, it means that the optimization type corresponding to the bitMask value in the aforementioned Table 2 has been applied. If the value of eoi_type is greater than 0 and the value of (eoi_type & bitMask) is 0, it means that the optimization type corresponding to the bitMask value is unspecified. If the value of eoi_type is 0, it means that an optimization determined by the application has been used.
[0194] The terms or names described below (e.g., names of syntax elements or variables, etc.) are merely examples, and the technical features of the present disclosure are not limited to the terms, etc. described below. For example, the image information described below may include various information according to the embodiments described in the present disclosure and may include information described in at least one of the tables described above.
[0195] The operations described below do not constitute an essential component of one embodiment, and at least some of the operations described below may be omitted. Furthermore, the operations described below do not constitute a sufficient component of one embodiment, and previously described operations may be added. Moreover, unless they contradict previously described operations, the operations described below form one embodiment integrally with previously described operations and do not form a separate embodiment distinct from previously described operations.
[0196] FIG. 7 is a diagram illustrating a method for decoding image information according to one embodiment of the present disclosure.
[0197] Terms or names (e.g., names of syntax elements or names of variables, etc.) are merely examples, and the technical features of the present disclosure are not limited to the terms, etc. described in FIG. 7. For example, the image information described in FIG. 7 may include various information according to the embodiments described in the present disclosure and may include information described in at least one of the tables described above.
[0198] The decoding method (S700) may include operations described below. The operations described below do not constitute an essential component of the decoding method according to one embodiment, and at least some of the operations described below may be omitted. Furthermore, the operations described below do not constitute a sufficient component of the decoding method according to one embodiment, and the previously described operations may be added. Moreover, unless the operations described below contradict the previously described operations, they form an embodiment integrally with the previously described operations and do not form a separate embodiment distinct from the previously described operations.
[0199] The decoding method (S700) can be executed by a decoding device including a memory and a processor electrically connected to the memory, for example, by a processor.
[0200] The decoding device can acquire SEI (supplemental enhancement information) messages (S710).
[0201] For example, a processor of a decoding device may acquire an SEI message. The SEI message may convey a specific type of information that assists in processes related to the decoding, display, or other purposes of image information. Here, the SEI message may not be necessary for the decoding process to determine the sample values of the decoded picture.
[0202] For example, the SEI message may include an Encoder Optimization Information (EOI) SEI (Supplemental Enhancement Information) message. According to one embodiment, a decoding device may obtain an Encoder Optimization Information SEI message associated with the optimization of a picture from a bitstream.
[0203] Encoder optimization information SEI messages may include information regarding SEI message handling and information regarding encoder optimization. For example, an encoder optimization information SEI message may provide encoder optimization information indicating whether the video is optimized for human viewing or machine analysis, and what type of optimization was applied during the preprocessing or encoding process. For example, encoder optimization may include optimization for human viewing, optimization for machine analysis, object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, or privacy optimization.
[0204] Encoder optimization information SEI messages can have various names. For example, an encoder optimization information SEI message may be referred to as an EOI SEI message, an EOI SEI related message, EOI SEI information, EOI SEI related information, encoder optimization information message, EOI message, EOI related message, EOI information, EOI related information, etc.
[0205] The encoder optimization information SEI message can take various forms. For example, the encoder optimization information SEI message may be a syntax element or a syntax structure containing one or more syntax elements. Additionally, the encoder optimization information SEI message may be a raw byte sequence payload (RBSP) containing one or more syntax elements or one or more syntax structures. For example, the encoder optimization information SEI message may be represented as encoder_optimization_info(payloadSize), but is not limited thereto.
[0206] According to one embodiment, an encoder optimization information SEI message may include an EOI identifier for identifying the encoder optimization information SEI message, EOI cancellation information and EOI persistence information associated with the persistence of the encoder optimization information SEI message, and EOI type information. Additionally, the encoder optimization information SEI message may further include human viewing optimization information, machine analysis optimization information, object-based optimization type information, temporal resampling type information, information regarding spatial resampling type, privacy protection method type information and / or privacy type information.
[0207] An EOI identifier may represent an identifier for distinguishing encoder optimization information SEI messages. An EOI identifier may be used to identify multiple encoder optimization information SEI messages transmitted within the same bitstream, and enables distinguishing specific coding conditions, algorithms, or purposes (e.g., image quality improvement, bit rate control, noise suppression, etc.) corresponding to the encoder optimization information contained in each encoder optimization information SEI message.
[0208] An EOI identifier can take various forms and be represented by various names. For example, an EOI identifier may be a syntax element or a syntax structure containing one or more syntax elements. For example, an EOI identifier that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits or a variable-length bit. An identifier that is a syntax element may be represented as eoi_id, etc., but is not limited thereto. For example, the value of an EOI identifier may be in the range of 0 to 63.
[0209] According to one embodiment, two or more encoder optimization information SEI messages having different EOI identifiers may exist simultaneously. That is, multiple encoder optimization information SEI messages having different EOI identifiers can be included within a single Access Unit or Current Layer to independently define the scope of application and persistence of each optimization information. Through this, it is possible to apply selective optimization, such as image quality enhancement, to only some pictures while applying spatial upsampling to all source pictures.
[0210] Accordingly, the decoding device can refer to the EOI identifier received from the encoding device to determine the association with a previous encoder optimization information SEI message having the same EOI identifier, or terminate the persistence of the optimization information corresponding to the identifier if EOI cancellation information is set. Therefore, the EOI identifier can clearly define the relationship between multiple encoder optimization information SEI messages and function as a reference standard for efficiently managing the scope of application and whether to update SEI messages having the same optimization purpose.
[0211] EOI cancellation information can indicate whether the encoder optimization information SEI message persists. EOI cancellation information can serve to control whether to explicitly terminate or maintain the validity of the encoder optimization information applied to the previous picture during the decoding process.
[0212] EOI cancellation information may take various forms and may be represented by various names. For example, EOI cancellation information may be a syntax element or a syntax structure containing one or more syntax elements. For example, EOI cancellation information that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits or a variable-length bit. EOI cancellation information that is a syntax element may be represented as eoi_cancel_flag, etc., but is not limited thereto.
[0213] For example, when EOI cancellation information is enabled, it may indicate that the persistence of an encoder optimization information SEI message having the same EOI identifier as the encoder optimization information SEI message of the current picture is terminated among encoder optimization information SEI messages for the previous picture of the current picture included in the same unit as the encoder optimization information SEI message. For instance, if the value of EOI cancellation information is 1, it may indicate that the persistence of an encoder optimization information SEI message having the same EOI identifier as the encoder optimization information SEI message that existed prior to the current picture within the same unit as the encoder optimization information SEI message is terminated.
[0214] For example, when EOI cancellation information is disabled, encoder optimization information according to the encoder optimization information SEI message may be applied to the current picture and subsequent picture included in the same unit as the encoder optimization information SEI message. For instance, if the value of the EOI cancellation information is 0, the encoder optimization information according to the encoder optimization information SEI message may indicate that it is applied to the current picture and subsequent picture included in the same unit as the encoder optimization information SEI message. However, this is not limited thereto, and alternatively, specifying that the value of the EOI cancellation information is 1 may be changed from specifying that the value of the EOI cancellation information is 0.
[0215] EOI persistence information may be information indicating persistence included in the encoder optimization information SEI message. For example, EOI persistence information may be used as a control element to specify the time range or validity period for which encoder optimization information is applied. EOI persistence information may be configured to distinguish whether the encoder optimization information included in the encoder optimization information SEI message is applied only to the current picture, or to one or more subsequent pictures after the current picture.
[0216] EOI persistence information may take various forms and be represented by various names. For example, EOI persistence information may be a syntax element or a syntax structure containing one or more syntax elements. For example, EOI persistence information that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits or a variable length bit. EOI persistence information that is a syntax element may be represented as eoi_persistence_flag, etc., but is not limited thereto.
[0217] For example, if EOI persistence information is disabled, encoder optimization information is applied only to the current picture and may not be applied to subsequent pictures included in the same unit as the encoder optimization information SEI message. For instance, if the value of EOI persistence information is 0, the encoder optimization information according to the encoder optimization information SEI message may be applied only to the current picture.
[0218] For example, when EOI persistence information is enabled, encoder optimization information may be continuously applied not only to the current picture but also to subsequent pictures. For instance, when the value of EOI persistence information is 1, encoder optimization information according to the encoder optimization information SEI message may be applied to the current picture and subsequent pictures included in the same unit as the encoder optimization information SEI message until a predefined termination condition is satisfied. However, this is not limited thereto, and alternatively, specifying that the value of EOI persistence information is 1 may be changed to specifying that the value of EOI persistence information is 0.
[0219] The condition for the termination of persistence of encoder optimization information may represent a condition for defining the point in time or situation in which the application scope of optimization information included in the encoder optimization information SEI message is terminated. Here, the condition for the termination of persistence of encoder optimization information may be one of the cases where a new coded layer-wise video sequence starts or the bitstream ends within the same unit containing the encoder optimization information SEI message. Alternatively, the condition for the termination of persistence of encoder optimization information may be when another encoder optimization information SEI message having the same EOI identifier as the current picture is included within the same unit. That is, it may be when a picture containing an encoder optimization information SEI message having the same EOI identifier is output according to the output order following the current picture in the same layer.
[0220] According to one embodiment, a plurality of encoder optimization information SEI messages may each have different EOI identifiers and persistence information. Accordingly, they can be configured to enable mutually independent and parallel control for each encoder optimization type.
[0221] Human viewing optimization information can indicate whether the optimization purpose involves human viewing. In other words, human viewing optimization information can indicate whether the optimized image is optimized for human viewing, suitable, unsuitable, or unknown.
[0222] Human viewing optimization information may take various forms and may be expressed by various names. For example, human viewing optimization information may be a syntax element or a syntax structure containing one or more syntax elements. For example, human viewing optimization information that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits or variable-length bits. Human viewing optimization information that is a syntax element may be expressed as eoi_for_human_viewing_idc or eoi_for_human_viewing_flag, but is not limited thereto.
[0223] Machine analysis optimization information can indicate whether the optimization objective involves machine analysis. In other words, machine analysis optimization information can indicate whether the optimized image is optimized for machine analysis, is suitable, unsuitable, or unknown.
[0224] Machine analysis optimization information may take various forms and may be represented by various names. For example, machine analysis optimization information may be a syntax element or a syntax structure containing one or more syntax elements. For example, machine analysis optimization information that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits or variable-length bits. Machine analysis optimization information that is a syntax element may be represented as eoi_for_machine_analysis_idc or eoi_for_machine_analysis_flag, but is not limited thereto.
[0225] EOI type information indicates the type of optimization applied. Encoder optimization may include object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, and / or privacy optimization. One or more of object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, and privacy optimization may be applied, and optimization type information may indicate the type of all applied optimizations.
[0226] EOI type information may take various forms and may be represented by various names. For example, optimization type information may be a syntax element or a syntax structure containing one or more syntax elements. For example, optimization type information that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits (e.g., 16 bits) or variable-length bits. Optimization type information that is a syntax element may be represented as eoi_type, eoi_type_flag, or eoi_type_idc, but is not limited thereto.
[0227] Based on EOI type information, multiple optimization type variables can be derived to indicate whether object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, and privacy protection optimization are applied, respectively. For example, the multiple optimization type variables may include, but are not limited to, EoiObjectBasedFlag, EoiTemporalResamplingFlag, EoiSpatialResamplingFlag, EoiTemporalQualityFlag, EoiSpatialQualityFlag, and EoiPrivacyProtectionFlag.
[0228] For example, each optimization type variable can represent an encoder optimization type corresponding to a bitmask value predefined according to Table 2 above. That is, EOI type information can represent an encoder optimization type corresponding to a predefined bitmask value.
[0229] For example, whether an encoder optimization type is applied to the current picture can be determined based on EOI type information and bitmask values. For instance, if the bit value corresponding to the bitmask value in the EOI type information is 1, the encoder optimization type corresponding to the bitmask value may be applied to the current picture. Here, the bit value corresponding to the bitmask value in the EOI type information may be represented by syntax elements (eoi_type & bitMask), but is not limited thereto. For instance, if the value of the EOI type information is 0, encoder optimization information pre-set by the application may be applied to the current picture.
[0230] According to one embodiment, if the value of the EOI type information is greater than 0 and the bit value corresponding to the bitmask value in the EOI type information is 0, the encoder optimization type corresponding to the bitmask value may be set to an unspecified state. This indicates that a specific optimization type is not explicitly applied and can be used as a control means to prevent semantic conflicts or duplicate applications with other parallel optimization information.
[0231] For example, two or more encoder optimization type SEI messages having different EOI type information may exist simultaneously. In this case, each encoder optimization type SEI message may correspond to a different encoder optimization type or purpose (e.g., image quality enhancement, spatial upsampling, noise suppression, etc.). Accordingly, even in situations where multiple optimization types are applied in parallel, the encoder optimization information SEI messages can be configured to prevent mutual interference or semantic ambiguity between encoder optimizations by distinguishing the EOI type information into an unapplied state and an unspecified state.
[0232] Object-based optimization type information is obtained when the optimization type indicated by the optimization type information includes object-based optimization, and may indicate the type of object-based optimization including blurring, quantization, and / or overwriting.
[0233] Object-based optimization type information may take various forms and may be represented by various names. For example, object-based optimization type information may be a syntax element or a syntax structure containing one or more syntax elements. For example, object-based optimization type information that is a syntax element may include multiple flags consisting of one bit or an indicator consisting of variable-length bits. Optimization type information that is a syntax element may be represented as eoi_object_based_flag or eoi_object_based_idc, but is not limited thereto.
[0234] Temporal resampling type information is obtained when the optimization type indicated by the optimization type information includes temporal resampling, and may indicate the type of temporal resampling. Temporal resampling may include temporal upsampling and temporal subsimplification. Temporal resampling type information may indicate whether the applied temporal resampling is temporal upsampling or temporal subsimplification. For example, temporal resampling type information with a value of 0 indicates that the applied temporal resampling is temporal subsampling. Temporal resampling type information with a value of 1 indicates that the applied temporal resampling is temporal upsampling. However, this is not limited thereto, and what is indicated by temporal resampling type information with a value of 1 may be interchangeable with what is indicated by temporal resampling type information with a value of 0.
[0235] Temporal resampling type information may take various forms and may be represented by various names. For example, temporal resampling type information may be a syntax element or a syntax structure containing one or more syntax elements. For example, temporal resampling type information that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits or variable-length bits. Temporal resampling type information that is a syntax element may be represented as eoi_temporal_resampling_type_flag or eoi_temporal_resampling_type_idc, but is not limited thereto.
[0236] Information regarding the spatial resampling type is obtained when the optimization type indicated by the optimization type information includes spatial resampling, and can indicate the type of spatial resampling. Spatial resampling includes resampling in the horizontal direction (width direction) and resampling in the vertical direction (height direction) in terms of space, and may include spatial upsampling and spatial subsampling in terms of sampling. In other words, spatial resampling may include upsampling / subsampling in the horizontal direction (width direction) and upsampling / subsampling in the vertical direction (height direction).
[0237] Information regarding spatial resampling types can be provided in various ways. For example, information regarding horizontal (width direction) upsampling / subsampling and vertical (height direction) upsampling / subsampling can be provided by providing information regarding the width and height of the original source picture. Additionally, resampling information indicating horizontal (width direction) upsampling / subsampling and resampling information indicating vertical (height direction) upsampling / subsampling can be provided directly.
[0238] For example, information regarding the spatial resampling type may include information providing the spatial resampling type (e.g., eoi_orig_pic_dimensions_flag), information regarding the width and height of the original source picture (e.g., eoi_orig_pic_width and eoi_orig_pic_height), and / or information regarding the spatial resampling type (e.g., eoi_spatial_resampling_type_flag).
[0239] Information regarding spatial resampling type may take various forms and may be expressed by various names. For example, information regarding spatial resampling type may be a syntax element or a syntax structure containing one or more syntax elements. For example, information regarding spatial resampling type that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits or variable-length bits. Information regarding spatial resampling type that is a syntax element may include syntax elements expressed as eoi_orig_pic_dimensions_flag, eoi_orig_pic_width, eoi_orig_pic_height, and eoi_spatial_resampling_type_flag, but is not limited thereto.
[0240] Information on the types of privacy protection methods may indicate the methods / algorithms used to apply privacy protection optimization. The methods / algorithms used to apply privacy protection optimization may include delegation, blurring, substitution, masking, etc., in the application.
[0241] Information on the type of privacy protection method may take various forms and may be expressed by various names. For example, information on the type of privacy protection method may be a syntax element or a syntax structure containing one or more syntax elements. For example, information on the type of privacy protection method that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits or variable-length bits. Information on the type of privacy protection method that is a syntax element may be expressed as eoi_privacy_protection_method_flag or eoi_privacy_protection_method_idc, but is not limited thereto.
[0242] The personal information type information may indicate the types of personal information protected by encoding optimization. The types of personal information protected by encoding optimization may include an individual's face, vehicle license plates, location information, etc.
[0243] Personal information type information may take various forms and may be expressed by various names. For example, personal information type information may be a syntax element or a syntax structure containing one or more syntax elements. For example, personal information type information that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits or variable-length bits. Personal information type information that is a syntax element may be expressed as eoi_privacy_info_type, eoi_privacy_info_type_flag, or eoi_privacy_info_type_idc, but is not limited thereto.
[0244] The decoding device can obtain encoder optimization information (S720).
[0245] For example, the processor of a decoding device can process encoder optimization information SEI messages. Based on processing the encoder optimization information SEI messages, the decoding device can obtain information regarding encoder optimization.
[0246] The encoder optimization information SEI message may include an EOI identifier for identifying the encoder optimization information SEI message, EOI cancellation information associated with the persistence of the encoder optimization information SEI message, and EOI persistence information. Additionally, the encoder optimization information SEI message may further include EOI type information indicating the encoder optimization type.
[0247] In this way, identifying multiple encoder optimization information SEI messages using EOI identifiers can provide various technical effects. According to one embodiment, two or more encoder optimization information SEI messages having different EOI identifiers may exist simultaneously. That is, multiple encoder optimization information SEI messages having different EOI identifiers can be included within a single Access Unit or Current Layer, allowing the scope of application and persistence of each encoder optimization information to be defined independently.
[0248] Furthermore, the decoding device can determine the association with a previous encoder optimization information SEI message having the same EOI identifier by referring to the EOI identifier received from the encoding device.
[0249] As such, multiple encoder optimization information SEI messages can have different EOI identifiers and persistence information, and accordingly, mutually independent and parallel recognition, processing, and management are possible for each encoder optimization type.
[0250] Furthermore, multiple encoder optimization information SEI messages having different EOI identifiers may exist simultaneously, and each SEI message may correspond to a different encoder optimization type or purpose. Accordingly, even in situations where different EOI type information is applied in parallel by multiple encoder optimization information SEI messages, by distinguishing between a not applied state and an unspecified state for each optimization type, conflicts in application scope and mutual interference between optimizations can be prevented.
[0251] FIG. 8 is a diagram illustrating a method for encoding image information according to one embodiment of the present disclosure.
[0252] The terms or names described in FIG. 8 (e.g., names of syntax elements or variables, etc.) are merely examples, and the technical features of the present disclosure are not limited to the terms, etc. described in FIG. 8. For example, the image information described in FIG. 8 may include various information according to the embodiments described in the present disclosure and may include information described in at least one of the tables described above.
[0253] The encoding method (S800) may include operations described below. The operations described below do not constitute an essential component of the decoding method according to one embodiment, and at least some of the operations described below may be omitted. Furthermore, the operations described below do not constitute a sufficient component of the encoding method according to one embodiment, and the previously described operations may be added. Moreover, unless the operations described below contradict the previously described operations, they form an embodiment integrally with the previously described operations and do not form a separate embodiment distinct from the previously described operations.
[0254] The encoding method (S800) may be executed by an encoding device including a memory and a processor electrically connected to the memory, for example, by a processor.
[0255] The encoding device can perform encoder optimization (S810).
[0256] For example, the processor of the encoding device can perform encoder optimization. For example, encoder optimization may include optimization for human viewing, optimization for machine analysis, object-based optimization, temporal resampling optimization, spatial resampling optimization, temporal quality optimization, spatial quality optimization, or privacy optimization.
[0257] The encoding device can encode video information (S820).
[0258] For example, the processor of the encoding device can generate information about encoder optimization based on the performed encoder optimization, generate an encoder optimization information (EOI) SEI (supplemental enhancement information) message based on the generated information about encoder optimization, and encode image information including the encoder optimization information SEI message.
[0259] The encoder optimization information SEI message may include information regarding SEI message handling and information regarding encoder optimization. For example, the encoder optimization information SEI message may provide encoder optimization information indicating whether the video is optimized for human viewing or machine analysis, and what type of optimization was applied during the preprocessing or encoding process.
[0260] Encoder optimization information SEI messages can have various names. For example, an encoder optimization information SEI message may be referred to as an EOI SEI message, an EOI SEI related message, EOI SEI information, EOI SEI related information, encoder optimization information message, EOI message, EOI related message, EOI information, EOI related information, etc.
[0261] The encoder optimization information SEI message can take various forms. For example, the encoder optimization information SEI message may be a syntax element or a syntax structure containing one or more syntax elements. Additionally, the encoder optimization information SEI message may be a raw byte sequence payload (RBSP) containing one or more syntax elements or one or more syntax structures. For example, the encoder optimization information SEI message may be represented as encoder_optimization_info(payloadSize), but is not limited thereto.
[0262] According to one embodiment, an encoder optimization information SEI message may include an EOI identifier for identifying the encoder optimization information SEI message, EOI cancellation information and EOI persistence information associated with the persistence of the encoder optimization information SEI message, and EOI type information. Additionally, the encoder optimization information SEI message may further include human viewing optimization information, machine analysis optimization information, object-based optimization type information, temporal resampling type information, information regarding spatial resampling type, privacy protection method type information and / or privacy type information.
[0263] The EOI identifier, EOI cancellation information associated with the persistence of the encoder optimization information SEI message, EOI persistence information, EOI type information, human viewing optimization information, machine analysis optimization information, object-based optimization type information, temporal resampling type information, information regarding spatial resampling type, privacy method type information and / or privacy type information may be identical to the EOI cancellation information associated with the persistence of the encoder optimization information SEI message, EOI persistence information, EOI type information, human viewing optimization information, machine analysis optimization information, object-based optimization type information, temporal resampling type information, information regarding spatial resampling type, privacy method type information and / or privacy type information described above in operation 710 of FIG. 7.
[0264] The description of the EOI identifier, EOI cancellation information associated with the persistence of the encoder optimization information SEI message, EOI persistence information, EOI type information, human viewing optimization information, machine analysis optimization information, object-based optimization type information, temporal resampling type information, information regarding spatial resampling type, privacy method type information and / or privacy type information refers to the description of the EOI identifier, EOI cancellation information associated with the persistence of the encoder optimization information SEI message, EOI persistence information, EOI type information, human viewing optimization information, machine analysis optimization information, object-based optimization type information, temporal resampling type information, information regarding spatial resampling type, privacy method type information and / or privacy type information described above in operation 710 of FIG. 7.
[0265] For example, EOI cancellation information may be generated based on whether the persistence of an encoder optimization information SEI message having the same EOI identifier as the encoder optimization information SEI message of the current picture is terminated among encoder optimization information SEI messages for the previous picture of the current picture included in the same unit as the encoder optimization information SEI message.
[0266] For example, EOI cancellation information can be generated based on whether the encoder optimization information according to the encoder optimization information SEI message is applied to the current picture and subsequent picture included in the same unit as the encoder optimization information SEI message.
[0267] For example, EOI persistence information can be generated based on whether encoder optimization information according to the encoder optimization information SEI message is applied only to the current picture.
[0268] For example, EOI persistence information may be generated based on whether encoder optimization information according to an encoder optimization information SEI message is applied to the current picture and subsequent picture included in the same unit as the encoder optimization information SEI message until a predefined termination condition is satisfied. Here, the termination condition may be at least one of when a new encoded layered image sequence (CLVS) starts within the same unit, when a bitstream ends, or when another encoder optimization information SEI message having the same EOI identifier as the current picture is included within the same unit.
[0269] In this way, the encoding device can encode video information including encoder optimization information SEI messages. The video information may simultaneously include two or more encoder optimization information SEI messages having different EOI identifiers.
[0270] Here, the encoder optimization information SEI message may include an EOI identifier for identifying the encoder optimization information SEI message.
[0271] In this way, identifying multiple encoder optimization information SEI messages using EOI identifiers can provide various technical effects. According to one embodiment, two or more encoder optimization information SEI messages having different EOI identifiers may exist simultaneously. That is, multiple encoder optimization information SEI messages having different EOI identifiers can be included within a single Access Unit or Current Layer, allowing the scope of application and persistence of each encoder optimization information to be defined independently.
[0272] In other words, as such, multiple encoder optimization information SEI messages can have different EOI identifiers and persistence information, and accordingly, mutually independent and parallel recognition, processing, and management are possible for each encoder optimization type.
[0273] In addition, the encoder optimization information SEI message may further include EOI type information indicating the encoder optimization type. Accordingly, even in situations where different EOI type information is applied in parallel by multiple encoder optimization information SEI messages, by distinguishing between a not applied state and an unspecified state for each optimization type, conflicts in the application scope and mutual interference between optimizations can be prevented.
[0274] A bitstream is generated based on video information encoded according to the encoding method (S800) described above, and the bitstream can be stored on a computer-readable storage medium.
[0275] In addition, a bitstream is generated based on video information encoded according to the encoding method (S800) described above, and the bitstream can be transmitted through a transmission unit and / or a transmission medium.
[0276] In this disclosure, as an example, the names of the syntax elements described above are all arbitrarily designated for clarity of explanation and are not intended to limit the names of the syntax elements. Additionally, each syntax element may be referred to as information. Furthermore, syntax elements may be obtained from a bitstream, but may also be derived from other syntax elements, and this may also be included in the embodiments of this disclosure.
[0277] In addition, the bitstream generated by the video encoding method may be stored on a non-transient computer-readable recording medium.
[0278] In addition, as another example, a bitstream generated by a video encoding method may be transmitted to another device (e.g., a video decoding device, etc.). In this case, the method of transmitting the bitstream may include a process of transmitting the bitstream.
[0279] The exemplary methods of the present disclosure are described as a series of operations for clarity of description, but this is not intended to limit the order in which the steps are performed, and if necessary, each step may be performed simultaneously or in a different order. To implement the method according to the present disclosure, additional steps may be included in addition to the steps exemplified, steps excluding some steps and including the remaining steps, or steps excluding some steps and including additional steps.
[0280] In the present disclosure, an image encoding device or an image decoding device performing a predetermined operation (step) may perform an operation (step) to check the conditions or circumstances for performing the said operation (step). For example, if it is stated that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or the image decoding device may perform an operation to check whether the said predetermined condition is satisfied, and then perform the said predetermined operation.
[0281] The various embodiments of the present disclosure are not intended to list all possible combinations but to describe representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combination of two or more.
[0282] The embodiments described in this disclosure may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each figure may be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for implementation may be stored in a digital storage medium.
[0283] In addition, the decoder (decoding device) and encoder (encoding device) to which the embodiment(s) of the present disclosure are applied may be included in multimedia broadcasting transmission and reception devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video conversation devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, Video on Demand (VoD) service providers, OTT video (Over the top video) devices, internet streaming service providers, 3D video devices, VR (virtual reality) devices, AR (argument reality) devices, video phone video devices, transportation terminals (e.g., vehicle (including autonomous vehicle) terminals, robot terminals, airplane terminals, ship terminals, etc.), and medical video devices, and may be used to process video signals or data signals. For example, OTT (Over the top video) devices may include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, DVRs (Digital Video Recorders), etc.
[0284] Additionally, the processing method to which the embodiment(s) of the present disclosure are applied may be produced in the form of a program that is executed by a computer and may be stored on a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of this document may also be stored on a computer-readable recording medium. A computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. A computer-readable recording medium may include, for example, a Blu-ray disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. Additionally, a computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission over the Internet). Additionally, a bitstream generated by an encoding method may be stored on a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0285] Additionally, the embodiment(s) of the present disclosure may be implemented as a computer program product by program code, and the program code may be executed on a computer by the embodiment(s) of the present disclosure. The program code may be stored on a carrier that is readable by a computer.
[0286] FIG. 9 is a drawing showing an example of a content streaming system to which embodiments of the present disclosure can be applied.
[0287] Referring to FIG. 9, a content streaming system to which an embodiment of the present disclosure is applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0288] The encoding server compresses content input from multimedia input devices, such as smartphones, cameras, and camcorders, into digital data to generate a bitstream, and transmits it to the streaming server. As another example, if multimedia input devices, such as smartphones, cameras, and camcorders, generate the bitstream directly, the encoding server can be omitted.
[0289] A bitstream may be generated by a video encoding method and / or video encoding device to which an embodiment of the present disclosure is applied, and a streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0290] A streaming server transmits multimedia data to a user device based on user requests made through a web server, and the web server can act as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server forwards the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which can perform the role of controlling commands and responses between devices within the content streaming system.
[0291] A streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a seamless streaming service, the streaming server may store the bitstream for a certain period of time.
[0292] Examples of user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs (head-mounted displays)), digital TVs, desktop computers, digital signage, etc.
[0293] Each server within the content streaming system can be operated as a distributed server, and in this case, data received from each server can be processed in a distributed manner.
[0294] FIG. 10 is a drawing showing another example of a content streaming system to which embodiments of the present disclosure may be applied.
[0295] Referring to FIG. 10, in an embodiment such as VCM, a task may be performed at a user terminal or at an external device (e.g., a streaming server, an analysis server, etc.) depending on the performance of the device, the user's request, the characteristics of the task to be performed, etc. In this way, in order to transmit information necessary for performing the task to an external device, the user terminal may generate a bitstream containing information necessary for performing the task (e.g., information such as the task, the neural network, and / or the purpose) directly or through an encoding server.
[0296] The analysis server can decode encoded information received from the user terminal (or from the encoding server) and then execute the task requested by the user terminal. The analysis server can transmit the results obtained through task execution back to the user terminal or to other associated service servers (e.g., web servers). For example, the analysis server can transmit the results obtained from performing a task to identify a fire to a fire-related server. The analysis server may include a separate control server, in which case the control server can play a role in controlling commands and responses between each device and server associated with the analysis server. Additionally, the analysis server may request desired information from the web server based on information regarding tasks that the user device intends to perform and tasks that it can perform. When the analysis server requests a desired service from the web server, the web server forwards it to the analysis server, and the analysis server can transmit the corresponding data to the user terminal. In this case, the control server of the content streaming system can perform the role of controlling commands and responses between each device within the streaming system.
[0297] An embodiment according to the present disclosure can be used to encode / decode video.
Claims
1. In a video decoding method performed by a decoding device, A step of obtaining an Encoder Optimization Information (SEI) supplemental enhancement information message associated with picture optimization from a bitstream; and It includes the step of obtaining encoder optimization information based on the above EOI SEI message, and An image decoding method characterized in that the above EOI SEI message includes an EOI identifier for identifying the above EOI SEI message, EOI cancellation information associated with the persistence of the above EOI SEI message, and EOI persistence information.
2. In Claim 1, An image decoding method characterized by the simultaneous existence of two or more EOI SEI messages corresponding to different EOI identifiers.
3. In Claim 1, A video decoding method characterized in that, when the value of the above EOI cancellation information is 1, the persistence of an EOI SEI message having the same EOI identifier as the EOI SEI message of the current picture among the EOI SEI messages for the previous picture of the current picture is terminated.
4. In Claim 3, A video decoding method characterized in that, when the value of the above EOI cancellation information is 0, the encoder optimization information according to the above EOI SEI message is applied to the current picture and the subsequent picture.
5. In Claim 3, A video decoding method characterized in that, when the value of the above EOI persistence information is 0, the encoder optimization information according to the above EOI SEI message is applied only to the current picture.
6. In Claim 3, A video decoding method characterized in that, when the value of the EOI persistence information is 1, the encoder optimization information according to the EOI SEI message is applied to the current picture and subsequent picture until a predefined termination condition is satisfied.
7. In Claim 6, An image decoding method characterized by the above termination condition being at least one of the following: when a new encoded layer image sequence (CLVS) starts within the same layer, when a bitstream ends, or when another EOI SEI message having the same EOI identifier as the current picture is included within the same layer.
8. In Claim 1, An image decoding method characterized in that the value of the above EOI identifier is in the range of 0 to 63.
9. In Claim 1, An image decoding method characterized in that the above EOI SEI message further includes EOI type information indicating an encoder optimization type corresponding to a predefined bitMask value.
10. In Claim 9, A video decoding method characterized by determining whether to apply an encoder optimization type corresponding to the bitmask value to the current picture based on the EOI type information and the bitmask value.
11. In Claim 10, A video decoding method characterized in that when the value of the EOI type information is greater than 0 and the bit value corresponding to the bitmask value in the value of the EOI type information is 0, the encoder optimization type corresponding to the bitmask value is unspecified.
12. A video encoding method performed by an encoding device, Step of performing encoder optimization; and The method includes the step of encoding video information containing an Encoder Optimization Information SEI (supplemental enhancement information) message generated based on the information regarding the above encoder optimization, and An image encoding method characterized in that the above EOI SEI message includes an EOI identifier for identifying the above EOI SEI message, EOI cancellation information associated with the persistence of the above EOI SEI message, and EOI persistence information.
13. In Claim 12, An image encoding method characterized by the above image information simultaneously including two or more EOI SEI messages according to different EOI identifiers.
14. In Claim 12, A video encoding method characterized by generating the above EOI cancellation information based on whether the persistence of an EOI SEI message having the same EOI identifier as the EOI SEI message of the current picture is terminated among the EOI SEI messages for the previous picture of the current picture.
15. In Claim 12, A video encoding method characterized in that the above EOI persistence information is generated based on whether encoder optimization information according to the above EOI SEI message is applied to the current picture and subsequent picture until a predefined termination condition is satisfied.
16. In Claim 12, An image encoding method characterized in that the value of the above EOI identifier is in the range of 0 to 63.
17. In Claim 12, A video encoding method characterized in that the above EOI SEI message further includes EOI type information indicating an encoder optimization type corresponding to a predefined bitmask value.
18. In Claim 17, A video encoding method characterized in that the bit value corresponding to the bitmask value in the above EOI type information or the value of the above EOI type information is generated based on whether the encoder optimization type corresponding to the bitmask value is unspecified in the current picture.
19. In a computer-readable storage medium for storing a bitstream, The above storage medium stores a bitstream for video information including an encoder optimization information SEI (supplemental enhancement information) message generated based on information regarding encoder optimization, and A storage medium characterized in that the above EOI SEI message includes an EOI identifier for identifying the above EOI SEI message, EOI cancellation information associated with the persistence of the above EOI SEI message, and EOI persistence information.
20. A method for transmitting data regarding an image, wherein the method comprises the steps of: performing encoder optimization and obtaining image information including an Encoder Optimization Information SEI (supplemental enhancement information) message generated based on information regarding the encoder optimization; and The method includes the step of transmitting the data including the above image information, A bitstream transmission method characterized in that the above EOI SEI message includes an EOI identifier for identifying the above EOI SEI message, EOI cancellation information associated with the persistence of the above EOI SEI message, and EOI persistence information.