Method for decoding image information, method for encoding image information, method related to bitstream, and computer-readable storage medium for storing bitstream

By deriving and encoding TDI SEI messages with purpose and persistence information, the method addresses high-resolution video transmission challenges, enhancing coding system reliability and efficiency.

WO2026084540A1PCT designated stage Publication Date: 2026-04-23LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2025-10-20
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

The increasing demand for high-resolution, high-quality video has led to higher transmission and storage costs due to the increase in transmitted information or bits, necessitating high-efficiency video compression technology to minimize confusion and malfunction in coding systems while improving reliability and efficiency.

Method used

The method involves deriving and encoding text description information (TDI) SEI messages with purpose, identifier, and persistence information to clarify the persistence of these messages, ensuring clear target TDI SEI message persistence is maintained or cancelled, thereby improving coding system reliability and efficiency.

Benefits of technology

This approach reduces confusion and malfunction in coding systems, enhances reliability, and improves data transmission efficiency by clarifying the persistence of TDI SEI messages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025016548_23042026_PF_FP_ABST
    Figure KR2025016548_23042026_PF_FP_ABST
Patent Text Reader

Abstract

A method according to the present disclosure comprises: acquiring image information including a text description information (TDI) supplemental enhancement information (SEI) message; and deriving text description information for at least one picture on the basis of the text description information SEI message, wherein the text description information SEI message includes description purpose information indicating a purpose of the text description information SEI message, identifier information for identifying the text description information SEI message, and persistence information indicating persistence of the text description information SEI message, and the persistence of the text description information SEI message is derived on the basis of the description purpose information and / or the identifier information.
Need to check novelty before this filing date? Find Prior Art

Description

A method for decoding image information, a method for encoding image information, a method relating to a bitstream, and a computer-readable storage medium for storing a bitstream

[0001] The present disclosure relates to a method for decoding image information, a method for encoding image information, a method for bitstreams, and a computer-readable storage medium for storing bitstreams.

[0002] Recently, the demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition), has been increasing across various fields. As video data becomes higher in resolution and quality, the relative amount of information or bits transmitted increases compared to conventional video data. This increase in transmitted information or bits leads to higher transmission and storage costs.

[0003] Accordingly, high-efficiency video compression technology is required to effectively transmit, store, and play back high-resolution, high-quality video information.

[0004] The present disclosure aims to suppress, prevent, or minimize confusion and / or malfunction in coding systems.

[0005] The present disclosure aims to improve the reliability of a coding system including an encoding device and a decoding device.

[0006] The present disclosure aims to improve the coding efficiency of a coding system including an encoding device and a decoding device.

[0007] The present disclosure aims to improve the data transmission efficiency of a coding system including an encoding device and a decoding device.

[0008] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure belongs from the description below.

[0009] A method for decoding image information according to one aspect of the present disclosure comprises: acquiring the image information including a text description information (TDI) SEI (supplemental enhancement information) message; and deriving text description information for at least one picture based on the text description information SEI message, wherein the text description information SEI message includes descriptive purpose information indicating the purpose of the text description information SEI message, identifier information identifying the text description information SEI message, and persistence information indicating the persistence of the text description information SEI message, and the persistence of the text description information SEI message is derived based on the descriptive purpose information and / or the identifier information.

[0010] According to one aspect of the present disclosure, an apparatus for decoding image information comprises a memory and a processor connected to the memory, wherein the processor acquires the image information including a text description information (TDI) SEI (supplemental enhancement information) message; and derives text description information for at least one picture based on the text description information SEI message, wherein the text description information SEI message includes descriptive purpose information indicating the purpose of the text description information SEI message, identifier information identifying the text description information SEI message, and persistence information indicating the persistence of the text description information SEI message, and the persistence of the text description information SEI message is derived based on the descriptive purpose information and / or the identifier information.

[0011] In a method or device for decoding the above image information, the persistence information may indicate the persistence of an SEI message that includes the same purpose as the description purpose information and / or the same identifier as the identifier information.

[0012] In a method or device for decoding the above image information, based on the persistence information having a value such as 0, an SEI message containing the same purpose as the description purpose information and / or the same identifier as the identifier information may be applied only to the current picture.

[0013] In a method or device for decoding the above image information, based on the persistence information having a value such as 1, an SEI message including the same purpose as the description purpose information and / or the same identifier as the identifier information may be applied to the current picture and at least one picture following the current picture.

[0014] In a method or device for decoding the above image information, based on the persistence information having a value such as 1, an SEI message containing the same purpose as the descriptive purpose information and / or the same identifier as the identifier information may be applied to at least one picture following the current picture until the current picture and the picture associated with the SEI message containing the same purpose as the descriptive purpose information and / or the same identifier as the identifier information are output.

[0015] In the method or device for decoding the above image information, the identifier information may have a value from 1 to 8,191.

[0016] A method for encoding image information according to one aspect of the present disclosure comprises generating text description information for at least one picture; and encoding the image information including a text description information (TDI) SEI (supplemental enhancement information) message generated based on the text description information for at least one picture, wherein the text description information SEI message includes description purpose information indicating the purpose of the text description information SEI message, identifier information identifying the text description information SEI message, and persistence information indicating the persistence of the text description information SEI message, and the persistence of the text description information SEI message is determined based on the description purpose information and / or the identifier information.

[0017] According to one aspect of the present disclosure, an apparatus for encoding image information comprises a memory and a processor connected to the memory, wherein the processor generates text description information for at least one picture; and encodes the image information including a text description information (TDI) SEI (supplemental enhancement information) message generated based on the text description information for at least one picture, wherein the text description information SEI message includes description purpose information indicating the purpose of the text description information SEI message, identifier information identifying the text description information SEI message, and persistence information indicating the persistence of the text description information SEI message, and the persistence of the text description information SEI message is determined based on the description purpose information and / or the identifier information.

[0018] In a method or device for encoding the above-mentioned image information, the persistence information may be generated based on the persistence of an SEI message containing the same purpose as the above-mentioned descriptive purpose information and / or the same identifier as the above-mentioned identifier information.

[0019] In a method or device for encoding the above-mentioned image information, the persistence information may have a value such as 0 based on the fact that the SEI message containing the same purpose as the above-mentioned descriptive purpose information and / or the same identifier as the above-mentioned identifier information is applied only to the current picture.

[0020] In a method or device for encoding the above image information, the SEI message containing the same purpose as the above description purpose information and / or the same identifier as the above identifier information may have a persistence information value such as 1 based on being applied to the current picture and at least one picture following the current picture.

[0021] In a method or device for encoding the above-mentioned image information, the persistence information may have a value such as 1 based on the fact that the SEI message containing the same purpose as the above-mentioned descriptive purpose information and / or the same identifier as the above-mentioned identifier information is applied to at least one picture following the current picture until the current picture and the picture associated with the SEI message containing the same purpose as the above-mentioned descriptive purpose information and / or the same identifier as the above-mentioned identifier information are output.

[0022] In the method or device for encoding the above image information, the identifier information may have a value from 1 to 8,191.

[0023] A method for a bitstream according to one aspect of the present disclosure comprises: generating text description information for at least one picture; generating a bitstream based on image information including a text description information (TDI) SEI (supplemental enhancement information) message generated based on the text description information for at least one picture; and transmitting data for the bitstream, wherein the text description information SEI message includes description purpose information indicating the purpose of the text description information SEI message, identifier information identifying the text description information SEI message, and persistence information indicating the persistence of the text description information SEI message, and the persistence of the text description information SEI message is determined based on the description purpose information and / or the identifier information.

[0024] According to one aspect of the present disclosure, an apparatus for a bitstream comprises: at least one processor that generates text description information for at least one picture and generates a bitstream based on image information including a text description information (TDI) SEI (supplemental enhancement information) message generated based on the text description information for at least one picture; and a transmission unit that transmits data for the bitstream, wherein the text description information SEI message includes description purpose information indicating the purpose of the text description information SEI message, identifier information identifying the text description information SEI message, and persistence information indicating the persistence of the text description information SEI message, and the persistence of the text description information SEI message is determined based on the description purpose information and / or the identifier information.

[0025] In the method or apparatus relating to the bitstream above, the persistence information may be generated based on the persistence of an SEI message containing the same purpose as the descriptive purpose information and / or the same identifier as the identifier information.

[0026] According to one aspect of the present disclosure, a computer-readable storage medium for storing a bitstream, wherein the storage medium stores the bitstream generated based on image information including a text description information (TDI) SEI (supplemental enhancement information) message generated based on text description information for at least one picture, the text description information SEI message includes description purpose information indicating the purpose of the text description information SEI message, identifier information identifying the text description information SEI message, and persistence information indicating the persistence of the text description information SEI message, and the persistence of the text description information SEI message is determined based on the description purpose information and / or the identifier information.

[0027] In the above computer-readable storage medium, the persistence information is generated based on the persistence of an SEI message containing the same purpose as the descriptive purpose information and / or the same identifier as the identifier information.

[0028] The features briefly summarized above regarding the present disclosure are merely exemplary aspects of the detailed description of the present disclosure that follows and do not limit the scope of the present disclosure.

[0029] According to the present disclosure, the target TDI SEI message and / or target text description for which persistence is maintained or cancelled becomes clear, and confusion or malfunction in a coding system due to the ambiguity of the target TDI SEI message and / or target text description for which persistence is maintained or cancelled can be suppressed, prevented, or minimized.

[0030] According to the present disclosure, the reliability of a coding system including an encoding device and a decoding device can be improved as the target TDI SEI message and / or target text description, for which persistence is maintained or cancelled, becomes clear.

[0031] According to the present disclosure, the coding efficiency of a coding system including an encoding device and a decoding device can be improved as the target TDI SEI message and / or target text description, for which persistence is maintained or cancelled, becomes clear.

[0032] According to the present disclosure, the data transmission efficiency of a coding system including an encoding device and a decoding device can be improved as the target TDI SEI message and / or target text description, for which persistence is maintained or cancelled, becomes clear.

[0033] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs from the description below.

[0034] FIG. 1 is a schematic diagram illustrating a video coding system to which an embodiment according to the present disclosure can be applied.

[0035] FIG. 2 is a schematic diagram showing an encoding device to which an embodiment according to the present disclosure can be applied.

[0036] FIG. 3 is a schematic diagram showing a decoding device to which an embodiment according to the present disclosure can be applied.

[0037] Figure 4 illustrates an exemplary hierarchical structure for a coded video / image.

[0038] FIG. 5 is a diagram illustrating a method for decoding image information according to one embodiment of the present disclosure.

[0039] FIG. 6 is a diagram illustrating a method for encoding image information according to one embodiment of the present disclosure.

[0040] FIG. 7 is a drawing illustrating an exemplary content streaming system to which an embodiment according to the present disclosure can be applied.

[0041] Hereinafter, embodiments of the present disclosure are described in detail with reference to the attached drawings so that those skilled in the art can easily implement them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein.

[0042] In describing the embodiments of the present disclosure, detailed descriptions of known configurations or functions are omitted if it is determined that such descriptions could obscure the essence of the present disclosure. Additionally, parts of the drawings unrelated to the description of the present disclosure have been omitted, and similar parts are denoted by similar reference numerals.

[0043] In the present disclosure, when a component is described as being "connected," "combined," or "joined" with another component, this may include not only a direct connection but also an indirect connection in which another component exists in between. Furthermore, when a component is described as "comprising" or "having" another component, this means that, unless specifically stated otherwise, it does not exclude the other component but may include an additional component.

[0044] In the present disclosure, terms such as first, second, etc. are used solely for the purpose of distinguishing one component from another and do not limit the order or importance of the components unless specifically stated otherwise. Accordingly, within the scope of the present disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and likewise, a second component in one embodiment may be referred to as a first component in another embodiment.

[0045] In this disclosure, distinct components are intended to clearly describe their respective features and do not imply that the components are separate. That is, multiple components may be integrated to form a single hardware or software unit, or a single component may be distributed to form multiple hardware or software units. Accordingly, such integrated or distributed embodiments are included within the scope of this disclosure, unless otherwise noted.

[0046] In the present disclosure, the components described in various embodiments do not necessarily mean essential components, and some may be optional components. Accordingly, embodiments consisting of a subset of the components described in one embodiment are also included within the scope of the present disclosure. Furthermore, embodiments including additional components in addition to the components described in various embodiments are also included within the scope of the present disclosure.

[0047] The present disclosure relates to the encoding and decoding of images. For example, the methods and embodiments disclosed in this document may be applied to methods disclosed in the VVC (versatile video coding) standard, EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard) or next-generation video / image coding standards (e.g., H.267 or H.268).

[0048] The present disclosure presents various embodiments relating to video / image coding, and unless otherwise stated, said embodiments may be performed in combination with one another.

[0049] Unless newly defined in this disclosure, the terms used herein may have the ordinary meanings commonly used in the technical field to which this disclosure belongs.

[0050] In this disclosure, "video" may refer to a set of images over time. In this disclosure, "picture" generally refers to a unit representing a single image at a specific time, and a slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles. A brick may represent a rectangular area of ​​rows of CTUs within a tile in a picture. In this document, tile groups and slices may be used interchangeably. For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.

[0051] In the present disclosure, "pixel" or "pel" may refer to the smallest unit constituting a picture (or image). Additionally, "sample" may be used as a term corresponding to pixel. A sample may generally represent a pixel or a pixel value, may represent only the pixel / pixel value of the luminance component, or may represent only the pixel / pixel value of the chroma component.

[0052] In this disclosure, "unit" may represent a basic unit of image processing. A unit may include at least one of a specific area of ​​a picture and information related to that area. A unit may include one luminance block and two chroma (e.g., cb, cr) blocks. Depending on the case, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows.

[0053] In the present disclosure, "current block" may mean one of "current coding block," "current coding unit," "block to be encoded," "block to be decoded," or "block to be processed." When prediction is performed, "current block" may mean "current prediction block" or "block to be predicted." When transformation (inverse transformation) / quantization (inverse quantization) is performed, "current block" may mean "current transformation block" or "block to be transformed." When filtering is performed, "current block" may mean "block to be filtered."

[0054] In the present disclosure, "current block" may mean a block comprising both a luminous component block and a chroma component block, or "luma block of the current block," unless explicitly stated as a chroma block. The luminous component block of the current block may be expressed by including an explicit description of a luminous component block, such as "luma block" or "current luminous block." Additionally, the chroma component block of the current block may be expressed by including an explicit description of a chroma component block, such as "chroma block" or "current chroma block."

[0055] In the present disclosure, " / " and "," may be interpreted as "and / or." For example, "A / B" and "A, B" may be interpreted as "A and / or B." Additionally, "A / B / C" and "A, B, C" may mean "at least one of A, B and / or C."

[0056] In the present disclosure, "or" may be interpreted as "and / or". For example, "A or B" may mean 1) "A" only, 2) "B" only, or 3) "A and B". Alternatively, in the present disclosure, "or" may mean "additionally or alternatively".

[0057] FIG. 1 is a schematic diagram illustrating a video / image coding system to which an embodiment according to the present disclosure can be applied.

[0058] Referring to FIG. 1, a video / image coding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image or data in the form of a file or streaming to the receiving device via a digital storage medium or a network.

[0059] The source device may include a video source, an encoding device, and a transmission unit. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.

[0060] A video source may acquire video / images through processes such as video / image capture, synthesis, or generation. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include, for example, a computer, a tablet, and a smartphone, etc., and may generate video / images (electronically). For example, virtual video / images may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process in which related data is generated.

[0061] The encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0062] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device in the form of a file or streaming via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit may include elements for creating a media file through a predetermined file format and elements for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.

[0063] The decoding device can decode video / images by performing a series of procedures such as inverse quantization, inverse transform, and prediction corresponding to the operation of the encoding device.

[0064] The renderer can render the decoded video / image. The rendered video / image can be displayed through the display unit.

[0065] FIG. 2 is a schematic diagram illustrating an encoding device to which an embodiment according to the present disclosure can be applied.

[0066] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a predictor (220), a residual processor (230), an entropy encoder (240), an adder (250), a filter (260), and a memory (270). The predictor (220) may include an inter-predictor (221) and an intra-predictor (222). The residual processor (230) may include a transformer (232), a quantizer (233), a dequantizer (234), and an inverse transformer (235). The residual processor (230) may further include a subtractor (231). The addition unit (250) may be referred to as a reconstructor or a reconstructed block generator. The above-described image segmentation unit (210), prediction unit (220), residual processing unit (230), entropy encoding unit (240), addition unit (250), and filtering unit (260) may be configured by one or more hardware components (e.g., an encoder chipset or processor) according to the embodiment. Additionally, the memory (270) may include a DPB (Decoded Picture Buffer) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.

[0067] The image segmentation unit (210) can divide an input image (or picture, frame) input to an encoding device (200) into one or more processing units. For example, the processing unit may be called a coding unit (CU). A coding unit may be recursively divided into a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, a single coding unit may be divided into multiple coding units of a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. For example, a quad-tree structure may be applied first, and a binary-tree structure and / or a ternary-tree structure may be applied later. Alternatively, a binary-tree structure may be applied first. A coding procedure according to the present disclosure may be performed based on the final coding unit that is no longer divided. In this case, based on coding efficiency according to image characteristics, the maximum coding unit may be used directly as the final coding unit, or, if necessary, the maximum coding unit may be recursively divided into lower-depth coding units so that a coding unit of the optimal size is used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). The prediction unit and the transformation unit may each be divided or partitioned from the final coding unit.The above prediction unit may be a unit of sample prediction, and the above transformation unit may be a unit that derives transformation coefficients and / or a unit that derives a residual signal from transformation coefficients.

[0068] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the context. In general, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chroma component. A sample may be used to refer to a single picture (or image) as a term corresponding to a pixel or pel.

[0069] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (predicted block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input image signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, as illustrated, the unit that subtracts the prediction signal (predicted block, prediction sample array) from the input image signal (original block, original sample array) within the encoding device (200) may be called a subtraction unit (231). The prediction unit (220) can perform a prediction for a block to be processed (hereinafter, current block) and generate a predicted block (predicted block) containing prediction samples for said current block. The prediction unit (220) can determine whether intra prediction is applied or inter prediction is applied in units of the current block or CU. The prediction unit (220) can generate various information regarding prediction, such as prediction mode information, as described below in the description of each prediction mode, and transmit it to the entropy encoding unit (240). The information regarding prediction can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.

[0070] The intra prediction unit (222) can predict the current block by referring to samples within the current picture. The referenced samples may be located near the current block or away from it, depending on the prediction mode. In intra prediction, the prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a Planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is merely an example, and depending on the settings, more or fewer directional prediction modes may be used. The intra prediction unit (222) may also determine the prediction mode applied to the current block by using the prediction mode applied to the surrounding blocks.

[0071] The inter prediction unit (221) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture containing the reference blocks and the reference picture containing the temporal neighboring blocks may be the same or different from each other. The temporal neighboring blocks may be referred to by names such as collocated reference block, collocated CU (colCU), etc. A reference picture containing the aforementioned temporal surrounding blocks may be called a collocated picture (colPic). For example, the inter prediction unit (221) may construct a list of motion information candidates based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of surrounding blocks as motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0072] The prediction unit (220) may generate a prediction signal based on various prediction methods and / or prediction techniques described below. For example, the prediction unit (220) may apply intra prediction or inter prediction for the prediction of the current block, as well as apply intra prediction and inter prediction simultaneously. A prediction method that applies intra prediction and inter prediction simultaneously for the prediction of the current block may be called combined inter and intra prediction (CIIP). Additionally, the prediction unit (220) may be based on an intra block copy (IBC) prediction mode or a palette mode for the prediction of the block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, for example, screen content coding (SCC). IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. Palette mode can be viewed as an example of intra-coding or intra-prediction. When palette mode is applied, sample values ​​within a picture can be signaled based on information regarding palette tables and palette indices.

[0073] The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or to generate a residual signal. The subtraction unit (231) can generate a residual signal (residual signal, residual block, residual sample array) by subtracting the prediction signal (predicted block, prediction sample array) output from the prediction unit (220) from the input image signal (original block, original sample array). The generated residual signal can be transmitted to the conversion unit (232).

[0074] The transformation unit (232) can generate transform coefficients by applying a transformation technique to a residual signal. For example, the transformation technique may include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transformation obtained based on a prediction signal generated using all previously reconstructed pixels. The transformation process may be applied to a block of pixels of the same size in a square, or to a block of variable size that is not square.

[0075] The quantization unit (233) can quantize the transformation coefficients and transmit them to the entropy encoding unit (240). The entropy encoding unit (240) can encode the quantized signal (information regarding the quantized transformation coefficients) and output it as a bitstream. The information regarding the quantized transformation coefficients may be called residual information. The quantization unit (233) can rearrange the block-shaped quantized transformation coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information regarding the quantized transformation coefficients based on the one-dimensional vector-shaped quantized transformation coefficients.

[0076] The entropy encoding unit (240) can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit (190) may encode information required for video / image restoration (e.g., values ​​of syntax elements) together or separately, in addition to quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may further include general constraint information. The signaling information, transmitted information, and / or syntax elements mentioned in the present disclosure may be included in the video / image information. The video / image information may be encoded through the encoding procedure described above and included in the bitstream.

[0077] The above bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmitting a signal output from the entropy encoding unit (240) and / or a storage unit (not shown) for storing it may be provided as an internal / external element of the encoding device (200), or the transmission unit may be provided as a component of the entropy encoding unit (240).

[0078] The quantized transformation coefficients output from the quantization unit (233) can be used to generate a residual signal. For example, a residual signal (residual block or residual samples) can be restored by applying inverse quantization and inverse transformation to the quantized transformation coefficients through the inverse quantization unit (234) and the inverse transformation unit (235).

[0079] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.

[0080] The adder (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit (221) or the intra prediction unit (222). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The adder (250) may be called a reconstructed unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after undergoing filtering as described below.

[0081] The filtering unit (260) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (260) can generate a modified restored picture by applying various filtering methods to the restored picture, and can store the modified restored picture in memory (270), specifically in the DPB of memory (170). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (240), as described below in the description of each filtering method. The information regarding filtering can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.

[0082] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter-prediction unit (221). Through this, the encoding device (200) can avoid prediction mismatches between the encoding device (200) and the decoding device when inter-prediction is applied, and can also improve encoding efficiency.

[0083] The DPB in memory (270) can store a modified restored picture to be used as a reference picture in the inter prediction unit (221). Memory (270) can store motion information of blocks from which motion information is derived (or encoded) in the current picture and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter prediction unit (221) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory (270) can store restoration samples of restored blocks in the current picture and transmit them to the intra prediction unit (222).

[0084] FIG. 3 is a schematic diagram illustrating a decoding device to which an embodiment according to the present disclosure can be applied.

[0085] As illustrated in FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-predictor (332) and an intra-predictor (331). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321). The aforementioned entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) may be configured by a single hardware component (e.g., a decoder chipset or a processor) according to an embodiment. Additionally, the memory (360) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.

[0086] When a bitstream containing video / image information is input, the decoding device (300) can restore the image by performing a process corresponding to the process performed by the encoding device (200) of FIG. 2. For example, the decoding device (300) can perform decoding using a processing unit applied in the encoding device (200). Thus, the processing unit for decoding may be, for example, a coding unit. The coding unit may be a coding tree unit, or a maximum coding unit may be obtained by dividing it according to a quad tree structure, a binary tree structure, and / or a binary tree structure. And, the restored image signal decoded and output through the decoding device (300) can be played back through a playback device (not shown).

[0087] The decoding device (300) can receive a signal output from the encoding device (200) of FIG. 2 in the form of a bitstream. The received signal can be decoded through an entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information necessary for image restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may further include general constraint information. The decoding device (300) can decode the picture based on the information regarding the parameter sets and / or the general constraint information. The signaling / received information and / or syntax elements described below can be obtained from the bitstream by decoding through the decoding procedure. For example, the entropy decoding unit (310) can decode information within the bitstream based on coding methods such as exponential coding, CAVLC, or CABAC, and output values ​​of syntax elements required for image restoration and quantized values ​​of transformation coefficients regarding residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using information on the syntax element to be decoded and decoding information of surrounding and decoding target blocks or information on symbols / bins decoded in the previous step, predict the probability of occurrence of the bin according to the determined context model, and perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin after determining the context model. Among the information decoded in the entropy decoding unit (310), information regarding prediction is provided to the prediction unit (330), and residual values ​​for which entropy decoding was performed in the entropy decoding unit (310), i.e., quantized transformation coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive residual signals (residual blocks, residual samples, residual sample array). Additionally, among the information decoded in the entropy decoding unit (310), information regarding filtering can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (300), or the receiving unit may be a component of the entropy decoding unit (310). Meanwhile, the decoding device according to the present document may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit (310), and the sample decoder may include at least one of the inverse quantization unit (321), inverse transform unit (322), adder (340), filtering unit (350), memory (360), inter prediction unit (332), and intra prediction unit (331).

[0088] In the inverse quantization unit (321), the quantized transformation coefficients can be inversely quantized to output transformation coefficients. The inverse quantization unit (321) can rearrange the quantized transformation coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed in the encoding device (200). The inverse quantization unit (321) can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain transformation coefficients.

[0089] In the inverse conversion unit (322), the conversion coefficients can be inversely converted to obtain a residual signal (residual block, residual sample array).

[0090] The prediction unit (330) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit may apply intra prediction or inter prediction for a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called combined inter and intra prediction (CIIP). Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, for example, screen content coding (SCC). IBC basically performs prediction within the current picture, but it can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. The palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and palette index can be included in the above video / image information and signaled.

[0091] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The description of the intra prediction unit (222) may be applied equally to the intra prediction unit (331). The referenced samples may be located in the neighborhood of the current block or located away from it, depending on the prediction mode. In intra prediction, the prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit (331) may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.

[0092] The inter prediction unit (332) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) may construct a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-prediction can be performed based on various prediction modes (techniques), and information regarding the prediction may include information indicating the mode (technique) of inter-prediction for the current block.

[0093] The adder (340) can generate a restoration signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the prediction signal (predicted block, predicted sample array) output from the prediction unit (330) (including the inter prediction unit (332) and / or intra prediction unit (331)). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the restoration block. The description of the adder (250) can be applied equally to the adder (340). The adder (340) may be called a restoration unit or a restoration block generation unit. The generated restoration signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after undergoing filtering as described below.

[0094] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.

[0095] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can generate a modified restored picture by applying various filtering methods to the restored picture, and can store the modified restored picture in memory (360), specifically in the DPB of memory (360). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0096] The (modified) restored picture stored in the DPB of the memory (360) can be used as a reference picture in the inter-prediction unit (332). The memory (360) can store motion information of blocks from which motion information within the current picture has been derived (or decoded) and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (332) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (360) can store restoration samples of blocks restored within the current picture and transmit them to the intra-prediction unit (331).

[0097] In this specification, the embodiments described in the filtering unit (260), inter prediction unit (221), and intra prediction unit (222) of the encoding device (200) may be applied to the filtering unit (350), inter prediction unit (332), and intra prediction unit (331) of the decoding device (300) in the same or corresponding manner.

[0098] Figure 4 illustrates an exemplary hierarchical structure for a coded video / image.

[0099] Referring to Figure 4, the coded image is divided into a Video Coding Layer (VCL) that handles the decoding processing of the image and the image itself, a subsystem that transmits and stores the encoded information, and a Network Abstraction Layer (NAL) that exists between the VCL and the subsystem and is responsible for network adaptation functions.

[0100] In VCL, VCL data containing compressed image data (slice data) can be generated, or parameter sets containing information such as Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS), or SEI (Supplemental Enhancement Information) messages that are additionally required in the decoding process of the image can be generated.

[0101] In NAL, a NAL unit can be created by adding header information (NAL unit header) to the Raw Byte Sequence Payload (RBSP) generated in VCL. In this case, the RBSP refers to slice data, parameter sets, SEI messages, etc. generated in VCL. The NAL unit header may include NAL unit type information specified according to the RBSP data included in the NAL unit.

[0102] As shown in FIG. 4, NAL units can be classified into VCL NAL units and Non-VCL NAL units depending on the RBSP generated in VCL. A VCL NAL unit may refer to a NAL unit containing information about an image (slice data), and a Non-VCL NAL unit may refer to a NAL unit containing information necessary to decode an image (parameter set or SEI message).

[0103] The aforementioned VCL NAL unit and Non-VCL NAL unit can be transmitted over a network by attaching header information according to the data specifications of the underlying system. For example, the NAL unit can be transformed into a data format of a specified specification, such as H.266 / VVC file format, RTP (Real-time Transport Protocol), TS (Transport Stream), etc., and transmitted over various networks.

[0104] As described above, the NAL unit type can be specified according to the RBSP data structure included in the NAL unit, and information about this NAL unit type can be stored in the NAL unit header and signaled.

[0105] For example, NAL units can be broadly classified into VCL NAL unit types and Non-VCL NAL unit types depending on whether they contain information about the image (slice data). VCL NAL unit types can be classified according to the properties and types of the picture included in the VCL NAL unit, while Non-VCL NAL unit types can be classified according to the types of parameter sets.

[0106] The following is an example of a NAL unit type specified according to the type of parameter set included in the Non-VCL NAL unit type.

[0107] - APS (Adaptation Parameter Set) NAL unit: Type for the NAL unit containing the APS

[0108] - DPS(Decoding Parameter Set) NAL unit: Type for the NAL unit containing the DPS

[0109] - VPS (Video Parameter Set) NAL unit: Type for the NAL unit containing the VPS

[0110] - SPS (Sequence Parameter Set) NAL unit: Type for the NAL unit containing the SPS

[0111] - PPS(Picture Parameter Set) NAL unit: Type for the NAL unit containing the PPS

[0112] The above-described NAL unit types have syntax information for the NAL unit type, and said syntax information can be stored in the NAL unit header and signaled. For example, said syntax information may be nal_unit_type, and NAL unit types may be specified by the nal_unit_type value.

[0113] A slice header (slice header syntax, slice header information) may include information / parameters that can be commonly applied to the slice. The APS (APS syntax) or PPS (PPS syntax) may include information / parameters that can be commonly applied to one or more slices or pictures. The SPS (SPS syntax) may include information / parameters that can be commonly applied to one or more sequences. The VPS (VPS syntax) may include information / parameters that can be commonly applied to multiple layers. The DPS (DPS syntax) may include information / parameters that can be commonly applied to the entire video. The DPS may include information / parameters related to the concatenation of a CVS (coded video sequence). In the present disclosure, High Level Syntax (HLS) may include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, or slice header syntax.

[0114] In the present disclosure, image / video information encoded in an encoding device and signaled in the form of a bitstream includes not only information related to partitioning within a picture, intra / inter prediction information, residual information, in-loop filtering information, etc., but may also include information included in the slice header, information included in the APS, information included in the PPS, information included in the SPS, information included in the VPS, and / or information included in the DPS.

[0115] The following descriptor of the present disclosure specifies the parsing process for each syntax element.

[0116] - ae(v): Context-adaptive arithmetic entropy encoding syntax element.

[0117] - b(8): A byte (8 bits) with an arbitrary bit sequence pattern. The parsing process for this descriptor is specified by the return value of the function read_bits(8).

[0118] - f(n): An n-bit fixed pattern bit sequence written starting from the left bit (left to right). The parsing process for this descriptor is specified by the return value of the function read_bits(n).

[0119] - i(n): A signed integer using n bits. If n is "v" in the syntax table, the number of bits depends on the value of other syntax elements. The parsing process of this descriptor is specified by the return value of the read_bits(n) function, which is interpreted as a 2-complement integer representation with the most significant bit written first.

[0120] - se(v): A zero-order Exp-Golomb-coded signed integer syntax element starting from the left bit. The parsing process of this descriptor is specified as the case where order k is 0.

[0121] - st(v): A null-terminated string encoded in Universal Code Character Set (UCS) Transport Format-8 (UTF-8) characters. The syntax resolution process is specified as follows: st(v) reads a sequence of bytes from the bitstream starting at the byte alignment position within the bitstream, from the current position to the next byte alignment position (excluding bytes equal to 0x00), and returns the bitstream pointer by (stringLength + 1) * 8 bit positions, where stringLength is equal to the number of bytes returned.

[0122] For reference, the st(v) syntax descriptor is used in this disclosure only when the current position in the bitstream is a byte-aligned position.

[0123] - tu(v): A truncated one-way operator using up to maxVal bits, where maxVal is defined in the semantics of the syntax element.

[0124] - u(n): An unsigned integer using n bits. If n is "v" in the syntax table, the number of bits depends on the values ​​of other syntax elements. The parsing process of this descriptor is determined by interpreting the return value of the read_bits(n) function as a binary representation of an unsigned integer with the most significant bit written first.

[0125] - ue(v): An unsigned integer Exp-Golomb coded syntax element of order 0 with the left bit first. The parsing process for this descriptor is specified as the case where order k is 0.

[0126] The SEI message related to the present disclosure will be described below.

[0127] Table 1 shows an example of a Text description information SEI message syntax according to one embodiment.

[0128] [Table 1]

[0129]

[0130] An example of text description information SEI message semantics according to one embodiment is described.

[0131] Text description information SEI messages provide text descriptions for one or more pictures.

[0132] txt_descr_purpose indicates the purpose of the text description SEI as specified in Table 1. The value of text_descr_purpose must be in the range (inclusive) from 0 to 5. Values ​​(inclusive) of text_descr_purpose from 6 to 255 are reserved for future use and must not exist in bitstreams conforming to the present disclosure. A decoder conforming to the present disclosure must accept text_descr_purpose values ​​within the range (inclusive) from 0 to 255.

[0133] Table 2 shows an example of the definition of txt_descr_purpose.

[0134] [Table 2]

[0135]

[0136] If txt_cancel_flag is 1, it indicates that the text descriptive information SEI message cancels the persistence of the previous text descriptive information SEI message with the same txt_descr_purpose in the output order applied to the current layer. If txt_cancel_flag is 0, it indicates that text descriptive information follows.

[0137] txt_descr_id represents the identifier value of this text description information SEI message. The value of txt_descr_id must be in the range from 1 to 16383 (inclusive). The value 0 is reserved. Text description SEI messages with different txt_descr_purpose values ​​use separate value ranges for txt_descr_id.

[0138] If txt_id_cancel_flag is 1, the text descriptive information SEI message indicates that the persistence of the previous text descriptive information SEI message with the same txt_descr_id and txt_descr_purpose values ​​as the current SEI is canceled in the output order applied to the current layer. If txt_id_cancel_flag is 0, it indicates that the text descriptive information syntax elements (txt_persistence_flag, txt_num_strings_minus1, txt_descr_string_lang[ i ], txt_descr_string[ i ]) follow.

[0139] txt_persistence_flag specifies the persistence of text information description SEI messages for the current layer.

[0140] If txt_persistence_flag is 0, it specifies that text description information is applied only to the currently decoded picture.

[0141] If txt_persistence_flag is 1, it specifies that the text description information SEI message applies to the currently decoded picture and persists to all subsequent pictures of the current layer in output order until one or more of the following conditions become true:

[0142] - New CLVS of the current layer have started.

[0143] - When the bitstream ends.

[0144] - When the picture of the current layer within the AU associated with the text description information SEI message having the same txt_descr_id is output following the current picture in the output order.

[0145] The value obtained by adding 1 to txt_num_strings_minus1 represents the number of following txt_descr_string_lang[ i ] and txt_descr_string[ i ] items.

[0146] txt_descr_string_lang[ i ] specifies the language of txt_descr_string[ i ]. The language of txt_descr_string[ i ] must be indicated by language tags defined in IETF RFC 5646. For the st(v) syntax analysis process, the length of txt_descr_string_lang[ i ] corresponding to the previously specified stringLength variable must be between 0 and 49 bytes.

[0147] For all m and n in the range from 0 to txt_num_strings_minus1, when m is different from n, the value of txt_descr_string_lang[m] must not be equal to the value of txt_descr_string_lang[n].

[0148] txt_descr_string[ i ] specifies the i-th text description information string, and its value is interpreted as specified in txt_descr_purpose.

[0149] When txt_descr_purpose is 0, the interpretation of the information conveyed by txt_descr_string is defined by the application.

[0150] When txt_descr_purpose is 1, txt_descr_string[ i ] specifies copyright information associated with picture(s) within the persistence scope of this SEI message.

[0151] When txt_descr_purpose is 2, txt_descr_string[ i ] specifies AI marking information associated with picture(s) within the persistence scope of this SEI message (if not an empty string).

[0152] For reference, if txt_descr_purpose is 2, the string may contain information about machine learning-based processing, the intended use of the decoded picture, or other aspects related to the picture.

[0153] If txt_descr_purpose is 3, txt_descr_string[ i ] specifies a plain text label description associated with the picture(s) within the persistence scope of this SEI message.

[0154] If txt_descr_purpose is 4, txt_descr_string[ i ] must specify content rating information corresponding to the U.S. and Canada Rating Region Table (RRT), which is relevant to the video within the persistence range of this SEI message.

[0155] If txt_descr_purpose is 5, txt_descr_string[ i ] must contain a tag URI of syntax and semantics identifying CLVS.

[0156] The conventional design of a Text Description (TDI) SEI message includes a persistence signal (txt_persistence_flag) and a Text Description identifier (txt_descr_id). While txt_persistence_flag should describe the persistence of Text Description information with a specific txt_descr_id, conventional semantics specifies that txt_persistence_flag describes the persistence of all TDI SEI messages.

[0157] In addition, conventional semantics include constraints on the invalid value range of txt_descr_id.

[0158] This can cause confusion during implementation and incorrect operation of the decoder.

[0159] One embodiment may be applied individually or in combination.

[0160] 1. Modify the semantics of txt_persistence_flag to describe persistence for TDI SEI messages with a specific txt_descr_id, rather than for all TDI SEI messages.

[0161] 2. Modify the semantics of txt_persistence_flag to describe persistence for TDI SEI messages with the same txt_descr_id and txt_descr_purpose, rather than for all TDI SEI messages.

[0162] 3. Modify the constraints on the range values ​​of txt_descr_id.

[0163] An example of text description information SEI message semantics according to one embodiment is described. Semantics different from the text description information SEI message semantics described above are described, and descriptions of semantics identical to the text description information SEI message semantics described above are replaced with descriptions of the text description information SEI message semantics described above. Specifically, among the text description information SEI message semantics described above, semantics regarding txt_persistence_flag are described, and descriptions of other semantics are replaced with descriptions of the text description information SEI message semantics described above.

[0164] txt_persistence_flag specifies the persistence of the text information description SEI message with the current txt_descr_id in the current layer.

[0165] If txt_persistence_flag is 0, it indicates that the text description information SEI message with the current txt_descr_id is applied only to the currently decoded picture.

[0166] If txt_persistence_flag is 1, the text description information SEI message with the current txt_descr_id is applied to the currently decoded picture and persists to all subsequent pictures of the current layer in output order until one or more of the following conditions become true:

[0167] - Until new CLVS of the current layer start.

[0168] - Until the bitstream ends.

[0169] - Until the picture of the current layer within the AU associated with the text description information SEI message having the same txt_descr_id is output following the current picture in output order.

[0170] An example of text description information SEI message semantics according to one embodiment is described. Semantics different from the text description information SEI message semantics described above are described, and descriptions of semantics identical to the text description information SEI message semantics described above are replaced with descriptions of the text description information SEI message semantics described above. Specifically, among the text description information SEI message semantics described above, semantics regarding txt_persistence_flag are described, and descriptions of other semantics are replaced with descriptions of the text description information SEI message semantics described above.

[0171] txt_persistence_flag specifies the persistence of text information description SEI messages with the current txt_descr_id and txt_descr_purpose for the current layer.

[0172] If txt_persistence_flag is 0, it specifies that the text description information SEI message with the current txt_descr_id and txt_descr_purpose applies only to the currently decoded picture.

[0173] If txt_persistence_flag is 1, the text description information SEI message with the current txt_descr_id and txt_descr_purpose applies to the currently decoded picture and persists to all subsequent pictures of the current layer in output order until one or more of the following conditions become true:

[0174] - Until new CLVS of the current layer start.

[0175] - Until the bitstream ends.

[0176] - Until the picture of the current layer within the AU associated with the text description information SEI message having the same txt_descr_id and txt_descr_purpose is output following the current picture in output order.

[0177] An example of text description information SEI message semantics according to one embodiment is described. Semantics different from the text description information SEI message semantics described above are described, and descriptions of semantics identical to the text description information SEI message semantics described above are replaced with descriptions of the text description information SEI message semantics described above. Specifically, among the text description information SEI message semantics described above, semantics regarding txt_descr_id are described, and descriptions of other semantics are replaced with descriptions of the text description information SEI message semantics described above.

[0178] txt_descr_id represents the identifier value of a text description information SEI message. The value of txt_descr_id must be in the range from 1 to 8,191 (inclusive). The value 0 is reserved. Text description SEI messages with different txt_descr_purpose values ​​use separate value ranges for txt_descr_id.

[0179] The terms or names described below (e.g., names of syntax elements or variables, etc.) are merely examples, and the technical features of the present disclosure are not limited to the terms, etc. described below. For example, the image information described below may include various information according to the embodiments described in the present disclosure and may include information described in at least one of the tables described above.

[0180] The operations described below do not constitute an essential component of one embodiment, and at least some of the operations described below may be omitted. Furthermore, the operations described below do not constitute a sufficient component of one embodiment, and previously described operations may be added. Moreover, unless they contradict previously described operations, the operations described below form one embodiment integrally with previously described operations and do not form a separate embodiment distinct from previously described operations.

[0181] FIG. 5 is a diagram illustrating a method for decoding image information according to one embodiment of the present disclosure.

[0182] Terms or names (e.g., names of syntax elements or names of variables, etc.) are merely examples, and the technical features of the present disclosure are not limited to the terms, etc. described in FIG. 5. For example, the image information described in FIG. 5 may include various information according to the embodiments described in the present disclosure and may include information described in at least one of the tables described above.

[0183] The decoding method (S500) may include operations described below. The operations described below do not constitute an essential component of the decoding method according to one embodiment, and at least some of the operations described below may be omitted. Furthermore, the operations described below do not constitute a sufficient component of the decoding method according to one embodiment, and the previously described operations may be added. Moreover, unless the operations described below contradict the previously described operations, they form an embodiment integrally with the previously described operations and do not form a separate embodiment distinguished from the previously described operations.

[0184] The decoding method (S500) can be executed by a decoding device including a memory and a processor electrically connected to the memory, for example, by a processor.

[0185] The decoding device can acquire image information (S510).

[0186] For example, the processor of a decoding device can acquire image information.

[0187] Image information may include SEI (supplemental enhancement information) messages. SEI messages may convey specific types of information that assist in processes related to the decoding, display, or other purposes of image information. Here, SEI messages may not be necessary for the decoding process to determine the sample values ​​of the decoded picture.

[0188] For example, the SEI message may include a text description information (TDI) SEI message that provides text description information (TDI) for at least one picture.

[0189] TDI SEI messages can take various forms. For example, a TDI SEI message may be a syntax element or a syntax structure containing one or more syntax elements. Additionally, a TDI SEI message may be a raw byte sequence payload (RBSP) containing one or more syntax elements or one or more syntax structures. For example, a TDI SEI message may be represented as text_description_info(payloadSize), but is not limited thereto.

[0190] Furthermore, the TDI SEI message may have various names such as TDI message, TDI related message, TDI related information, TXT SEI message, TXT message, TXT related message, TXT related information, text description information message, etc., and the names are not limited.

[0191] The text description information provided by the TDI SEI message may include basic information related to the text description, such as information regarding the number of text descriptions, information regarding the language of the text descriptions, and text descriptions.

[0192] Information regarding the number of text descriptions can indicate the number of text descriptions and information regarding their language provided by the TDI SEI message.

[0193] Information regarding the number of text descriptions may take various forms and may be expressed by various names. For example, information regarding the number of text descriptions may be a syntax element or a syntax structure containing one or more syntax elements. For example, information regarding the number of text descriptions that is a syntax element may include an integer composed of multiple bits (e.g., 8 bits). Information regarding the number of text descriptions that is a syntax element may be expressed as txt_num_strings_minus1 or tdi_num_strings_minus1, but is not limited thereto.

[0194] Information regarding the language of the text description is provided in the form of a string and can represent the language of the text description.

[0195] Information regarding the language of a text description may take various forms and may be expressed by various names. For example, information regarding the language of a text description may be a syntax element or a syntax structure containing one or more syntax elements. For example, information regarding the language of a text description that is a syntax element may be in the form of a string. Information regarding the language of a text description that is a syntax element may be expressed as txt_descr_string_lang[ i ] or tdi_descr_string_lang[ i ], but is not limited thereto.

[0196] Text descriptions are provided in the form of strings and may include various texts that are interpreted according to the purpose of the text description.

[0197] Text descriptions can take various forms and may be represented by various names. For example, a text description may be a syntax element or a syntax structure containing one or more syntax elements. For example, a text description that is a syntax element may be in the form of a string. A text description that is a syntax element may be represented as txt_descr_string[ i ] or tdi_descr_string[ i ], but is not limited thereto.

[0198] Text description information may include additional information related to the text description information in addition to basic information regarding the text description. For example, the additional information may include information regarding the purpose of the text description, cancellation information of the TDI SEI message, identifier information of the text description, cancellation information of the text description, persistence information of the text description, etc.

[0199] Information regarding the purpose of the text description may indicate the purpose of the text description. The purpose of the text description may include the definition of the application, copyright information, AI marking information, general comment information, content rating information, tag URIs for identifying bitstreams, etc.

[0200] Information regarding the purpose of a text description may take various forms and may be expressed by various names. For example, information regarding the purpose of a text description may be a syntax element or a syntax structure containing one or more syntax elements. For example, information regarding the purpose of a text description that is a syntax element may include an integer composed of multiple bits (e.g., 8 bits). Information regarding the purpose of a text description that is a syntax element may be expressed as txt_descr_purpose or tdi_descr_purpose, but is not limited thereto.

[0201] Cancellation information of a TDI SEI message may indicate whether the persistence of all previous TDI SEI messages is canceled. For example, cancellation information of a TDI SEI message with a value of 1 may indicate that the persistence of all previous TDI SEI messages is canceled. Cancellation information of a TDI SEI message with a value of 0 may indicate that the current TDI SEI message contains text description information. However, this is not limited thereto, and alternatively, what is indicated by a cancellation information value of 1 may be changed from what is indicated by a cancellation information value of 0.

[0202] Cancellation information of a TDI SEI message may take various forms and may be represented by various names. For example, cancellation information of a TDI SEI message may be a syntax element or a syntax structure containing one or more syntax elements. For example, cancellation information of a TDI SEI message that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits or a variable length. Cancellation information of a TDI SEI message that is a syntax element may be represented as txt_cancel_flag, tdi_cancel_flag, tdi_purpose_cancel_flag, etc., but is not limited thereto.

[0203] The identifier information of a text description may include an identifier for identifying the text description. The identifier information of a text description may have a value in the range of 1 to 8,191 (inclusive).

[0204] Identifier information of a text description may take various forms and may be represented by various names. For example, identifier information of a text description may be a syntax element or a syntax structure containing one or more syntax elements. For example, identifier information of a text description that is a syntax element may include an integer composed of multiple bits (e.g., 13 bits). Identifier information of a text description that is a syntax element may be represented as txt_descr_id or tdi_descr_id, but is not limited thereto.

[0205] Cancellation information for a text description may indicate whether the persistence of a previous TDI SEI message is canceled, having the same purpose as the purpose indicated by the information regarding the purpose of the text description included in the current TDI SEI message and / or the same identifier as the identifier indicated by the identifier information included in the current TDI SEI message. For example, cancellation information for a text description with a value of 1 may indicate that the persistence of a previous TDI SEI message is canceled, having the same purpose as the purpose indicated by the information regarding the purpose of the text description included in the current TDI SEI message and / or the same identifier as the identifier indicated by the identifier information included in the current TDI SEI message. Cancellation information for a text description with a value of 0 may indicate that the current TDI SEI message contains text description information. However, this is not limited thereto, and alternatively, what is indicated by a cancellation information value of 1 may be changed from what is indicated by a cancellation information value of 0.

[0206] Thus, by limiting the TDI SEI message whose persistence is canceled by the cancellation information of the text description to a previous TDI SEI message having the same purpose as the purpose indicated by the information regarding the purpose of the text description included in the current TDI SEI message and / or the same identifier as the identifier indicated by the identifier information included in the current TDI SEI message, the target TDI SEI message whose persistence is canceled becomes clear, and confusion or malfunction in the decoding device caused by the ambiguity of the target TDI SEI message whose persistence is canceled can be suppressed, prevented, or minimized. Furthermore, the TDI SEI message whose persistence is canceled by the cancellation information of the TDI SEI message and the TDI SEI message whose persistence is canceled by the cancellation information of the text description can be clearly distinguished.

[0207] Cancellation information of a text description may take various forms and may be expressed by various names. For example, cancellation information of a text description may be a syntax element or a syntax structure containing one or more syntax elements. For example, cancellation information of a text description that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more or variable-length bits. Cancellation information of a text description that is a syntax element may be expressed as txt_id_cancel_flag or tdi_id_cancel_flag, but is not limited thereto.

[0208] The persistence information of a text description may indicate whether a TDI SEI message is persistent, having the same purpose as the purpose indicated by the information regarding the purpose of the text description included in the TDI SEI message and / or the same identifier as the identifier indicated by the identifier information included in the TDI SEI message. For example, the persistence information of a text description with a value of 0 may indicate that a TDI SEI message is applicable only up to the current picture, having the same purpose as the purpose indicated by the information regarding the purpose of the text description included in the current TDI SEI message and / or the same identifier as the identifier indicated by the identifier information included in the current TDI SEI message. The persistence information of a text description with a value of 1 may indicate that a TDI SEI message having the same purpose as the purpose indicated by the information regarding the purpose of the text description included in the current TDI SEI message and / or the same identifier as the identifier indicated by the identifier information included in the current TDI SEI message is persisted until i) a new CLVS is started, ii) a bitstream is terminated, or iii) a picture associated with a new TDI SEI message having the same purpose as the purpose indicated by the information regarding the purpose of the text description included in the current TDI SEI message and / or the same identifier as the identifier indicated by the identifier information included in the current TDI SEI message is output. However, this is not limited thereto, and alternatively, what is indicated by the persistence information having a value of 1 may be changed to what is indicated by the persistence information having a value of 0.

[0209] In this way, the TDI SEI message whose persistence is maintained or canceled by the persistence information of the text description is limited to a TDI SEI message having the same purpose as the purpose indicated by the information regarding the purpose of the text description included in the current TDI SEI message and / or the same identifier as the identifier indicated by the identifier information included in the current TDI SEI message, thereby making the target TDI SEI message and / or target text description whose persistence is maintained or canceled clear, and can suppress, prevent, or minimize confusion or malfunction in the decoding device caused by the ambiguity of the target TDI SEI message and / or target text description whose persistence is maintained or canceled.

[0210] Persistence information of a text description may take various forms and may be expressed by various names. For example, persistence information of a text description may be a syntax element or a syntax structure containing one or more syntax elements. For example, persistence information of a text description that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits or a variable length. Persistence information of a text description that is a syntax element may be expressed as txt_persistence_flag or tdi_persistence_flag, but is not limited thereto.

[0211] The decoding device can derive text description information (S520).

[0212] For example, a processor of a decoding device can process a TDI SEI message. The decoding device can derive text description information based on processing the TDI SEI message. Specifically, the decoding device can obtain a text description and its persistence based on processing the TDI SEI message.

[0213] In this way, the decoding device can obtain a text description and its persistence based on TDI SEI messages. For example, the decoding device can obtain information regarding a text description and its persistence based on information regarding the purpose of the text description, cancellation information of the TDI SEI message, identifier information of the text description, cancellation information of the text description, and persistence information of the text description.

[0214] The decoding device can obtain information regarding the text description and its language based on cancellation information of the TDI SEI message and / or cancellation information of the text description. For example, the decoding device can obtain information regarding the text description and its language based on the fact that the cancellation information of the TDI SEI message does not indicate cancellation of persistence and the cancellation information of the text description does not indicate cancellation of persistence.

[0215] Additionally, the decoding device can obtain a text description regarding the current picture based on text description information. For example, the decoding device can obtain a text description written in the language indicated by the information regarding the language of the text description included in the text description information.

[0216] The decoding device can obtain information regarding the persistence of text description information based on information regarding the purpose of the text description, identifier information of the text description, and / or persistence information of the text description. Based on the persistence information of the text description, the decoding device can determine whether a TDI SEI message has persistence that has the same purpose as indicated by the information regarding the purpose of the text description included in the current TDI SEI message and / or the same identifier as indicated by the identifier information of the text description included in the current TDI SEI message.

[0217] As described above, since the text description information whose persistence is maintained or canceled by the persistence information of the text description is limited to a TDI SEI message having the same purpose as the purpose indicated by the information regarding the purpose of the text description included in the current TDI SEI message and / or the same identifier as the identifier indicated by the identifier information included in the current TDI SEI message, the target TDI SEI message and / or text description whose persistence is maintained or canceled becomes clear, and confusion or malfunction in the decoding device due to the ambiguity of the target TDI SEI message and / or target text description whose persistence is maintained or canceled can be suppressed, prevented, or minimized.

[0218] By doing so, the reliability of a coding system including an encoding device and a decoding device can be further improved.

[0219] By doing so, the coding efficiency of a coding system including an encoding device and a decoding device can be further improved.

[0220] By doing so, the data transmission efficiency of a coding system including an encoding device and a decoding device can be further improved.

[0221] FIG. 6 is a diagram illustrating a method for encoding image information according to one embodiment of the present disclosure.

[0222] The terms or names described in FIG. 6 (e.g., names of syntax elements or variables, etc.) are merely examples, and the technical features of the present disclosure are not limited to the terms, etc. described in FIG. 6. For example, the image information described in FIG. 6 may include various information according to the embodiments described in the present disclosure and may include information described in at least one of the tables described above.

[0223] The encoding method (S600) may include operations described below. The operations described below do not constitute an essential component of the decoding method according to one embodiment, and at least some of the operations described below may be omitted. Furthermore, the operations described below do not constitute a sufficient component of the encoding method according to one embodiment, and the previously described operations may be added. Moreover, unless the operations described below contradict the previously described operations, they form an embodiment integrally with the previously described operations and do not form a separate embodiment distinct from the previously described operations.

[0224] The encoding method (S600) can be executed by an encoding device including a memory and a processor electrically connected to the memory, for example, by a processor.

[0225] The encoding device can generate text description information (S610).

[0226] For example, the processor of the encoding device may obtain information regarding at least one text description and a picture to which the text description applies through various paths and / or various means. Additionally, the processor of the encoding device may generate text description information based on the obtained information regarding at least one text description and a picture to which the text description applies.

[0227] Text description information may include basic information related to text descriptions, such as information regarding the number of at least one text description, information regarding the language of each of at least one text description, and at least one text description.

[0228] Information regarding the number of text descriptions can indicate the number of at least one text description and information regarding that language.

[0229] Information regarding the number of text descriptions may take various forms and may be expressed by various names. For example, information regarding the number of text descriptions may be a syntax element or a syntax structure containing one or more syntax elements. For example, information regarding the number of text descriptions that is a syntax element may include an integer composed of multiple bits (e.g., 8 bits). Information regarding the number of text descriptions that is a syntax element may be expressed as txt_num_strings_minus1 or tdi_num_strings_minus1, but is not limited thereto.

[0230] Information regarding the language of the text description is provided in the form of a string and can represent the language of the text description.

[0231] Information regarding the language of a text description may take various forms and may be expressed by various names. For example, information regarding the language of a text description may be a syntax element or a syntax structure containing one or more syntax elements. For example, information regarding the language of a text description that is a syntax element may be in the form of a string. Information regarding the language of a text description that is a syntax element may be expressed as txt_descr_string_lang[ i ] or tdi_descr_string_lang[ i ], but is not limited thereto.

[0232] Text descriptions are provided in the form of strings and may include various texts that are interpreted according to the purpose of the text description.

[0233] Text descriptions can take various forms and may be represented by various names. For example, a text description may be a syntax element or a syntax structure containing one or more syntax elements. For example, a text description that is a syntax element may be in the form of a string. A text description that is a syntax element may be represented as txt_descr_string[ i ] or tdi_descr_string[ i ], but is not limited thereto.

[0234] The encoding device can encode video information (S620).

[0235] For example, the processor of the encoding device can generate text description information (TDI) SEI (supplemental enhancement information) containing text description information, generate image information containing TDI SEI messages, and encode the image information.

[0236] The encoding device can generate TDI SEI messages. Here, the TDI SEI message may have various names such as TDI message, TDI related message, TDI related information, TXT SEI message, TXT message, TXT related message, TXT related information, text description information message, etc., and the names are not limited.

[0237] Additionally, the TDI SEI message may take various forms. For example, the TDI SEI message may be a syntax element or a syntax structure containing one or more syntax elements. Additionally, the TDI SEI message may be a raw byte sequence payload (RBSP) containing one or more syntax elements or one or more syntax structures. For example, the TDI SEI message may be represented as text_description_info(payloadSize), but is not limited thereto.

[0238] Text description information may include additional information related to the text description information in addition to basic information regarding the text description. For example, the text description information may include information regarding the purpose of the text description, cancellation information of the TDI SEI message, identifier information of the text description, cancellation information of the text description, persistence information of the text description, etc.

[0239] Information regarding the purpose of the text description may indicate the purpose of the text description. The purpose of the text description may include the definition of the application, copyright information, AI marking information, general comment information, content rating information, tag URIs for identifying bitstreams, etc.

[0240] Information regarding the purpose of a text description may take various forms and may be expressed by various names. For example, information regarding the purpose of a text description may be a syntax element or a syntax structure containing one or more syntax elements. For example, information regarding the purpose of a text description that is a syntax element may include an integer composed of multiple bits (e.g., 8 bits). Information regarding the purpose of a text description that is a syntax element may be expressed as txt_descr_purpose or tdi_descr_purpose, but is not limited thereto.

[0241] Cancellation information of a TDI SEI message may indicate whether the persistence of all previous TDI SEI messages is canceled. For example, cancellation information of a TDI SEI message with a value of 1 may indicate that the persistence of all previous TDI SEI messages is canceled. Cancellation information of a TDI SEI message with a value of 0 may indicate that the current TDI SEI message contains text description information. However, this is not limited thereto, and alternatively, what is indicated by a cancellation information value of 1 may be changed from what is indicated by a cancellation information value of 0.

[0242] Cancellation information of a TDI SEI message may take various forms and may be represented by various names. For example, cancellation information of a TDI SEI message may be a syntax element or a syntax structure containing one or more syntax elements. For example, cancellation information of a TDI SEI message that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits or a variable length. Cancellation information of a TDI SEI message that is a syntax element may be represented as txt_cancel_flag, tdi_cancel_flag, tdi_purpose_cancel_flag, etc., but is not limited thereto.

[0243] The identifier information of a text description may include an identifier for identifying the text description. The identifier information of a text description may have a value in the range of 1 to 8,191 (inclusive).

[0244] Identifier information of a text description may take various forms and may be represented by various names. For example, identifier information of a text description may be a syntax element or a syntax structure containing one or more syntax elements. For example, identifier information of a text description that is a syntax element may include an integer composed of multiple bits (e.g., 13 bits). Identifier information of a text description that is a syntax element may be represented as txt_descr_id or tdi_descr_id, but is not limited thereto.

[0245] Cancellation information for a text description may indicate whether the persistence of a previous TDI SEI message is canceled, having the same purpose as the purpose indicated by the information regarding the purpose of the text description included in the current TDI SEI message and / or the same identifier as the identifier indicated by the identifier information included in the current TDI SEI message. For example, cancellation information for a text description with a value of 1 may indicate that the persistence of a previous TDI SEI message is canceled, having the same purpose as the purpose indicated by the information regarding the purpose of the text description included in the current TDI SEI message and / or the same identifier as the identifier indicated by the identifier information included in the current TDI SEI message. Cancellation information for a text description with a value of 0 may indicate that the current TDI SEI message contains text description information. However, this is not limited thereto, and alternatively, what is indicated by a cancellation information value of 1 may be changed from what is indicated by a cancellation information value of 0.

[0246] Thus, by limiting the TDI SEI message whose persistence is canceled by the cancellation information of the text description to a previous TDI SEI message having the same purpose as the purpose indicated by the information regarding the purpose of the text description included in the current TDI SEI message and / or the same identifier as the identifier indicated by the identifier information included in the current TDI SEI message, the target TDI SEI message whose persistence is canceled becomes clear, and confusion or malfunction in the decoding device caused by the ambiguity of the target TDI SEI message whose persistence is canceled can be suppressed, prevented, or minimized. Furthermore, the TDI SEI message whose persistence is canceled by the cancellation information of the TDI SEI message and the TDI SEI message whose persistence is canceled by the cancellation information of the text description can be clearly distinguished.

[0247] Cancellation information of a text description may take various forms and may be expressed by various names. For example, cancellation information of a text description may be a syntax element or a syntax structure containing one or more syntax elements. For example, cancellation information of a text description that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more or variable-length bits. Cancellation information of a text description that is a syntax element may be expressed as txt_id_cancel_flag or tdi_id_cancel_flag, but is not limited thereto.

[0248] The persistence information of a text description may indicate whether a TDI SEI message is persistent, having the same purpose as the purpose indicated by the information regarding the purpose of the text description included in the TDI SEI message and / or the same identifier as the identifier indicated by the identifier information included in the TDI SEI message. For example, the persistence information of a text description with a value of 0 may indicate that a TDI SEI message is applicable only up to the current picture, having the same purpose as the purpose indicated by the information regarding the purpose of the text description included in the current TDI SEI message and / or the same identifier as the identifier indicated by the identifier information included in the current TDI SEI message. The persistence information of a text description with a value of 1 may indicate that a TDI SEI message having the same purpose as the purpose indicated by the information regarding the purpose of the text description included in the current TDI SEI message and / or the same identifier as the identifier indicated by the identifier information included in the current TDI SEI message is persisted until i) a new CLVS is started, ii) a bitstream is terminated, or iii) a picture associated with a new TDI SEI message having the same purpose as the purpose indicated by the information regarding the purpose of the text description included in the current TDI SEI message and / or the same identifier as the identifier indicated by the identifier information included in the current TDI SEI message is output. However, this is not limited thereto, and alternatively, what is indicated by the persistence information having a value of 1 may be changed to what is indicated by the persistence information having a value of 0.

[0249] In this way, the TDI SEI message whose persistence is maintained or canceled by the persistence information of the text description is limited to a TDI SEI message having the same purpose as the purpose indicated by the information regarding the purpose of the text description included in the current TDI SEI message and / or the same identifier as the identifier indicated by the identifier information included in the current TDI SEI message, thereby making the target TDI SEI message and / or target text description whose persistence is maintained or canceled clear, and can suppress, prevent, or minimize confusion or malfunction in the decoding device caused by the ambiguity of the target TDI SEI message and / or target text description whose persistence is maintained or canceled.

[0250] Persistence information of a text description may take various forms and may be expressed by various names. For example, persistence information of a text description may be a syntax element or a syntax structure containing one or more syntax elements. For example, persistence information of a text description that is a syntax element may include a flag consisting of one bit or an indicator consisting of two or more bits or a variable length. Persistence information of a text description that is a syntax element may be expressed as txt_persistence_flag or tdi_persistence_flag, but is not limited thereto.

[0251] The encoding device can generate image information based on TDI SEI messages.

[0252] Image information may include various SEI messages, including TDI SEI messages. SEI messages may convey specific types of information that assist in processes related to the decoding, display, or other purposes of image information. Here, SEI messages may not be necessary for the decoding process to determine the sample values ​​of the decoded picture.

[0253] The encoding device can encode video information containing various SEI messages.

[0254] In this way, the encoding device can generate a TDI SEI message based on a text description and its persistence. For example, the encoding device can generate information regarding the purpose of the text description, cancellation information of the TDI SEI message, identifier information of the text description, cancellation information of the text description, and persistence information of the text description based on information regarding the text description and its persistence.

[0255] The encoding device can generate text description information based on a text description regarding the current picture. For example, the encoding device can generate text description information such that a text description written in the language indicated by the information regarding the language of the text description included in the text description information is derived.

[0256] Additionally, the encoding device can generate cancellation information for the TDI SEI message and / or cancellation information for the text description to signal information regarding the text description and its language depending on the persistence of the text description.

[0257] The encoding device can generate information regarding the purpose of a text description, identifier information of a text description, and / or persistence information of a text description based on the persistence of a text description. The encoding device can generate persistence information of a text description based on whether a TDI SEI message has persistence that has the same purpose as indicated by the information regarding the purpose of the text description included in the current TDI SEI message and / or the same identifier as indicated by the identifier information of the text description included in the current TDI SEI message.

[0258] As described above, since the text description information whose persistence is maintained or canceled by the persistence information of the text description is limited to a TDI SEI message having the same purpose as the purpose indicated by the information regarding the purpose of the text description included in the current TDI SEI message and / or the same identifier as the identifier indicated by the identifier information included in the current TDI SEI message, the target TDI SEI message and / or target text description whose persistence is maintained or canceled becomes clear, and confusion or malfunction in the decoding device due to the ambiguity of the target TDI SEI message and / or target text description whose persistence is maintained or canceled can be suppressed, prevented, or minimized.

[0259] By doing so, the reliability of a coding system including an encoding device and a decoding device can be further improved.

[0260] By doing so, the coding efficiency of a coding system including an encoding device and a decoding device can be further improved.

[0261] By doing so, the data transmission efficiency of a coding system including an encoding device and a decoding device can be further improved.

[0262] A bitstream is generated based on video information encoded according to the encoding method (S600) described above, and the bitstream can be stored on a computer-readable storage medium.

[0263] In addition, a bitstream is generated based on video information encoded according to the encoding method (S600) described above, and the bitstream can be transmitted through a transmission unit and / or a transmission medium.

[0264] FIG. 7 is a drawing illustrating an exemplary content streaming system to which an embodiment according to the present disclosure can be applied.

[0265] As illustrated in FIG. 7, a content streaming system to which an embodiment of the present disclosure is applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0266] The above encoding server compresses content input from multimedia input devices, such as smartphones, cameras, and camcorders, into digital data to generate a bitstream and transmits it to the streaming server. As another example, if multimedia input devices, such as smartphones, cameras, and camcorders, generate the bitstream directly, the encoding server may be omitted.

[0267] The bitstream may be generated by a video encoding method and / or encoding device to which an embodiment of the present disclosure is applied, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0268] The streaming server transmits multimedia data to a user device based on a user request through a web server, and the web server can act as a medium to inform the user of available services. When a user requests a desired service from the web server, the web server transmits it to the streaming server, and the streaming server can transmit multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server can perform the role of controlling commands and responses between each device within the content streaming system.

[0269] The streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, to provide a seamless streaming service, the streaming server can store the bitstream for a certain period of time.

[0270] Examples of the above user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs (head-mounted displays)), digital TVs, desktop computers, digital signage, etc.

[0271] Each server within the above-mentioned content streaming system can be operated as a distributed server, and in this case, data received from each server can be processed in a distributed manner.

[0272] The scope of the present disclosure includes software or machine-executable instructions (e.g., operating system, application, firmware, program, etc.) that enable an operation according to a method of various embodiments to be executed on a device or computer, and a non-transitory computer-readable medium on which such software or instructions, etc. are stored and executable on a device or computer.

[0273] An embodiment according to the present disclosure can be used to encode / decode images.

Claims

1. Regarding the method of decoding video information, Acquire the image information including text description information (TDI) SEI (supplemental enhancement information) messages; The method includes deriving text description information for at least one picture based on the above text description information SEI message, wherein The above text description information SEI message includes description purpose information indicating the purpose of the above text description information SEI message, identifier information identifying the above text description information SEI message, and persistence information indicating the persistence of the above text description information SEI message. The persistence of the above text description information SEI message is a method derived based on the above description purpose information and / or the above identifier information.

2. In Paragraph 1, A method for indicating the persistence of an SEI message, wherein the above persistence information includes the same purpose as the above descriptive purpose information and / or the same identifier as the above identifier information.

3. In Paragraph 1, A method in which an SEI message containing the same purpose as the above-mentioned descriptive purpose information and / or the same identifier as the above-mentioned identifier information is applied only to the current picture, based on the above-mentioned persistence information having a value such as 0.

4. In Paragraph 1, A method in which an SEI message containing the same purpose as the descriptive purpose information and / or the same identifier as the identifier information is applied to a current picture and at least one picture following the current picture, based on the persistence information having a value equal to 1.

5. In Paragraph 1, A method in which, based on the persistence information having a value equal to 1, an SEI message containing the same purpose as the descriptive purpose information and / or the same identifier as the identifier information is applied to at least one picture following the current picture until the picture associated with the SEI message containing the same purpose as the descriptive purpose information and / or the same identifier as the identifier information is output.

6. In Paragraph 1, The above identifier information has a value from 1 to 8,191.

7. Regarding the method of encoding video information, Generate text description information for at least one picture; The method includes encoding the image information including a text description information (TDI) SEI (supplemental enhancement information) message generated based on the text description information for at least one picture, and The above text description information SEI message includes description purpose information indicating the purpose of the above text description information SEI message, identifier information identifying the above text description information SEI message, and persistence information indicating the persistence of the above text description information SEI message. A method for determining the persistence of the text description information SEI message based on the above description purpose information and / or the above identifier information.

8. In Paragraph 7, A method for generating persistence information based on the persistence of an SEI message containing the same purpose as the above-described purpose information and / or the same identifier as the above-described identifier information.

9. In Paragraph 7, A method in which the persistence information has a value such as 0 based on the fact that an SEI message containing the same purpose as the above-described purpose information and / or the same identifier as the above-described identifier information applies only to the current picture.

10. In Paragraph 7 A method in which the persistence information has a value such as 1, based on the application of an SEI message containing the same purpose as the above-described purpose information and / or the same identifier as the above-described identifier information to the current picture and at least one picture following the current picture.

11. In Paragraph 7, A method in which the persistence information has a value such as 1, based on the fact that an SEI message containing the same purpose as the above-described purpose information and / or the same identifier as the above-described identifier information is applied to at least one picture following the current picture until the current picture and the picture associated with the SEI message containing the same purpose as the above-described purpose information and / or the same identifier as the above-described identifier information are output.

12. In Paragraph 7, The above identifier information has a value from 1 to 8,191.

13. Regarding methods concerning bitstreams, Generate text description information for at least one picture; Generating a bitstream based on image information including a text description information (TDI) SEI (supplemental enhancement information) message generated based on the text description information for at least one picture; It includes transmitting data regarding the above bitstream, and The above text description information SEI message includes description purpose information indicating the purpose of the above text description information SEI message, identifier information identifying the above text description information SEI message, and persistence information indicating the persistence of the above text description information SEI message. A method for determining the persistence of the text description information SEI message based on the above description purpose information and / or the above identifier information.

14. In Paragraph 13, A method for generating persistence information based on the persistence of an SEI message containing the same purpose as the above-described purpose information and / or the same identifier as the above-described identifier information.

15. In a computer-readable storage medium for storing a bitstream, The above storage medium stores the bitstream generated based on image information including a text description information (TDI) SEI (supplemental enhancement information) message generated based on the text description information for at least one picture, wherein The above text description information SEI message includes description purpose information indicating the purpose of the above text description information SEI message, identifier information identifying the above text description information SEI message, and persistence information indicating the persistence of the above text description information SEI message. A storage medium in which the persistence of the text description information SEI message is determined based on the above description purpose information and / or the above identifier information.

16. In Paragraph 15, A storage medium in which the persistence information is generated based on the persistence of an SEI message containing the same purpose as the above-described purpose information and / or the same identifier as the above-described identifier information.