Method for decoding image information, method for encoding image information, method for bitstream, and computer-readable storage medium
Patent Information
- Application Number
- PCT/KR2026/004563
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-28
- Filing Date
- 2026-03-23
- Publication Date
- 2026-10-01
Smart Images

Figure KR2026004563_01102026_PF_FP_ABST
Abstract
Description
Method for decoding image information, method for encoding image information, method relating to a bitstream and a computer-readable storage medium
[0001] The present disclosure relates to a method for decoding image information, a method for encoding image information, a method for bitstreams, and a computer-readable storage medium.
[0002] Recently, the demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition), has been increasing across various fields. As video data becomes higher in resolution and quality, the relative amount of information or bits transmitted increases compared to conventional video data. This increase in the amount of transmitted information or bits leads to higher transmission and storage costs.
[0003] Accordingly, high-efficiency video compression technology is required to effectively transmit, store, and play back high-resolution, high-quality video information.
[0004] The present disclosure aims to improve the reliability of a coding system including an encoding device and a decoding device.
[0005] The present disclosure aims to improve the coding efficiency of a coding system including an encoding device and a decoding device.
[0006] The present disclosure aims to improve the data transmission efficiency of a coding system including an encoding device and a decoding device.
[0007] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure belongs from the description below.
[0008] A method for decoding image information according to one aspect of the present disclosure comprises: obtaining at least one SEI message including an error recovery SEI (supplemental enhancement information) message from a bitstream; and, based on the error recovery SEI message, obtaining a range of POC (picture order count) values for at least one temporal sublayer in which no active entry exists in a reference picture list for a current picture or a sub-picture of the current picture.
[0009] According to one aspect of the present disclosure, an apparatus for decoding image information comprises a memory and at least one processor connected to the memory, wherein the at least one processor obtains at least one SEI message including an error recovery SEI (supplemental enhancement information) message from a bitstream; and, based on the error recovery SEI message, obtains a range of POC (picture order count) values for at least one temporal sublayer in which no active entry exists in a reference picture list for a current picture or a sub-picture of the current picture.
[0010] In a method or device for decoding the above image information, the last picture having a specific temporal identifier used as a reference picture for the current picture or a sub-picture of the current picture can be derived based on the POC information of the error recovery SEI message corresponding to the specific temporal identifier.
[0011] In a method or device for decoding the above image information, the error recovery SEI message may include information regarding the number of at least one temporal sublayer.
[0012] In a method or apparatus for decoding the above image information, the error recovery SEI message may include POC information for deriving the last picture having the specific temporal identifier used as a reference picture for the current picture or the sub-picture of the current picture for each of the at least one temporal sublayer.
[0013] In a method or device for decoding the above image information, the current picture having the specific temporal identifier or the sub-picture of the current picture may not have an active entry of a reference picture list having the specific temporal identifier and a POC value within the range of the POC value.
[0014] In a method or device for decoding the above image information, the range of the POC value can be derived based on the difference between the POC value based on the POC information and the POC value of the current picture.
[0015] A method for encoding image information according to one aspect of the present disclosure comprises: generating a range of picture order count (POC) values for at least one temporal sublayer in which no active entry exists in a reference picture list for a current picture or a subpicture of said current picture; and encoding an error recovery supplemental enhancement information (SEI) message based on the range of POC values for said at least one temporal sublayer.
[0016] According to one aspect of the present disclosure, an apparatus for encoding image information comprises a memory and at least one processor connected to the memory, wherein the at least one processor generates a range of picture order count (POC) values for at least one temporal sublayer in which no active entry exists in a reference picture list for a current picture or a subpicture of the current picture; and, based on the range of POC values for the at least one temporal sublayer, encodes an error recovery supplemental enhancement information (SEI) message.
[0017] In a method or device for encoding the above image information, the last picture having a specific temporal identifier used as a reference picture for the current picture or a sub-picture of the current picture can be derived based on the POC information of the error recovery SEI message corresponding to the specific temporal identifier.
[0018] In the method or device for encoding the above-mentioned image information, the error recovery SEI message may include information regarding the number of the at least one temporal sublayer.
[0019] In a method or device for encoding the above-mentioned image information, the error recovery SEI message may include POC information for deriving the last picture having the specific temporal identifier used as a reference picture for the current picture or the sub-picture of the current picture for each of the at least one temporal sublayer.
[0020] In a method or device for encoding the above image information, the current picture having the specific temporal identifier or the sub-picture of the current picture may not have an active entry of a reference picture list having the specific temporal identifier and a POC value within the range of the POC value.
[0021] In a method or device for encoding the above-mentioned image information, the range of the POC value can be derived based on the difference between the POC value based on the POC information and the POC value of the current picture.
[0022] A method for a bitstream according to one aspect of the present disclosure comprises generating the bitstream, wherein the bitstream is generated based on generating a picture order count (POC) range for at least one temporal sublayer in which no active entry exists in a reference picture list for a current picture or a subpicture of the current picture, and encoding an error recovery supplemental enhancement information (SEI) message based on the POC range for the at least one temporal sublayer, and transmitting the bitstream.
[0023] An apparatus for a bitstream according to one aspect of the present disclosure comprises at least one processor that generates the bitstream, wherein the bitstream is generated based on generating a picture order count (POC) range for at least one temporal sublayer in which no active entry exists in a reference picture list for a current picture or a subpicture of the current picture, and encoding an error recovery supplemental enhancement information (SEI) message based on the POC range for the at least one temporal sublayer, and a transmission unit that transmits the bitstream.
[0024] According to one aspect of the present disclosure, in a computer-readable storage medium, the storage medium stores a bitstream, and the bitstream is generated based on generating a picture order count (POC) range for at least one temporal sublayer, wherein no active entry exists in a reference picture list for a current picture or a subpicture of the current picture, and encoding an error recovery supplemental enhancement information (SEI) message based on the POC range for the at least one temporal sublayer.
[0025] The features briefly summarized above regarding the present disclosure are merely exemplary aspects of the detailed description of the present disclosure that follows and do not limit the scope of the present disclosure.
[0026] According to the present disclosure, the reliability of a coding system including an encoding device and a decoding device can be improved.
[0027] According to the present disclosure, the coding efficiency of a coding system including an encoding device and a decoding device can be improved.
[0028] According to the present disclosure, the data transmission efficiency of a coding system including an encoding device and a decoding device can be improved.
[0029] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure pertains from the description below.
[0030] FIG. 1 is a schematic diagram illustrating a video coding system to which an embodiment according to the present disclosure can be applied.
[0031] FIG. 2 is a schematic diagram showing an encoding device to which an embodiment according to the present disclosure can be applied.
[0032] FIG. 3 is a schematic diagram showing a decoding device to which an embodiment according to the present disclosure can be applied.
[0033] Figure 4 illustrates an exemplary hierarchical structure for a coded video / image.
[0034] FIGS. 5 to 9 are drawings for explaining the operation by SEI message according to one embodiment of the present disclosure.
[0035] FIG. 10 is a diagram illustrating a method for decoding image information according to one embodiment of the present disclosure.
[0036] FIG. 11 is a drawing illustrating a method for encoding image information according to one embodiment of the present disclosure.
[0037] FIG. 12 is a drawing illustrating an exemplary content streaming system to which an embodiment according to the present disclosure can be applied.
[0038] Hereinafter, embodiments of the present disclosure are described in detail with reference to the attached drawings so that those skilled in the art can easily implement them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein.
[0039] In describing the embodiments of the present disclosure, detailed descriptions of known configurations or functions are omitted if it is determined that such descriptions could obscure the essence of the present disclosure. Additionally, parts of the drawings unrelated to the description of the present disclosure have been omitted, and similar parts are denoted by similar reference numerals.
[0040] In the present disclosure, when a component is described as being "connected," "combined," or "joined" with another component, this may include not only a direct connection but also an indirect connection in which another component exists in between. Furthermore, when a component is described as "comprising" or "having" another component, this means that, unless specifically stated otherwise, it does not exclude the other component but may include an additional component.
[0041] In the present disclosure, terms such as first, second, etc. are used solely for the purpose of distinguishing one component from another and do not limit the order or importance of the components unless specifically stated otherwise. Accordingly, within the scope of the present disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and likewise, a second component in one embodiment may be referred to as a first component in another embodiment.
[0042] In this disclosure, distinct components are intended to clearly describe their respective features and do not imply that the components are separate. That is, multiple components may be integrated to form a single hardware or software unit, or a single component may be distributed to form multiple hardware or software units. Accordingly, such integrated or distributed embodiments are included within the scope of this disclosure, unless otherwise noted.
[0043] In the present disclosure, the components described in various embodiments do not necessarily mean essential components, and some may be optional components. Accordingly, embodiments consisting of a subset of the components described in one embodiment are also included within the scope of the present disclosure. Furthermore, embodiments including additional components in addition to the components described in various embodiments are also included within the scope of the present disclosure.
[0044] The present disclosure relates to the encoding and decoding of images. For example, the methods and embodiments disclosed in this document may be applied to methods disclosed in the VVC (versatile video coding) standard, EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard) or next-generation video / image coding standards (e.g., H.267 or H.268).
[0045] The present disclosure presents various embodiments relating to video / image coding, and unless otherwise stated, said embodiments may be performed in combination with one another.
[0046] Unless newly defined in this disclosure, the terms used herein may have the ordinary meanings commonly used in the technical field to which this disclosure belongs.
[0047] In this disclosure, "video" may refer to a set of images over time. In this disclosure, "picture" generally refers to a unit representing a single image at a specific time, and a slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles. A brick may represent a rectangular area of rows of CTUs within a tile in a picture. In this document, tile groups and slices may be used interchangeably. For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.
[0048] In the present disclosure, "pixel" or "pel" may refer to the smallest unit constituting a picture (or image). Additionally, "sample" may be used as a term corresponding to pixel. A sample may generally represent a pixel or a pixel value, may represent only the pixel / pixel value of the luminance component, or may represent only the pixel / pixel value of the chroma component.
[0049] In this disclosure, "unit" may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to that area. A unit may include one luminance block and two chroma (e.g., cb, cr) blocks. Depending on the case, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows.
[0050] In the present disclosure, "current block" may mean one of "current coding block," "current coding unit," "block to be encoded," "block to be decoded," or "block to be processed." When prediction is performed, "current block" may mean "current prediction block" or "block to be predicted." When transformation (inverse transformation) / quantization (inverse quantization) is performed, "current block" may mean "current transformation block" or "block to be transformed." When filtering is performed, "current block" may mean "block to be filtered."
[0051] In the present disclosure, "current block" may mean a block comprising both a luminous component block and a chroma component block, or "luma block of the current block," unless explicitly stated as a chroma block. The luminous component block of the current block may be expressed by explicitly including the explicit description of a luminous component block, such as "luma block" or "current luminous block." Additionally, the chroma component block of the current block may be expressed by explicitly including the explicit description of a chroma component block, such as "chroma block" or "current chroma block."
[0052] In the present disclosure, " / " and "," may be interpreted as "and / or." For example, "A / B" and "A, B" may be interpreted as "A and / or B." Additionally, "A / B / C" and "A, B, C" may mean "at least one of A, B and / or C."
[0053] In the present disclosure, "or" may be interpreted as "and / or". For example, "A or B" may mean 1) "A" only, 2) "B" only, or 3) "A and B". Alternatively, in the present disclosure, "or" may mean "additionally or alternatively".
[0054] FIG. 1 is a schematic diagram illustrating a video / image coding system to which an embodiment according to the present disclosure can be applied.
[0055] Referring to FIG. 1, a video / image coding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image or data in the form of a file or streaming to the receiving device via a digital storage medium or a network.
[0056] The source device may include a video source, an encoding device, and a transmission unit. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.
[0057] A video source may acquire video / images through processes such as video / image capture, synthesis, or generation. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include, for example, a computer, a tablet, and a smartphone, etc., and may generate video / images (electronically). For example, virtual video / images may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process in which related data is generated.
[0058] The encoding device can encode input video / image information. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0059] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device in the form of a file or streaming via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit may include elements for creating a media file through a predetermined file format and elements for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0060] The decoding device can decode video / images by performing a series of procedures such as inverse quantization, inverse transform, and prediction corresponding to the operation of the encoding device.
[0061] The renderer can render the decoded video / image. The rendered video / image can be displayed through the display unit.
[0062] FIG. 2 is a schematic diagram illustrating an encoding device to which an embodiment according to the present disclosure can be applied.
[0063] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a predictor (220), a residual processor (230), an entropy encoder (240), an adder (250), a filter (260), and a memory (270). The predictor (220) may include an inter-predictor (221) and an intra-predictor (222). The residual processor (230) may include a transformer (232), a quantizer (233), a dequantizer (234), and an inverse transformer (235). The residual processor (230) may further include a subtractor (231). The addition unit (250) may be referred to as a reconstructor or a reconstructed block generator. The above-described image segmentation unit (210), prediction unit (220), residual processing unit (230), entropy encoding unit (240), addition unit (250), and filtering unit (260) may be configured by one or more hardware components (e.g., an encoder chipset or processor) according to the embodiment. Additionally, the memory (270) may include a DPB (Decoded Picture Buffer) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.
[0064] The image segmentation unit (210) can divide an input image (or picture, frame) input to an encoding device (200) into one or more processing units. For example, the processing unit may be called a coding unit (CU). A coding unit may be recursively divided into a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, a single coding unit may be divided into multiple coding units of a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. For example, a quad-tree structure may be applied first, and a binary-tree structure and / or a ternary-tree structure may be applied later. Alternatively, a binary-tree structure may be applied first. A coding procedure according to the present disclosure may be performed based on the final coding unit that is no longer divided. In this case, based on coding efficiency according to image characteristics, the maximum coding unit may be used directly as the final coding unit, or, if necessary, the maximum coding unit may be recursively divided into lower-depth coding units so that a coding unit of the optimal size is used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). The prediction unit and the transformation unit may each be divided or partitioned from the final coding unit.The above prediction unit may be a unit of sample prediction, and the above transformation unit may be a unit that derives transformation coefficients and / or a unit that derives a residual signal from transformation coefficients.
[0065] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the context. In general, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chroma component. A sample may be used to refer to a single picture (or image) as a term corresponding to a pixel or pel.
[0066] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (predicted block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input image signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, as illustrated, the unit that subtracts the prediction signal (predicted block, prediction sample array) from the input image signal (original block, original sample array) within the encoding device (200) may be called a subtraction unit (231). The prediction unit (220) can perform a prediction for a block to be processed (hereinafter, current block) and generate a predicted block (predicted block) containing prediction samples for said current block. The prediction unit (220) can determine whether intra prediction is applied or inter prediction is applied in units of the current block or CU. The prediction unit (220) can generate various information regarding prediction, such as prediction mode information, as described below in the description of each prediction mode, and transmit it to the entropy encoding unit (240). The information regarding prediction can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.
[0067] The intra prediction unit (222) can predict the current block by referring to samples within the current picture. The referenced samples may be located near the current block or away from it, depending on the prediction mode. In intra prediction, the prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a Planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is merely an example, and depending on the settings, more or fewer directional prediction modes may be used. The intra prediction unit (222) may also determine the prediction mode applied to the current block by using the prediction mode applied to the surrounding blocks.
[0068] The inter prediction unit (221) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture containing the reference blocks and the reference picture containing the temporal neighboring blocks may be the same or different from each other. The temporal neighboring blocks may be referred to by names such as collocated reference block, collocated CU (colCU), etc. A reference picture containing the aforementioned temporal surrounding blocks may be called a collocated picture (colPic). For example, the inter prediction unit (221) may construct a list of motion information candidates based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of surrounding blocks as motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0069] The prediction unit (220) may generate a prediction signal based on various prediction methods and / or prediction techniques described below. For example, the prediction unit (220) may apply intra prediction or inter prediction for the prediction of the current block, as well as apply intra prediction and inter prediction simultaneously. A prediction method that applies intra prediction and inter prediction simultaneously for the prediction of the current block may be called combined inter and intra prediction (CIIP). Additionally, the prediction unit (220) may be based on an intra block copy (IBC) prediction mode or a palette mode for the prediction of the block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, for example, screen content coding (SCC). IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. Palette mode can be viewed as an example of intra-coding or intra-prediction. When palette mode is applied, sample values within a picture can be signaled based on information regarding palette tables and palette indices.
[0070] The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or to generate a residual signal. The subtraction unit (231) can generate a residual signal (residual signal, residual block, residual sample array) by subtracting the prediction signal (predicted block, prediction sample array) output from the prediction unit (220) from the input image signal (original block, original sample array). The generated residual signal can be transmitted to the conversion unit (232).
[0071] The transformation unit (232) can generate transform coefficients by applying a transformation technique to a residual signal. For example, the transformation technique may include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transformation obtained based on a prediction signal generated using all previously reconstructed pixels. The transformation process may be applied to a block of pixels of the same size in a square, or to a block of variable size that is not square.
[0072] The quantization unit (233) can quantize the transformation coefficients and transmit them to the entropy encoding unit (240). The entropy encoding unit (240) can encode the quantized signal (information regarding the quantized transformation coefficients) and output it as a bitstream. The information regarding the quantized transformation coefficients may be called residual information. The quantization unit (233) can rearrange the block-shaped quantized transformation coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information regarding the quantized transformation coefficients based on the one-dimensional vector-shaped quantized transformation coefficients.
[0073] The entropy encoding unit (240) can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit (190) may encode information required for video / image restoration (e.g., values of syntax elements) together or separately, in addition to quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may further include general constraint information. The signaling information, transmitted information, and / or syntax elements mentioned in the present disclosure may be included in the video / image information. The video / image information may be encoded through the encoding procedure described above and included in the bitstream.
[0074] The above bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmitting a signal output from the entropy encoding unit (240) and / or a storage unit (not shown) for storing it may be provided as an internal / external element of the encoding device (200), or the transmission unit may be provided as a component of the entropy encoding unit (240).
[0075] The quantized transformation coefficients output from the quantization unit (233) can be used to generate a residual signal. For example, a residual signal (residual block or residual samples) can be restored by applying inverse quantization and inverse transformation to the quantized transformation coefficients through the inverse quantization unit (234) and the inverse transformation unit (235).
[0076] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.
[0077] The adder (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit (221) or the intra prediction unit (222). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The adder (250) may be called a reconstructed unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after undergoing filtering as described below.
[0078] The filtering unit (260) can improve subjective / objective quality by applying filtering to the restored signal. For example, the filtering unit (260) can generate a modified restored picture by applying various filtering methods to the restored picture, and can store the modified restored picture in memory (270), specifically in the DPB of memory (170). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (240), as described below in the description of each filtering method. The information regarding filtering can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.
[0079] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter-prediction unit (221). Through this, the encoding device (200) can avoid prediction mismatches between the encoding device (200) and the decoding device when inter-prediction is applied, and can also improve encoding efficiency.
[0080] The DPB in memory (270) can store a modified restored picture to be used as a reference picture in the inter prediction unit (221). Memory (270) can store motion information of blocks from which motion information is derived (or encoded) in the current picture and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter prediction unit (221) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory (270) can store restoration samples of restored blocks in the current picture and transmit them to the intra prediction unit (222).
[0081] FIG. 3 is a schematic diagram illustrating a decoding device to which an embodiment according to the present disclosure can be applied.
[0082] As illustrated in FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-predictor (332) and an intra-predictor (331). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321). The aforementioned entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) may be configured by a single hardware component (e.g., a decoder chipset or a processor) according to an embodiment. Additionally, the memory (360) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.
[0083] When a bitstream containing video / image information is input, the decoding device (300) can restore the image by performing a process corresponding to the process performed by the encoding device (200) of FIG. 2. For example, the decoding device (300) can perform decoding using a processing unit applied in the encoding device (200). Thus, the processing unit for decoding may be, for example, a coding unit. The coding unit may be a coding tree unit, or a maximum coding unit may be obtained by dividing it according to a quad tree structure, a binary tree structure, and / or a binary tree structure. And, the restored image signal decoded and output through the decoding device (300) can be played back through a playback device (not shown).
[0084] The decoding device (300) can receive a signal output from the encoding device (200) of FIG. 2 in the form of a bitstream. The received signal can be decoded through an entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information necessary for image restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may further include general constraint information. The decoding device (300) can decode the picture based on the information regarding the parameter sets and / or the general constraint information. The signaling / received information and / or syntax elements described below can be obtained from the bitstream by decoding through the decoding procedure. For example, the entropy decoding unit (310) can decode information within the bitstream based on coding methods such as exponential coding, CAVLC, or CABAC, and output values of syntax elements required for image restoration and quantized values of transformation coefficients regarding residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using information on the syntax element to be decoded and decoding information of surrounding and decoding target blocks or information on symbols / bins decoded in the previous step, predict the probability of occurrence of the bin according to the determined context model, and perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin after determining the context model. Among the information decoded in the entropy decoding unit (310), information regarding prediction is provided to the prediction unit (330), and residual values for which entropy decoding was performed in the entropy decoding unit (310), i.e., quantized transformation coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive residual signals (residual blocks, residual samples, residual sample array). Additionally, among the information decoded in the entropy decoding unit (310), information regarding filtering can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (300), or the receiving unit may be a component of the entropy decoding unit (310). Meanwhile, the decoding device according to the present document may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit (310), and the sample decoder may include at least one of the inverse quantization unit (321), inverse transform unit (322), adder (340), filtering unit (350), memory (360), inter prediction unit (332), and intra prediction unit (331).
[0085] In the inverse quantization unit (321), the quantized transformation coefficients can be inversely quantized to output transformation coefficients. The inverse quantization unit (321) can rearrange the quantized transformation coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed in the encoding device (200). The inverse quantization unit (321) can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain transformation coefficients.
[0086] In the inverse conversion unit (322), the conversion coefficients can be inversely converted to obtain a residual signal (residual block, residual sample array).
[0087] The prediction unit (330) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit may apply intra prediction or inter prediction for a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called combined inter and intra prediction (CIIP). Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, for example, screen content coding (SCC). IBC basically performs prediction within the current picture, but it can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. The palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and palette index can be included in the above video / image information and signaled.
[0088] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The description of the intra prediction unit (222) may be applied equally to the intra prediction unit (331). The referenced samples may be located in the neighborhood of the current block or located away from it, depending on the prediction mode. In intra prediction, the prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit (331) may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0089] The inter prediction unit (332) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) may construct a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-prediction can be performed based on various prediction modes (techniques), and information regarding the prediction may include information indicating the mode (technique) of inter-prediction for the current block.
[0090] The adder (340) can generate a restoration signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the prediction signal (predicted block, predicted sample array) output from the prediction unit (330) (including the inter prediction unit (332) and / or intra prediction unit (331)). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the restoration block. The description of the adder (250) can be applied equally to the adder (340). The adder (340) may be called a restoration unit or a restoration block generation unit. The generated restoration signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after undergoing filtering as described below.
[0091] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0092] The filtering unit (350) can improve subjective / objective quality by applying filtering to the restored signal. For example, the filtering unit (350) can generate a modified restored picture by applying various filtering methods to the restored picture, and can store the modified restored picture in memory (360), specifically in the DPB of memory (360). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0093] The (modified) restored picture stored in the DPB of the memory (360) can be used as a reference picture in the inter-prediction unit (332). The memory (360) can store motion information of blocks from which motion information within the current picture has been derived (or decoded) and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (332) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (360) can store restoration samples of blocks restored within the current picture and transmit them to the intra-prediction unit (331).
[0094] In this specification, the embodiments described in the filtering unit (260), inter prediction unit (221), and intra prediction unit (222) of the encoding device (200) may be applied to the filtering unit (350), inter prediction unit (332), and intra prediction unit (331) of the decoding device (300) in the same or corresponding manner.
[0095] Figure 4 illustrates an exemplary hierarchical structure for a coded video / image.
[0096] Referring to Figure 4, the coded image is divided into a Video Coding Layer (VCL) that handles the decoding processing of the image and the image itself, a subsystem that transmits and stores the encoded information, and a Network Abstraction Layer (NAL) that exists between the VCL and the subsystem and is responsible for network adaptation functions.
[0097] In VCL, VCL data containing compressed image data (slice data) can be generated, or parameter sets containing information such as Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS), or SEI (Supplemental Enhancement Information) messages that are additionally required in the decoding process of the image can be generated.
[0098] In NAL, a NAL unit can be created by adding header information (NAL unit header) to the Raw Byte Sequence Payload (RBSP) generated in VCL. In this case, the RBSP refers to slice data, parameter sets, SEI messages, etc. generated in VCL. The NAL unit header may include NAL unit type information specified according to the RBSP data included in the NAL unit.
[0099] As shown in FIG. 4, NAL units can be classified into VCL NAL units and Non-VCL NAL units depending on the RBSP generated in VCL. A VCL NAL unit may refer to a NAL unit containing information about an image (slice data), and a Non-VCL NAL unit may refer to a NAL unit containing information necessary to decode an image (parameter set or SEI message).
[0100] The aforementioned VCL NAL unit and Non-VCL NAL unit can be transmitted over a network by attaching header information according to the data specifications of the underlying system. For example, the NAL unit can be transformed into a data format of a specified specification, such as H.266 / VVC file format, RTP (Real-time Transport Protocol), TS (Transport Stream), etc., and transmitted over various networks.
[0101] As described above, the NAL unit type can be specified according to the RBSP data structure included in the NAL unit, and information about this NAL unit type can be stored in the NAL unit header and signaled.
[0102] For example, NAL units can be broadly classified into VCL NAL unit types and Non-VCL NAL unit types depending on whether they contain information about the image (slice data). VCL NAL unit types can be classified according to the properties and types of the picture included in the VCL NAL unit, while Non-VCL NAL unit types can be classified according to the types of parameter sets.
[0103] The following is an example of a NAL unit type specified according to the type of parameter set included in the Non-VCL NAL unit type.
[0104] - APS (Adaptation Parameter Set) NAL unit: Type for the NAL unit containing the APS
[0105] - DPS(Decoding Parameter Set) NAL unit: Type for the NAL unit containing the DPS
[0106] - VPS (Video Parameter Set) NAL unit: Type for the NAL unit containing the VPS
[0107] - SPS (Sequence Parameter Set) NAL unit: Type for the NAL unit containing the SPS
[0108] - PPS(Picture Parameter Set) NAL unit: Type for the NAL unit containing the PPS
[0109] The above-described NAL unit types have syntax information for the NAL unit type, and said syntax information can be stored in the NAL unit header and signaled. For example, said syntax information may be nal_unit_type, and NAL unit types may be specified by the nal_unit_type value.
[0110] A slice header (slice header syntax, slice header information) may include information / parameters that can be commonly applied to the slice. The APS (APS syntax) or PPS (PPS syntax) may include information / parameters that can be commonly applied to one or more slices or pictures. The SPS (SPS syntax) may include information / parameters that can be commonly applied to one or more sequences. The VPS (VPS syntax) may include information / parameters that can be commonly applied to multiple layers. The DPS (DPS syntax) may include information / parameters that can be commonly applied to the entire video. The DPS may include information / parameters related to the concatenation of a CVS (coded video sequence). In the present disclosure, High Level Syntax (HLS) may include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, or slice header syntax.
[0111] In the present disclosure, image / video information encoded by an encoding device and signaled in the form of a bitstream includes not only information related to picture partitioning, intra / inter prediction information, residual information, in-loop filtering information, etc., but may also include information included in the slice header, information included in the APS, information included in the PPS, information included in the SPS, information included in the VPS, and / or information included in the DPS.
[0112] The following descriptor of the present disclosure specifies the parsing process for each syntax element:
[0113] - ae(v): context-adaptive arithmetic entropy-coded syntax element.
[0114] - b(8): A byte (8 bits) with an arbitrary bit sequence pattern. The parsing process for this descriptor is specified by the return value of the function read_bits(8).
[0115] - f(n): A fixed-pattern bit string using n bits written with the left bit first. The parsing process for this descriptor is specified by the return value of the function read_bits(n).
[0116] - i(n): A signed integer using n bits. In the syntax table, if n is "v", the number of bits depends on the values of other syntax elements. The parsing process for this descriptor is specified by the return value of the function read_bits(n), which is interpreted as a two's complement integer representation where the most significant bit is written first.
[0117] - se(v): A signed integer 0th order Exp-Golomb-coded syntax element with the left bit coming first. The parsing process for this descriptor is specified as having order k of 0.
[0118] - st(v): A null-terminated string encoded in Universal Coded Character Set (UCS) Transfer Format-8 (UTF-8) characters as specified in ISO / IEC 10646. The parsing process is specified as follows: st(v) reads and returns a sequence of bytes from the bitstream starting at the byte-aligned position of the bitstream, from the current position to a point that does not contain the next byte-aligned byte, such as 0x00, and moves the bitstream pointer by (stringLength + 1) * 8 bit positions, where stringLength is equal to the number of bytes returned.
[0119] For reference, the st(v) syntax descriptor is used in this specification only when the current position of the bitstream is a byte alignment position.
[0120] - tu(v): A truncated unary code using up to maxVal bits, using maxVal defined in the semantics of the syntax element.
[0121] - u(n): An unsigned integer using n bits. In the syntax table, if n is "v", the number of bits varies depending on the values of other syntax elements. The parsing process for this descriptor is specified by the return value of the function read_bits(n), which is interpreted as the binary representation of the unsigned integer with the most significant bit written first.
[0122] - ue(v): An unsigned integer 0-th order Exp-Golomb-coded syntax element with the left bit coming first. The parsing process for this descriptor is specified as having order k of 0.
[0123] The SEI message related to the present disclosure will be described below.
[0124] Table 1 shows an example of the syntax of an error recovery SEI message according to one embodiment.
[0125] [Table 1]
[0126]
[0127] Below, an example of the semantics of an error recovery SEI message according to one embodiment is described.
[0128] The error recovery SEI message indicates a range of picture order count values for which there are no active entries in the reference picture list for the current picture and subsequent pictures in the decoding order, or their subpictures.
[0129] For reference, in a system where missing encoded pictures or subpictures of encoded pictures are detected and reported to the encoding system via external means, an error recovery SEI message can be used to determine which encoded pictures or subpictures are expected to be suitable for display.
[0130] The use of this SEI message requires the definition of the following variables:
[0131] - Picture order count of the current picture displayed as CurrPoc
[0132] - Number of subpictures displayed as NumSubpics
[0133] An er_ref_subpic_info_present_flag of 0 indicates that an er_ref_pic_delta_poc_minus1 syntax element exists. An er_ref_subpic_info_present_flag of 0 indicates that an er_num_subpics_minus2 syntax element exists.
[0134] er_ref_pic_delta_poc_minus1 + 1 represents the difference in picture order count between the reference picture and the current picture. The value of er_ref_pic_delta_poc must be in the range between 0 and CurrPoc - 1 (including the boundary value).
[0135] er_num_subpics_minus2 + 2 specifies the number of subpictures where the picture order count difference is signaled.
[0136] The bitstream conformity requirement is that when er_num_subpics_minus2 exists, NumSubpics must be equal to er_num_subpics_minus2 + 2.
[0137] er_ref_subpic_delta_poc_minus1[ i ] + 1 represents the difference in picture order count between the reference picture of the i-th subpicture and the current picture. The value of er_ref_subpic_delta_poc_minus1[ i ] must be in the range between 0 and CurrPoc - 1 (including boundary values). If it does not exist, the value of er_ref_subpic_delta_poc_minus1[ i ] is inferred to be equal to er_ref_pic_delta_poc_minus1.
[0138] The variable SubpicRefPoc[ i ] is set to equal CurrPoc - ( er_ref_subpic_delta_poc_minus1[ i ] + 1 ) for i in the range between 0 .. NumSubpics - 1 (including boundary values).
[0139] The function PicOrderCnt( picX ) is specified as follows:
[0140] - PicOrderCnt(picX) = PicOrderCntVal of picture picX
[0141] When an error recovery SEI message exists, the bitstream conformity requirement is that for i in the range between 0 .. NumSubpics - 1 (including boundary values), the subpicture i of the current picture and all subsequent encoded pictures in CLVS in decoding order must not have a slice that references the reference picture refPic in the active entry of the reference picture list where SubpicRefPoc[ i ] < PicOrderCnt( refPic ) < CurrPoc.
[0142] FIGS. 5 to 9 are drawings for explaining the operation by SEI message according to one embodiment of the present disclosure.
[0143] The design of the error recovery SEI message according to one embodiment may not be optimal because it is too strict, especially when temporal scalability is used. FIGS. 5 and 6 illustrate problems related to the proposed SEI when processing lost pictures. In a low-latency configuration having three temporal sublayers, the following encoding structure is assumed. As shown in FIG. 5, the pictures at Tid 0 each refer to a previous Tid 0 picture, the pictures at Tid 1 each refer to a previous Tid 0 picture and a previous Tid 1 picture, and finally, the pictures at Tid 2 each refer to a previous Tid 0 picture, a previous Tid 1 picture, and a previous Tid 2 picture.
[0144] If pic 11, which is the Tid 2 picture, is lost and a message from the decoder is received by the encoder before the encoding of picture 15, the encoder must change the reference pictures, such as pictures 15, 16, and 17, as illustrated in FIG. 6. When using an SEI according to one embodiment, in this example, since the encoder is allowed to use reference pictures only from picture 10 or earlier pictures, pictures 12, 13, and 14 automatically become unusable as reference pictures and must all be changed. This is clearly not optimal and is not a desirable idea in practice. Pictures 12 and 13 are still suitable for reference.
[0145] If the signaling is based on a temporal sublayer, only the pictures in the temporal sublayer where the picture is lost need to be changed, as illustrated in FIGS. 7, 8, and 9. Since the lost picture is in Tid 2, the pictures in Tid 0 and Tid 1 are not actually affected. Using the above example, the encoder actually has several options for changing the reference to the picture in Tid 2 (i.e., picture 17), as illustrated in FIGS. 7, 8, and 9:
[0146] Possibility 1: Instead of referring to pic 11 as shown in Fig. 7, pic 17 may refer to the previous Tid 2 picture (e.g., pic 8).
[0147] Possibility 2: Instead of referring to pic 11 as shown in Fig. 8, pic 17 may refer to the next closest available reference picture (e.g., pic 13).
[0148] Possibility 3: Instead of referring to pic 11 as shown in Fig. 9, pic 17 may refer to the next closest available reference picture (e.g., pic 10) located before the lost picture.
[0149] Any of the three possibilities above can provide better compression performance by better utilizing pictures that are considered unusable for reference by the SEI message according to one embodiment.
[0150] If signaling is designed to consider the temporal sublayer, for the example presented above, the ER SEI includes the following POC values:
[0151] - Tid 0: 12
[0152] - Tid 1: 13
[0153] - Tid 2: 8
[0154] In one embodiment, the following items may be applied individually or in combination.
[0155] - Modify the signaling of Picture Order Count (POC) information for referenced pictures within Error Recovery (ER) SEI messages to one per temporal sub-layer.
[0156] - For each temporal sublayer, the signaled POC specifies that any pictures within the same temporal sublayer that have a POC greater than the signaled POC and smaller than the POC of the current picture (i.e., the picture associated with the SEI message) are not used as reference pictures for the current picture and subsequent pictures in the decoding order.
[0157] - The two points above can also be extended to sub-pictures.
[0158] An example of the syntax and semantics of an error recovery SEI message according to one embodiment is described.
[0159] Table 2 shows an example of the syntax of an error recovery SEI message according to one embodiment.
[0160] [Table 2]
[0161]
[0162] Below, an example of the semantics of an error recovery SEI message according to one embodiment is described.
[0163] The error recovery SEI message indicates a range of picture order count values for different temporal sublayers that have no active entries in the reference picture list for the current picture and subsequent pictures in the decoding order, or their subpictures.
[0164] For reference, in a system where missing encoded pictures or subpictures of encoded pictures are detected and reported to the encoding system via external means, an error recovery SEI message can be used to determine which encoded pictures or subpictures are expected to be suitable for display.
[0165] The use of this SEI message requires the definition of the following variables:
[0166] - Picture order count of the current picture, displayed as CurrPoc.
[0167] - The number of subpictures displayed as NumSubpics.
[0168] er_max_sublayers_minus1 plus 1 specifies the maximum number of temporal sublayers for which the picture order count value is specified.
[0169] An er_ref_subpic_info_present_flag of 0 indicates that an er_ref_pic_delta_poc_minus1 syntax element exists. An er_ref_subpic_info_present_flag of 0 indicates that an er_num_subpics_minus2 syntax element exists.
[0170] er_ref_pic_delta_poc_minus1[ i ] plus 1 specifies the picture order count difference value to derive the last pictures with temporal id i that precede the current picture in decoding order, which can be used as reference pictures for the current picture and subsequent pictures in decoding order. The value of er_ref_pic_delta_poc_minus1[ i ] must be in the range between 0 and CurrPoc - 1 (including boundary values).
[0171] It is a bitstream conformity requirement that there are no pictures in the active reference picture list of the current picture and subsequent pictures in the decoding order that have Tid i and a picture order count in the range between CurrPoc - er_ref_pic_delta_poc_minus1[ i ] and CurrPoc - 1 (including boundary values).
[0172] er_num_subpics_minus2 plus 2 specifies the number of subpictures where the picture order count difference is signaled.
[0173] The bitstream conformity requirement is that when er_num_subpics_minus2 exists, NumSubpics must be equal to er_num_subpics_minus2 + 2.
[0174] er_ref_subpic_delta_poc_minus1[ i ][ j ] plus 1 specifies the picture order count difference value for deriving the last pictures with temporal id i that precede the current picture in decoding order, which can be used as the i-th subpicture of the current picture and as a reference picture for subsequent pictures in decoding order. The value of er_ref_subpic_delta_poc_minus1[ i ][ j ] must be in the range between 0 and CurrPoc - 1 (including boundary values).
[0175] When it does not exist, the value of er_ref_subpic_delta_poc_minus1[ i ][ j ] is inferred to be the same as er_ref_pic_delta_poc_minus1[ j ].
[0176] The bitstream conformity requirement is that there must be no picture in the active reference picture list of the i-th subpicture of the current picture and subsequent pictures in decoding order that has Tid j and a picture order count between CurrPoc - er_ref_subpic_delta_poc_minus1[ i ][ j ] and CurrPoc - 1 (including boundary values).
[0177] An example of the syntax and semantics of an error recovery SEI message according to one embodiment is described.
[0178] Table 3 shows an example of the syntax of an error recovery SEI message according to one embodiment.
[0179] [Table 3]
[0180]
[0181] Below, an example of the semantics of an error recovery SEI message according to one embodiment is described.
[0182] The error recovery SEI message indicates a range of picture order count values for one or more temporal sublayers that have no active entries in the reference picture list for the current picture and subsequent pictures in the decoding order, or their subpictures.
[0183] For reference, in a system where missing encoded pictures or subpictures of encoded pictures are detected and reported to the encoding system via external means, an error recovery SEI message can be used to determine which encoded pictures or subpictures are expected to be suitable for display.
[0184] The use of this SEI message requires the definition of the following variables:
[0185] - Picture order count of the current picture, displayed as CurrPoc.
[0186] - The number of subpictures displayed as NumSubpics.
[0187] er_num_sublayers_minus1 plus 1 specifies the number of temporal sublayers where the picture order count value is displayed.
[0188] An er_ref_subpic_info_present_flag of 0 indicates that an er_ref_pic_delta_poc_minus1 syntax element exists. An er_ref_subpic_info_present_flag of 0 indicates that an er_num_subpics_minus2 syntax element exists.
[0189] er_ref_pic_delta_poc_minus1[ i ] plus 1 specifies the picture order count difference value for deriving the last pictures with temporal Id i that precede the current picture in decoding order, which can be used as reference pictures for the current picture and subsequent pictures in decoding order. The value of er_ref_pic_delta_poc_minus1[ i ] must be in the range between 0 and CurrPoc - 1 (including boundary values).
[0190] For each value i in the range between 0 and er_num_sublayers_minus1 (including boundary values) and each value j in the range between 0 and i (including boundary values), it is a bitstream conformance requirement that all subsequent encoded pictures with temporal sublayer identifier i in the current picture and decoding order within the CLVS must not have an active entry in the reference picture list with temporal sublayer identifier j, and whose picture order count is in the range between CurrPoc - er_ref_pic_delta_poc_minus1[ j ] and CurrPoc - 1 (including boundary values).
[0191] er_num_subpics_minus2 plus 2 specifies the number of subpictures where the picture order count difference is signaled.
[0192] The bitstream conformity requirement is that when er_num_subpics_minus2 exists, NumSubpics must be equal to er_num_subpics_minus2 + 2.
[0193] er_ref_subpic_delta_poc_minus1[ i ][ j ] plus 1 specifies the picture order count difference value for deriving the last pictures with temporal Id i that precede the current picture in decoding order, which can be used as the i-th subpicture of the current picture and as a reference picture for subsequent pictures in decoding order. The value of er_ref_subpic_delta_poc_minus1[ i ][ j ] must be in the range between 0 and CurrPoc - 1 (including boundary values).
[0194] When it does not exist, the value of er_ref_subpic_delta_poc_minus1[ i ][ j ] is inferred to be the same as er_ref_pic_delta_poc_minus1[ j ].
[0195] For each i value in the range between 0 and er_num_subpics_minus2 + 1 (including boundary values), each j value in the range between 0 and er_num_sublayers_minus1 (including boundary values), and each k value in the range between 0 and j (including boundary values), it is a bitstream conformance requirement that all subsequent encoded pictures with a temporal sublayer identifier j in the current picture and decoding order within the CLVS have a picture order count in the range between CurrPoc - er_ref_pic_subpic_delta_poc_minus1[ i ][ k ] and CurrPoc - 1 (including boundary values) and do not have an active entry in the reference picture list of the i-th subpicture with a temporal sublayer identifier k.
[0196] An example of the syntax and semantics of an error recovery SEI message according to one embodiment is described.
[0197] Table 4 shows an example of the syntax of an error recovery SEI message according to one embodiment.
[0198] [Table 4]
[0199]
[0200] Below, an example of the semantics of an error recovery SEI message according to one embodiment is described.
[0201] The error recovery SEI message indicates a range of picture order count values for different temporal sublayers that have no active entries in the reference picture list for the current picture and subsequent pictures in the decoding order, or their subpictures.
[0202] For reference, in a system where missing encoded pictures or subpictures of encoded pictures are detected and reported to the encoding system via external means, an error recovery SEI message can be used to determine which encoded pictures or subpictures are expected to be suitable for display.
[0203] The use of this SEI message requires the definition of the following variables:
[0204] - Picture order count of the current picture, displayed as CurrPoc.
[0205] - The number of subpictures displayed as NumSubpics.
[0206] er_max_sublayers_minus1 plus 1 specifies the maximum number of temporal sublayers for which the picture order count value is specified.
[0207] An er_ref_subpic_info_present_flag of 0 indicates that an er_ref_pic_delta_poc_minus1 syntax element exists. An er_ref_subpic_info_present_flag of 0 indicates that an er_num_subpics_minus2 syntax element exists.
[0208] er_ref_pic_delta_poc_minus1[ i ] plus 1 specifies the picture order count difference value for deriving the last pictures with temporal Id i that precede the current picture in decoding order, which can be used as reference pictures for the current picture and subsequent pictures in decoding order. The value of er_ref_pic_delta_poc_minus1[ i ] must be in the range between 0 and CurrPoc - 1 (including boundary values).
[0209] For any j between 0 and i (including boundary values), the current picture with temporal sublayer identifier i and all subsequent encoded pictures in CLVS in decoding order must not have active entries in the reference picture list with picture order counts in the range of CurrPoc - er_ref_pic_delta_poc_minus1[ j ] to CurrPoc - 1 (including boundary values) as a bitstream conformance requirement.
[0210] er_num_subpics_minus2 plus 2 specifies the number of subpictures where the picture order count difference is signaled.
[0211] The bitstream conformity requirement is that when er_num_subpics_minus2 exists, NumSubpics must be equal to er_num_subpics_minus2 + 2.
[0212] er_ref_subpic_delta_poc_minus1[ i ][ j ] plus 1 specifies the picture order count difference value for deriving the last pictures with temporal Id i that precede the current picture in decoding order, which can be used as the i-th subpicture of the current picture and as a reference picture for subsequent pictures in decoding order. The value of er_ref_subpic_delta_poc_minus1[ i ][ j ] must be in the range between 0 and CurrPoc - 1 (including boundary values).
[0213] When it does not exist, the value of er_ref_subpic_delta_poc_minus1[ i ][ j ] is inferred to be the same as er_ref_pic_delta_poc_minus1[ j ].
[0214] For any k between 0 and j (including boundary values), the bitstream conformance requirement is that the current picture with temporal sublayer identifier j and all subsequent encoded pictures in the CLVS in decoding order must not have an active entry in the reference picture list of the i-th subpicture whose picture order count is in the range between CurrPoc - er_ref_pic_subpic_delta_poc_minus1[ i ][ k ] and CurrPoc - 1 (including boundary values).
[0215] FIG. 10 is a diagram illustrating a method for decoding image information according to one embodiment of the present disclosure.
[0216] The decoding method (S1000) may include operations described below.
[0217] The terms or names described below (e.g., names of syntax elements or variables, etc.) are merely examples, and the technical features of the present disclosure are not limited to the terms or names described below. For example, the image information described below may include various information according to the embodiments described in the present disclosure and may include information described in at least one of the tables described above.
[0218] The operations described below do not constitute an essential component of the decoding method according to one embodiment, and at least some of the operations described below may be omitted. Furthermore, the operations described below do not constitute a sufficient component of the decoding method according to one embodiment, and the previously described operations may be added.
[0219] The sequence of operations illustrated in the drawings regarding the operations described below is merely an example, and the operations described below may be performed in a different order unless it contradicts the operations to be described.
[0220] The operations described below form a single embodiment integrally with the configurations and / or operations described above, unless they conflict with the configurations and / or operations described above, and do not form a separate embodiment distinct from the configurations and / or operations described above.
[0221] The terms "first," "second," "third," etc., used below are merely distinguishing indicators for identifying specific messages or information among multiple pieces of information or messages that may be included within the CLVS, and are not intended to limit the order, importance, or relative priority of these messages or information, or to restrict them to specific embodiments. For example, "first message (or first information)" refers only to at least one message (or information) included in at least one picture unit within the CLVS, and does not imply that said message (or information) must be encoded or decoded first.
[0222] The decoding method (S1000) can be executed by a decoding device including a memory and a processor electrically connected to the memory, for example, by a processor.
[0223] The decoding device can acquire at least one SEI (supplemental enhancement information) message (S1010).
[0224] For example, a processor of a decoding device may acquire supplemental enhancement information (SEI) messages from a bitstream. At least one SEI message may convey a specific type of information that assists in processes related to the decoding, display, or other purposes of image information. Here, the SEI message may not be necessary for the decoding process to determine the sample values of the decoded picture.
[0225] For example, at least one SEI message may include an error recovery SEI message.
[0226] The error recovery SEI message may provide a range of POC values for temporal sublayers for which there are no active entries in the reference picture list for the current picture and subsequent pictures in the decoding order, or their subpictures. For example, missing pictures or their subpictures may be detected and reported to the encoding device, and the encoding device may provide information regarding pictures or their subpictures suitable for display through the error recovery SEI message.
[0227] The error recovery SEI message may have various names such as error recovery message, error recovery related information, error recovery information SEI message, error recovery information, and error recovery SEI, and such names are not limited.
[0228] Error recovery SEI messages can take various forms. For example, an error recovery SEI message may be a syntax element or a syntax structure containing one or more syntax elements. Additionally, an error recovery SEI message may be a raw byte sequence payload (RBSP) containing one or more syntax elements or one or more syntax structures. For example, an error recovery SEI message may be represented as error_recovery or error_recovery_info, but is not limited thereto.
[0229] Error recovery SEI messages may include sublayer count information, picture POC information, subpicture count information, and subpicture POC information.
[0230] The number of sublayers information may represent the number of temporal sublayers or the maximum number of temporal sublayers for which POC information is provided by the error recovery SEI message. For example, the value of the number of sublayers information plus 1 may be equal to the number of temporal sublayers or the maximum number of temporal sublayers.
[0231] Sublayer count information may take various forms and may be represented by various names. For example, sublayer count information may be a syntax element or a syntax structure containing one or more syntax elements. For example, sublayer count information may consist of bits of a fixed length (e.g., 3 bits). Sublayer count information may be represented as er_num_sublayers_minus1 or er_max_sublayers_minus1, but is not limited thereto.
[0232] Picture POC information may represent a POC difference value for deriving a last picture having a specific temporal identifier that can be used as a reference picture for the current picture or the subsequent picture of the current picture (hereinafter referred to as 'subsequent picture'). Here, the specific temporal identifier may include identifiers of temporal sublayers. The value of the picture POC information may be within the range between 0 and the POC value of the current picture minus 1. The current picture or subsequent picture whose temporal identifier is the specific temporal identifier must have a POC value within the range between the POC value of the current picture minus the value of the POC information and the POC value of the current picture minus 1, and the temporal identifier must not have an active entry in the reference picture list that is the specific temporal identifier.
[0233] Picture POC information may take various forms and may be represented by various names. For example, picture POC information may be a syntax element or a syntax structure containing one or more syntax elements. For example, picture POC information may consist of bits of variable length. Picture POC information may be represented as er_ref_pic_delta_poc_minus1, etc., but is not limited thereto.
[0234] The subpicture count information may represent the number of subpictures for which POC information is provided by the error recovery SEI message. For example, the value of the subpicture count information plus 2 may be equal to the number of subpictures.
[0235] Subpicture count information may take various forms and may be represented by various names. For example, subpicture count information may be a syntax element or a syntax structure containing one or more syntax elements. For example, subpicture count information may consist of bits of variable length. Subpicture count information may be represented as er_num_subpics_minus2, etc., but is not limited thereto.
[0236] Subpicture POC information may represent a POC difference value for deriving a last picture having a specific temporal identifier that can be used as a reference picture for subpictures of the current picture or a subsequent picture. Here, the specific temporal identifier may include identifiers of temporal sublayers. The value of the subpicture POC information may be within the range between 0 and the current picture's POC value minus 1. A subpicture of the current picture or a subsequent picture whose temporal identifier is the specific temporal identifier must have a POC value within the range between the current picture's POC value minus the value of the POC information and the current picture's POC value minus 1, and the temporal identifier must not have an active entry in the reference picture list that is the specific temporal identifier. If the subpicture POC information does not exist, the value of the subpicture POC information may be inferred to be the same as the value of the picture POC information.
[0237] Subpicture POC information may take various forms and may be represented by various names. For example, subpicture POC information may be a syntax element or a syntax structure containing one or more syntax elements. For example, subpicture POC information may consist of bits of variable length. Subpicture POC information may be represented as er_ref_subpic_delta_poc_minus1, etc., but is not limited thereto.
[0238] As such, the error recovery SEI message according to one embodiment can independently provide information regarding the reference picture for each temporal layer.
[0239] The decoding device can obtain information related to the reference picture (S1020).
[0240] For example, the processor of the decoding device can obtain POC information for pictures that cannot be used as reference pictures for the current picture, subsequent picture, or their subpictures based on an error recovery SEI message. Specifically, the processor can obtain a range of POC values for temporal sublayers for which there are no active entries in the reference picture list for the current picture, subsequent picture, or their subpicture. In other words, the last picture having a specific temporal identifier used as a reference picture for the current picture, subsequent picture, or their subpicture can be derived based on the POC information of the error recovery SEI message corresponding to the specific temporal identifier.
[0241] Error recovery SEI messages may include picture POC information or subpicture POC information.
[0242] Picture POC information is used to derive the last picture having a specific temporal identifier used as a reference picture for the current picture or a subsequent picture for each of at least one temporal sublayer, and may represent a POC difference value for deriving the last picture having a specific temporal identifier that can be used as a reference picture for the current picture or a subsequent picture.
[0243] Picture POC information, for a current picture or subsequent picture whose temporal identifier is a specific temporal identifier, the POC value is within the range of the current picture's POC value minus the value of the POC information minus the current picture's POC value minus 1, and the temporal identifier does not have an active entry in the reference picture list that is the same as the temporal identifier of the current picture or subsequent picture.
[0244] Additionally, for each of at least one temporal sublayer, the subpicture POC information is used to derive the last picture having a specific temporal identifier used as a reference picture for the subpicture of the current picture or the subsequent picture, and there is a POC difference value for deriving the last picture having a specific temporal identifier that can be used as a reference picture for the subpicture of the current picture or the subsequent picture.
[0245] Subpicture POC information, for a subpicture of a current picture or a subsequent picture whose temporal identifier is a specific temporal identifier, the POC value is within the range between the current picture's POC value minus the value of the POC information and the current picture's POC value minus 1, and the temporal identifier does not have an active entry in the reference picture list that is identical to the temporal identifier of the current picture or the subsequent picture.
[0246] In this way, the processor of the decoding device can obtain a range of POC values of a picture that cannot be used as an active entry of the reference picture list based on an error recovery SEI message.
[0247] As such, the error recovery SEI message according to one embodiment can provide information regarding the reference picture independently for each temporal layer. In other words, beyond simply determining whether a reference is possible based on the Picture Order Count (POC) range, an independent recovery range can be set for each temporal layer (Tid).
[0248] Through this, temporal scalability can be optimized. Previously, all reference pictures within a specific POC range were considered unusable, so even if only the data in the upper layer (High Tid) was lost, lossless pictures in the lower layer (Base Layer) could be excluded from the reference list. In one embodiment, the value of the POC information can be set differently for each Tid, so that even if the data in the upper layer is lost, the reference structure of the lower layer is maintained and decoding can continue.
[0249] In addition, the discarding of unnecessary data can be prevented. If a reference failure occurs only in a specific temporal layer, pictures from other layers can still be used as references. This can improve the user experience (QoE) by allowing the decoder to display as many pictures as possible from available layers instead of freezing the screen or waiting for an entire refresh (IDR) when an error occurs.
[0250] In addition, precise control at the subpicture level is possible. In a 360-degree video or multi-view streaming environment, when a problem occurs only in a specific frame rate (Temporal Layer) of a specific area (Subpicture), data in the remaining area or basic frame rate can be restored and output normally without being affected.
[0251] In addition, clarity regarding bitstream compatibility is enhanced. By ensuring that the POC value is within the range between the current picture's POC value minus the value of the POC information minus 1, and that the temporal identifier cannot be included in the reference list, the error handling operations between the encoder and the decoder can be synchronized and malfunctions can be prevented.
[0252] As a result, the reliability of the coding system can be improved, and the coding efficiency and data transmission efficiency of the coding system can be enhanced.
[0253] FIG. 11 is a drawing illustrating a method for encoding image information according to one embodiment of the present disclosure.
[0254] The encoding method (S1100) may include operations described below.
[0255] The terms or names described below (e.g., names of syntax elements or variables, etc.) are merely examples, and the technical features of the present disclosure are not limited to the terms or names described below. For example, the image information described below may include various information according to the embodiments described in the present disclosure and may include information described in at least one of the tables described above.
[0256] The operations described below do not constitute an essential component of the decoding method according to one embodiment, and at least some of the operations described below may be omitted. Furthermore, the operations described below do not constitute a sufficient component of the decoding method according to one embodiment, and the previously described operations may be added.
[0257] The sequence of actions illustrated in the drawings regarding the actions described below is merely an example, and the actions described below may be performed in a different order as long as it does not contradict the causal relationship of the actions to be described.
[0258] The operations described below form a single embodiment integrally with the configurations and / or operations described above, unless they conflict with the configurations and / or operations described above, and do not form a separate embodiment distinct from the configurations and / or operations described above.
[0259] The terms "first," "second," "third," etc., used below are merely distinguishing indicators for identifying specific messages or information among multiple pieces of information or messages that may be included within the CLVS, and are not intended to limit the order, importance, or relative priority of these messages or information, or to restrict them to specific embodiments. For example, "first message (or first information)" refers only to at least one message (or information) included in at least one picture unit within the CLVS, and does not imply that said message (or information) must be encoded or decoded first.
[0260] The encoding method (S1100) can be executed by an encoding device including a memory and a processor electrically connected to the memory, for example, by a processor.
[0261] The encoding device can generate information related to the reference picture (S1110).
[0262] For example, the processor of the encoding device may generate POC information for pictures that cannot be used as reference pictures for the current picture, subsequent picture, or subpicture.
[0263] The decoding device may detect missing encoded pictures or subpictures of encoded pictures and report them to the encoding device via external means. Based on the report from the decoding device, the processor of the encoding device may identify the missing pictures and their server pictures and prevent the missing pictures and their server pictures from being used as reference pictures. Specifically, the processor may generate a range of POC values for temporal sublayers for which there are no active entries in the reference picture list for the current picture, subsequent picture, or their subpicture.
[0264] The processor can generate picture POC information or subpicture POC information based on the range of POC values for temporal sublayers that do not have active entries in the reference picture list.
[0265] Picture POC information is used to derive the last picture having a specific temporal identifier used as a reference picture for the current picture or a subsequent picture for each of at least one temporal sublayer, and may represent a POC difference value for deriving the last picture having a specific temporal identifier that can be used as a reference picture for the current picture or a subsequent picture.
[0266] Picture POC information, for a current picture or subsequent picture whose temporal identifier is a specific temporal identifier, the POC value is within the range of the current picture's POC value minus the value of the POC information minus the current picture's POC value minus 1, and the temporal identifier does not have an active entry in the reference picture list that is the same as the temporal identifier of the current picture or subsequent picture.
[0267] Additionally, for each of at least one temporal sublayer, the subpicture POC information is used to derive the last picture having a specific temporal identifier used as a reference picture for the subpicture of the current picture or the subsequent picture, and there is a POC difference value for deriving the last picture having a specific temporal identifier that can be used as a reference picture for the subpicture of the current picture or the subsequent picture.
[0268] Subpicture POC information, for a subpicture of a current picture or a subsequent picture whose temporal identifier is a specific temporal identifier, the POC value is within the range between the current picture's POC value minus the value of the POC information and the current picture's POC value minus 1, and the temporal identifier does not have an active entry in the reference picture list that is identical to the temporal identifier of the current picture or the subsequent picture.
[0269] Thus, the processor of the encoding device can generate picture POC information or subpicture POC information regarding the range of POC values of a picture that cannot be used as an active entry of a reference picture list, based on missing encoded pictures or subpictures of encoded pictures.
[0270] The encoding device can encode at least one SEI message including an error recovery SEI message (S1120).
[0271] For example, the processor of the encoding device can generate an error recovery SEI message containing picture POC information or subpicture POC information based on picture POC information or subpicture POC information, and encode at least one SEI message containing the error recovery SEI message.
[0272] At least one SEI message may convey a specific type of information that assists in processes related to the decoding, display, or other purposes of image information. Here, the SEI message may not be necessary for the decoding process to determine the sample values of the decoded picture.
[0273] The error recovery SEI message can provide a range of POC values for a temporal sublayer that has no active entry in the reference picture list for the current picture and subsequent pictures in the decoding order, or subpictures of those pictures.
[0274] The error recovery SEI message may have various names such as error recovery message, error recovery related information, error recovery information SEI message, error recovery information, and error recovery SEI, and such names are not limited.
[0275] Error recovery SEI messages can take various forms. For example, an error recovery SEI message may be a syntax element or a syntax structure containing one or more syntax elements. Additionally, an error recovery SEI message may be a raw byte sequence payload (RBSP) containing one or more syntax elements or one or more syntax structures. For example, an error recovery SEI message may be represented as error_recovery or error_recovery_info, but is not limited thereto.
[0276] Error recovery SEI messages may include sublayer count information, picture POC information, subpicture count information, and subpicture POC information.
[0277] The sublayer count information, picture POC information, subpicture count information, and subpicture POC information may be the same as the sublayer count information, picture POC information, subpicture count information, and subpicture POC information described earlier together with FIG. 5. The description of the sublayer count information, picture POC information, subpicture count information, and subpicture POC information is replaced by the description of the sublayer count information, picture POC information, subpicture count information, and subpicture POC information described earlier together with FIG. 5.
[0278] As such, the error recovery SEI message according to one embodiment can independently provide information regarding the reference picture for each temporal layer.
[0279] An error recovery SEI message according to one embodiment may include information regarding a reference picture independently for each temporal layer. In other words, beyond simply determining whether a reference is possible based on the Picture Order Count (POC) range, an independent recovery range may be set for each temporal layer (Tid).
[0280] Through this, temporal scalability can be optimized. Previously, all reference pictures within a specific POC range were considered unusable, so even if only the data in the upper layer (High Tid) was lost, lossless pictures in the lower layer (Base Layer) could be excluded from the reference list. In one embodiment, the value of the POC information can be set differently for each Tid, so that even if the data in the upper layer is lost, the reference structure of the lower layer is maintained and decoding can continue.
[0281] In addition, the discarding of unnecessary data can be prevented. If a reference failure occurs only in a specific temporal layer, pictures from other layers can still be used as references. This can improve the user experience (QoE) by allowing the decoder to display as many pictures as possible from available layers instead of freezing the screen or waiting for an entire refresh (IDR) when an error occurs.
[0282] In addition, precise control at the subpicture level is possible. In a 360-degree video or multi-view streaming environment, when a problem occurs only in a specific frame rate (Temporal Layer) of a specific area (Subpicture), data in the remaining area or basic frame rate can be restored and output normally without being affected.
[0283] In addition, clarity regarding bitstream compatibility is enhanced. By ensuring that the POC value is within the range between the current picture's POC value minus the value of the POC information minus 1, and that the temporal identifier cannot be included in the reference list, the error handling operations between the encoder and the decoder can be synchronized and malfunctions can be prevented.
[0284] As a result, the reliability of the coding system can be improved, and the coding efficiency and data transmission efficiency of the coding system can be enhanced.
[0285] At least one SEI message encoded according to the encoding method (S1100) described above can be output in the form of a bitstream. In other words, the bitstream can be generated based on at least one SEI message encoded according to the encoding method (S1100) described above.
[0286] A bitstream generated based on at least one SEI message encoded according to the encoding method (S1100) described above can be stored on a computer-readable storage medium.
[0287] A bitstream generated based on at least one SEI message encoded according to the encoding method (S1100) described above can be transmitted through a transmission unit and / or a transmission medium.
[0288] FIG. 12 is a drawing illustrating an exemplary content streaming system to which an embodiment according to the present disclosure can be applied.
[0289] As illustrated in FIG. 12, a content streaming system to which an embodiment of the present disclosure is applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0290] The above encoding server compresses content input from multimedia input devices, such as smartphones, cameras, and camcorders, into digital data to generate a bitstream and transmits it to the streaming server. As another example, if multimedia input devices, such as smartphones, cameras, and camcorders, generate the bitstream directly, the encoding server may be omitted.
[0291] The bitstream may be generated by a video encoding method and / or encoding device to which an embodiment of the present disclosure is applied, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0292] The streaming server transmits multimedia data to a user device based on a user request through a web server, and the web server can act as a medium to inform the user of available services. When a user requests a desired service from the web server, the web server transmits it to the streaming server, and the streaming server can transmit multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server can perform the role of controlling commands and responses between each device within the content streaming system.
[0293] The streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, to provide a seamless streaming service, the streaming server can store the bitstream for a certain period of time.
[0294] Examples of the above user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs (head-mounted displays)), digital TVs, desktop computers, digital signage, etc.
[0295] Each server within the above-mentioned content streaming system can be operated as a distributed server, and in this case, data received from each server can be processed in a distributed manner.
[0296] The scope of the present disclosure includes software or machine-executable instructions (e.g., operating system, application, firmware, program, etc.) that enable an operation according to a method of various embodiments to be executed on a device or computer, and a non-transitory computer-readable medium on which such software or instructions, etc. are stored and executable on a device or computer.
[0297] An embodiment according to the present disclosure can be used to encode / decode images.
Claims
1. Regarding the method of decoding video information, Obtain at least one SEI message containing an error recovery SEI (supplemental enhancement information) message from a bitstream; A method comprising obtaining a range of POC (picture order count) values for at least one temporal sublayer, based on the above error recovery SEI message, where no active entry exists in the reference picture list for the current picture or the sub-picture of the current picture.
2. In Paragraph 1, A method in which the last picture having a specific temporal identifier used as a reference picture for the current picture or a sub-picture of the current picture is derived based on the POC information of the error recovery SEI message corresponding to the specific temporal identifier.
3. In Paragraph 1, A method in which the above error recovery SEI message includes information regarding the number of at least one temporal sublayer.
4. In Paragraph 1, A method comprising the above error recovery SEI message including POC information for deriving the last picture having the specific temporal identifier used as a reference picture for the current picture or the sub-picture of the current picture for each of the at least one temporal sublayer.
5. In Paragraph 4, A method in which a current picture having the specific temporal identifier or a sub-picture of the current picture does not have an active entry of a reference picture list having the specific temporal identifier, having a POC value within the range of the POC value.
6. In Paragraph 5, A method for deriving the range of the above POC value based on the difference between the POC value based on the above POC information and the POC value of the above current picture.
7. Regarding the method of encoding video information, Generating a range of POC (picture order count) values for at least one temporal sublayer where no active entry exists in the reference picture list for the current picture or the sub-picture of the said current picture; A method comprising encoding an error recovery SEI (supplemental enhancement information) message based on a range of POC values for at least one temporal sublayer.
8. In Paragraph 7, A method in which the last picture having a specific temporal identifier used as a reference picture for the current picture or a sub-picture of the current picture is derived based on the POC information of the error recovery SEI message corresponding to the specific temporal identifier.
9. In Paragraph 7, A method in which the above error recovery SEI message includes information regarding the number of at least one temporal sublayer.
10. In Paragraph 7, A method comprising the above error recovery SEI message including POC information for deriving the last picture having the specific temporal identifier used as a reference picture for the current picture or the sub-picture of the current picture for each of the at least one temporal sublayer.
11. In Paragraph 10, A method in which a current picture having the specific temporal identifier or a sub-picture of the current picture does not have an active entry of a reference picture list having the specific temporal identifier, having a POC value within the range of the POC value.
12. In Paragraph 11, A method for deriving the range of the above POC value based on the difference between the POC value based on the above POC information and the POC value of the above current picture.
13. A computer-readable storage medium for storing a bitstream generated based on the method according to paragraph 7.
14. Regarding methods concerning bitstreams, The bitstream is generated based on generating a POC (picture order count) range for at least one temporal sublayer in which no active entry exists in the reference picture list for the current picture or the sub-picture of the current picture, and encoding an error recovery SEI (supplemental enhancement information) message based on the POC range for the at least one temporal sublayer. A method including transmitting the above bitstream.