Image encoding / decoding method and device, and recording medium storing bitstream
The method addresses inefficiencies in managing display overlays and resampling for high-resolution videos by signaling and configuring overlay information, enhancing encoding/decoding efficiency and quality in video compression.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2025-10-17
- Publication Date
- 2026-04-23
AI Technical Summary
Existing video compression technologies struggle to efficiently handle high-resolution and high-quality video formats, particularly in managing display overlays and resampling information, which affects encoding and decoding efficiency.
A method and apparatus for signaling and configuring display overlay information, including the number of overlays, offset parameters, and resampling information, within the bitstream, allowing for precise control and efficient encoding/decoding of high-resolution videos.
Enhances the encoding and decoding efficiency of high-resolution videos by accurately managing display overlays and resampling, improving the overall quality and performance of video compression.
Smart Images

Figure KR2025016442_23042026_PF_FP_ABST
Abstract
Description
Video encoding / decoding method and device, and a recording medium storing a bitstream
[0001] The present invention relates to a video encoding / decoding method and apparatus, and a recording medium storing a bitstream.
[0002] Recently, the demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition) video, has been increasing across various application fields, and accordingly, high-efficiency video compression technologies are being discussed.
[0003] Various image compression technologies exist, such as inter-prediction technology that predicts pixel values in the current picture from previous or subsequent pictures, intra-prediction technology that predicts pixel values in the current picture using pixel information within the current picture, and entropy coding technology that assigns short codes to values with high frequency and long codes to values with low frequency; by utilizing these image compression technologies, image data can be effectively compressed for transmission or storage.
[0004] The present disclosure provides a method and apparatus for configuring display overlay information.
[0005] The present disclosure provides a method and apparatus for signaling display overlay information.
[0006] The present disclosure provides a method and apparatus for configuring information regarding the number of display overlays, offset parameters for the overlays, and / or resampling information.
[0007] The present disclosure provides a method and apparatus for signaling information about the number of display overlays, offset parameters for the overlays, and / or resampling information.
[0008] An image decoding method and apparatus according to the present disclosure receive a bitstream including an encoded video picture and can restore the encoded video picture included in the bitstream. The bitstream may include information regarding the number of display overlays.
[0009] In the image decoding method and apparatus according to the present disclosure, the number of display overlays may be specified as 1 or more based on the value of information regarding the number of display overlays.
[0010] In the image decoding method and apparatus according to the present disclosure, information regarding the number of display overlays can be obtained from the network abstraction layer (NAL) unit of the bitstream.
[0011] In the image decoding method and apparatus according to the present disclosure, the value obtained by adding 1 to the value of the information regarding the number of display overlays can specify the number of display overlays.
[0012] In the image decoding method and apparatus according to the present disclosure, the value of the information regarding the number of the display overlays may have a value within the range of 0 to 31.
[0013] In the image decoding method and apparatus according to the present disclosure, based on the number of display overlays being one or more, the bitstream may include an offset parameter that specifies the horizontal and vertical positions for the upper-left sample of at least one display overlay.
[0014] In the image decoding method and apparatus according to the present disclosure, the bitstream includes an offset parameter existence flag indicating whether the offset parameter exists, and the offset parameter can be signaled from the bitstream based on the value of the offset parameter existence flag.
[0015] In the image decoding method and apparatus according to the present disclosure, based on the number of display overlays being one or more, the bitstream may include resampling information that specifies the width and height of a sample array of at least one display overlay.
[0016] In the image decoding method and apparatus according to the present disclosure, the bitstream includes a resampling enabled flag indicating whether a display overlay component is capable of resampling, and the resampling information can be signaled from the bitstream based on the value of the offset parameter flag.
[0017] A video encoding method and apparatus according to the present disclosure may receive a video picture to be encoded, encode the received video picture to generate video information regarding the video picture, generate information regarding the number of display overlays, and generate a bitstream including the video information and the information regarding the number of display overlays. Based on the value of the information regarding the number of display overlays, the number of display overlays may be specified as one or more. The information regarding the number of display overlays may be encoded in a network abstraction layer (NAL) unit of the bitstream.
[0018] A computer-readable digital storage medium is provided that stores encoded video / image information that causes an image decoding method to be performed by a decoding device according to the present disclosure.
[0019] A computer-readable digital storage medium is provided that stores video / image information generated according to the image encoding method according to the present disclosure.
[0020] A method and apparatus for transmitting video / image information generated according to the image encoding method according to the present disclosure are provided.
[0021] According to the present disclosure, a display overlay that is one within the displayed picture can also be expressed.
[0022] According to the present disclosure, if there are offset parameters and resampling information within the display overlay information, signaling can be performed regardless of the number of display overlays.
[0023] FIG. 1 illustrates a video / image coding system according to the present disclosure.
[0024] FIG. 2 shows a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and to which encoding of a video / image signal is performed.
[0025] FIG. 3 shows a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and to which decoding of a video / image signal is performed.
[0026] FIG. 4 illustrates a method for restoring a video picture performed in a decoding device (300) according to the present disclosure.
[0027] FIG. 5 illustrates a schematic configuration of a decoding device (300) that performs a method for restoring a video picture according to the present disclosure.
[0028] FIG. 6 illustrates a method for generating a bitstream performed in an encoding device (200) according to the present disclosure.
[0029] FIG. 7 illustrates a schematic configuration of an encoding device (200) that performs a method for generating a bitstream according to the present disclosure.
[0030] FIG. 8 shows an example of a content streaming system to which embodiments of the present disclosure can be applied.
[0031] The present disclosure is susceptible to various modifications and may have various embodiments; specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, and it should be understood that it includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the present disclosure. Similar reference numerals have been used for similar components in the description of each drawing.
[0032] Terms such as "first," "second," etc., may be used to describe various components, but said components should not be limited by said terms. Such terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present disclosure, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.
[0033] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.
[0034] The terms used in this application are used merely to describe specific embodiments and are not intended to limit the disclosure. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, terms such as “comprising” or “having” are intended to specify the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0035] The present disclosure relates to video / video coding. For example, the methods / embodiments disclosed herein may be applied to methods disclosed in the VVC (versatile video coding) standard. Additionally, the methods / embodiments disclosed herein may be applied to methods disclosed in the EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard), or next-generation video / video coding standards (e.g., H.267 or H.268).
[0036] This specification presents various embodiments regarding video / image coding, and unless otherwise noted, said embodiments may be performed in combination with one another.
[0037] In this specification, "video" may refer to a set of images over time. "Picture" generally refers to a unit representing a single image of a specific time period, and "slice" or "tile" is a unit that constitutes a part of a picture in coding. A slice or tile may contain one or more coding tree units (CTUs). A picture may consist of one or more slices or tiles. A tile is a rectangular area composed of multiple CTUs within a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of CTUs having a height equal to the height of the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of CTUs having a height specified by the picture parameter set and a width equal to the width of the picture. CTUs within a tile are arranged continuously according to the CTU raster scan, whereas tiles within a picture may be arranged continuously according to the tile's raster scan. A single slice may include an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that can be exclusively contained in a single NAL unit. Meanwhile, a single picture may be divided into two or more subpictures. A subpicture may be a rectangular area of one or more slices within a picture.
[0038] A pixel, or pel, can refer to the smallest unit that constitutes a picture (or image). Additionally, the term 'sample' may be used as a counterpart to pixel. A sample generally represents a pixel or its value, and it may represent only the pixel / pixel value of the luminance (luma) component or only the pixel / pixel value of the chroma component.
[0039] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to that area. A unit may include one luminance block and two chroma (e.g., cb, cr) blocks. Depending on the case, the term unit may be used interchangeably with terms such as block or area. In general, an MxN block may include samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows.
[0040] In this specification, "A or B" may mean "only A," "only B," or "both A and B." Alternatively, in this specification, "A or B" may be interpreted as "A and / or B." For example, in this specification, "A, B or C" may mean "only A," "only B," "only C," or "any combination of A, B and C."
[0041] A slash ( / ) or a comma used in this specification may mean "and / or." For example, "A / B" may mean "A and / or B." Accordingly, "A / B" may mean "only A," "only B," or "both A and B." For example, "A, B, C" may mean "A, B or C."
[0042] In this specification, "at least one of A and B" may mean "only A," "only B," or "both A and B." Additionally, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted as synonymous with "at least one of A and B."
[0043] Additionally, in this specification, "at least one of A, B and C" may mean "only A," "only B," "only C," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" may mean "at least one of A, B and C."
[0044] Additionally, parentheses used in this specification may mean "for example." Specifically, where indicated as "prediction (intra-prediction)," "intra-prediction" may be proposed as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra-prediction," and "intra-prediction" may be proposed as an example of "prediction." Furthermore, even where indicated as "prediction (i.e., intra-prediction)," "intra-prediction" may be proposed as an example of "prediction."
[0045] Technical features described individually within a single drawing in this specification may be implemented individually or simultaneously.
[0046] FIG. 1 illustrates a video / image coding system according to the present disclosure.
[0047] Referring to FIG. 1, the video / image coding system may include a first device (source device) and a second device (receiving device).
[0048] A source device can transmit encoded video / image information or data in the form of a file or streaming to a receiving device via a digital storage medium or a network. The source device may include a video source, an encoding device, and a transmission unit. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. A transmitter may be included in the encoding device. A receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.
[0049] A video source may acquire video / images through processes such as video / image capture, synthesis, or generation. The video source may include a video / image capture device and / or a video / image generation device. A video / image capture device may include one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device may include a computer, a tablet, a smartphone, etc., and may generate video / images (electronically). For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process in which related data is generated.
[0050] The encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0051] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device in the form of a file or streaming via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit may include elements for creating a media file through a predetermined file format and elements for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0052] The decoding device can decode video / images by performing a series of procedures such as inverse quantization, inverse transform, and prediction corresponding to the operation of the encoding device.
[0053] The renderer can render the decoded video / image. The rendered video / image can be displayed through the display unit.
[0054] FIG. 2 shows a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and to which encoding of a video / image signal is performed.
[0055] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a predictor (220), a residual processor (230), an entropy encoder (240), an adder (250), a filter (260), and a memory (270). The predictor (220) may include an inter-predictor (221) and an intra-predictor (222). The residual processor (230) may include a transformer (232), a quantizer (233), a dequantizer (234), and an inverse transformer (235). The residual processor (230) may further include a subtractor (231). The addition unit (250) may be referred to as a reconstructor or a reconstructed block generator. The above-described image segmentation unit (210), prediction unit (220), residual processing unit (230), entropy encoding unit (240), addition unit (250), and filtering unit (260) may be configured by one or more hardware components (e.g., an encoding device chipset or processor) according to the embodiment. Additionally, the memory (270) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.
[0056] The image segmentation unit (210) can divide an input image (or picture, frame) input to an encoding device (200) into one or more processing units. For example, the processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure.
[0057] For example, a single coding unit may be divided into multiple coding units with a deeper depth based on a quad tree structure, a binary tree structure, and / or a terrestrial structure. In this case, for example, the quad tree structure may be applied first and the binary tree structure and / or terrestrial structure may be applied later. Alternatively, the binary tree structure may be applied before the quad tree structure. A coding procedure according to the present specification may be performed based on a final coding unit that is no longer divided. In this case, based on coding efficiency according to image characteristics, the maximum coding unit may be used directly as the final coding unit, or, if necessary, the coding unit may be recursively divided into coding units of a lower depth so that a coding unit of the optimal size may be used as the final coding unit. Here, the term "coding procedure" may include procedures such as prediction, transformation, and restoration described below.
[0058] As another example, the processing unit may further include a Prediction Unit (PU) or a Transform Unit (TU). In this case, the Prediction Unit and the Transform Unit may each be divided or partitioned from the aforementioned final coding unit. The Prediction Unit may be a unit for sample prediction, and the Transform Unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from transformation coefficients.
[0059] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the context. In general, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chroma component. A sample may be used to refer to a single picture (or image) as a term corresponding to a pixel or pel.
[0060] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (prediction block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input video signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input video signal (original block, original sample array) within the encoding device (200) may be called a subtraction unit (231).
[0061] The prediction unit (220) performs a prediction for a block to be processed (hereinafter referred to as the current block) and can generate a predicted block containing prediction samples for the current block. The prediction unit (220) can determine whether intra prediction is applied or inter prediction is applied at the current block or CU level. The prediction unit (220) can generate various information regarding the prediction, such as prediction mode information, as described below in the description of each prediction mode, and transmit it to the entropy encoding unit (240). The information regarding the prediction can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.
[0062] The intra prediction unit (222) can predict the current block by referring to samples within the current picture. The referenced samples may be located near the current block or at a certain distance from the current block depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The non-directional mode may include at least one DC mode or a planar mode. The directional mode may include 33 directional modes or 65 directional modes depending on the degree of fineness of the prediction direction. However, this is merely an example, and depending on the settings, more or fewer directional modes may be used. The intra prediction unit (222) may determine the prediction mode applied to the current block by using the prediction mode applied to the surrounding blocks.
[0063] The inter prediction unit (221) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture containing the reference blocks and the reference picture containing the temporal neighboring blocks may be the same or different. The above temporal surrounding blocks may be referred to by names such as collocated reference block, collocated CU (colCU), etc., and the reference picture containing the above temporal surrounding blocks may be referred to as a collocated picture (colPic). For example, the inter prediction unit (221) may construct a list of motion information candidates based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of surrounding blocks as motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vectors of surrounding blocks are used as motion vector predictors, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0064] The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit may apply intra prediction or inter prediction for prediction of a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called a combined inter and intra prediction (CIIP) mode. Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in screen content coding (SCC) for games. IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may utilize at least one of the inter prediction techniques described in this specification. The palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, sample values within the picture can be signaled based on information regarding the palette table and palette index. The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or to generate a residual signal.
[0065] The transformation unit (232) can generate transform coefficients by applying a transformation technique to a residual signal. For example, the transformation technique may include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transformation obtained based on a prediction signal generated using all previously restored pixels. Additionally, the transformation process may be applied to a pixel block of the same size in a square, or to a block of variable size that is not square.
[0066] The quantization unit (233) quantizes the transformation coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information regarding the quantized transformation coefficients) and output it as a bitstream. The information regarding the quantized transformation coefficients may be called residual information. The quantization unit (233) can rearrange the block-shaped quantized transformation coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information regarding the quantized transformation coefficients based on the one-dimensional vector-shaped quantized transformation coefficients.
[0067] The entropy encoding unit (240) can perform various encoding methods such as exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit (240) may encode information required for video / image restoration (e.g., values of syntax elements, etc.) together or separately, in addition to quantized transform coefficients.
[0068] Encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream at the level of a Network Abstraction Layer (NAL) unit. The video / image information may further include information regarding various parameter sets, such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Additionally, the video / image information may further include general constraint information. In this specification, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the encoding procedure described above and included in the bitstream. The bitstream may be transmitted over a network or stored on a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) that transmits the signal output from the entropy encoding unit (240) and / or a storage unit (not shown) that stores it may be configured as internal / external elements of the encoding device (200), or the transmission unit may be included in the entropy encoding unit (240).
[0069] The quantized transform coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, a residual signal (residual block or residual samples) can be restored by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (234) and the inverse transformation unit (235). The adder (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter-prediction unit (221) or the intra-prediction unit (222). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The adder (250) may be called a reconstruction unit or a reconstruction block generation unit. The generated restoration signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after filtering as described below. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.
[0070] The filtering unit (260) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (260) can generate a modified restored picture by applying various filtering methods to the restored picture, and can store the modified restored picture in memory (270), specifically in the DPB of memory (270). The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (240). The information regarding filtering can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.
[0071] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter-prediction unit (221). Through this, when inter-prediction is applied, the encoding device can avoid prediction mismatches between the encoding device (200) and the decoding device, and can also improve encoding efficiency.
[0072] The DPB of the memory (270) can store the modified restored picture to be used as a reference picture in the inter-prediction unit (221). The memory (270) can store motion information of blocks from which motion information is derived (or encoded) within the current picture and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (221) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (270) can store restoration samples of the blocks restored within the current picture and transmit them to the intra-prediction unit (222).
[0073] FIG. 3 shows a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and to which decoding of a video / image signal is performed.
[0074] Referring to FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-predictor (332) and an intra-predictor (331). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321).
[0075] The aforementioned entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) may be configured by a single hardware component (e.g., a decoding device chipset or processor) according to an embodiment. Additionally, the memory (360) may include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.
[0076] When a bitstream containing video / image information is input, the decoding device (300) can restore the image in correspondence with the process in which the video / image information is processed by the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied by the encoding device. Accordingly, the processing unit for decoding may be a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a binary tree structure. One or more conversion units may be derived from the coding unit. And, the restored image signal decoded and output through the decoding device (300) can be played back through a playback device.
[0077] The decoding device (300) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through an entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may further include general constraint information. The decoding device can decode the picture based on information regarding the parameter sets and / or the general constraint information. The signaling / receiving information and / or syntax elements described below in this specification may be decoded through the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit (310) can decode information within the bitstream based on coding methods such as exponential chord coding, CAVLC, or CABAC, and output the values of syntax elements required for image restoration and the quantized values of transformation coefficients regarding residuals. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, determine a context model using information on the syntax element to be decoded and decoding information of surrounding and decoding target blocks or information on symbols / bins decoded in the previous step, predict the probability of occurrence of the bin according to the determined context model, and perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin after determining the context model. Among the information decoded in the entropy decoding unit (310), information regarding prediction is provided to the prediction unit (inter prediction unit (332) and intra prediction unit (331)), and the residual value for which entropy decoding was performed in the entropy decoding unit (310), i.e., quantized transformation coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive residual signals (residual blocks, residual samples, residual sample array). Additionally, among the information decoded in the entropy decoding unit (310), information regarding filtering can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (300), or the receiving unit may be a component of the entropy decoding unit (310).
[0078] Meanwhile, the decoding device according to the present specification may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoding device (video / image / picture information decoding device) and a sample decoding device (video / image / picture sample decoding device). The information decoding device may include the entropy decoding unit (310), and the sample decoding device may include at least one of the inverse quantization unit (321), inverse transform unit (322), adder (340), filtering unit (350), memory (360), inter prediction unit (332), and intra prediction unit (331).
[0079] In the inverse quantization unit (321), the quantized transformation coefficients can be inversely quantized to output transformation coefficients. The inverse quantization unit (321) can rearrange the quantized transformation coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain transformation coefficients.
[0080] In the inverse conversion unit (322), the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).
[0081] The prediction unit (320) can perform a prediction for the current block and generate a predicted block containing prediction samples for the current block. The prediction unit (320) can determine whether an intra prediction or an inter prediction is applied to the current block based on information regarding the prediction output from the entropy decoding unit (310), and can determine a specific intra / inter prediction mode.
[0082] The prediction unit (320) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit (320) may apply intra prediction or inter prediction for prediction of a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called a combined inter and intra prediction (CIIP) mode. Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, such as SCC (screen content coding). IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may utilize at least one of the inter prediction techniques described in this specification. The palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, information regarding the palette table and palette index can be included in the above video / image information and signaled.
[0083] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block according to the prediction mode, or may be located at a certain distance from the current block. In intra prediction, the prediction modes may include one or more non-directional modes and a plurality of directional modes. The intra prediction unit (331) may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0084] The inter prediction unit (332) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) may construct a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-prediction can be performed based on various prediction modes, and information regarding the prediction may include information indicating the inter-prediction mode for the current block.
[0085] The adder (340) can generate a restoration signal (restoration picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter prediction unit (332) and / or the intra prediction unit (331)). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the prediction block can be used as the restoration block.
[0086] The addition unit (340) may be called a restoration unit or a restoration block generation unit. The generated restoration signal may be used for intra-predicting the next block to be processed within the current picture, may be output after filtering as described below, or may be used for inter-predicting the next picture. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0087] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can generate a modified restored picture by applying various filtering methods to the restored picture, and can transmit the modified restored picture to memory (360), specifically to the DPB of memory (360). The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0088] The (modified) restored picture stored in the DPB of the memory (360) can be used as a reference picture in the inter prediction unit (332). The memory (360) can store motion information of blocks from which motion information within the current picture has been derived (or decoded) and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter prediction unit (332) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (360) can store restoration samples of blocks restored within the current picture and transmit them to the intra prediction unit (331).
[0089] In this specification, the embodiments described in the filtering unit (260), inter prediction unit (221), and intra prediction unit (222) of the encoding device (200) may be applied to the filtering unit (350), inter prediction unit (332), and intra prediction unit (331) of the decoding device (300) in the same or corresponding manner.
[0090] Display overlay information (DOI) supplemental enhancement information (SEI) message
[0091] The display overlay information (DOI) SEI message provides metadata that enables the formation of a target display picture formed by overlaying multiple aligned display overlays in a specific order. The display overlays include a texture and optionally an alpha channel, each contained within a cropped decoded picture, a subpicture, and / or a constituent rectangle.
[0092] To use these SEI messages, the following variables are defined. Here, i is the layer identifier of a layer that may exist in the current CVS (coded video sequence).
[0093] - The arrays of picture width and picture height in units of luma samples can be represented as PicWidthInLumaSamples[i] and PicHeightInLumaSamples[i], respectively.
[0094] - The chroma format specifier can be represented by ChromaFormatIdc[i].
[0095] - An array of subpicture counts can be represented as NumSubpics[i].
[0096] - Arrays of widths and heights of subpictures can be denoted as SubPicWidth[i][j] and SubPicHeight[i][j], respectively, where j is the subpicture index of 0,...,NumSubpics[i] - 1.
[0097] Table 1 shows an example of a DOI SEI message.
[0098] display_overlays_info( payloadSize ) {Descriptordoi_idu(6)doi_cancel_flagu(1)if( !doi_cancel_flag ) {doi_persistence_flagu(1)doi_num_display_overlays_minus2ue(v)doi_target_pic_size_present_flagu(1)if( doi_target_pic_size_present_flag )doi_target_pic_width_minus1u(16)doi_target_pic_height_minus1u(16)}doi_nuh_layer_id_present_flagu(1)doi_pic_partition_flagu(1)if( doi_pic_partition_flag ) { doi_partition_type_flagu(1)doi_partition_id_len_minus1u(4)}doi_offset_params_present_flagu(1)if( doi_offset_params_present_flag )doi_offset_param_length_minus1u(4)doi_resampling_enabled_flagu(1)if( doi_resampling_enabled_flag )doi_size_param_length_minus1u(4)for( i = 0; i < doi_num_display_overlays_minus2 + 2;i++ ) { if ( doi_nuh_layer_id_present_flag ) doi_nuh_layer_id[ i ]u(6)if ( doi_pic_partition_flag ) doi_partition_id[ i ]u(v)doi_alpha_present_flag[ i ]u(1)if ( doi_alpha_present_flag[ i ] ) {if (doi_nuh_layer_id_present_flag) doi_alpha_nuh_layer_id[ i ]u(6)if (doi_pic_partition_flag) doi_alpha_partition_id[ i ]u(v)}if( i > 0 ) {if (doi_offset_params_present_flag )doi_top_left_x[ i ]u(v)doi_top_left_y[ i ]u(v)}if( doi_resampling_enabling_flag ) {doi_width_minus1[ i ]u(v)doi_height_minus1[ i ]u(v)}}}}};
[0099] Describe the syntax elements exemplified in Table 1. doi_id specifies the identifier of the DOI SEI message.
[0100] If doi_cancel_flag is 1, it indicates that the SEI message cancels the persistence of the previous DOI SEI message with the same doi_id in the output order. If doi_cancel_flag is 0, it indicates that display overlay information is followed.
[0101] doi_persistence_flag specifies the persistence of DOI SEI messages for CVS. If doi_persistence_flag is 0, it specifies that DOI SEI messages apply only to the current access unit (AU). If doi_persistence_idc is 1, it specifies that DOI SEI messages apply to the current AU and persist to all subsequent AUs in output order until one or more of the following conditions are met.
[0102] - A new CVS begins.
[0103] - The bitstream ends.
[0104] - Pictures of the current AU containing DOI SEI messages with the same doi_id value are output after the current picture in output order.
[0105] The value of doi_num_display_overlays_minus2 plus 2 specifies the number of display overlays for which information is conveyed in SEI messages. The value of doi_num_display_overlays_minus2 ranges from 0 to 30.
[0106] If doi_target_pic_size_present_flag is 1, it indicates that doi_target_pic_width_minus1 and doi_target_pic_width_minus1 syntax elements exist.
[0107] If doi_target_pic_size_present_flag is 0, it indicates that doi_target_pic_width_minus1 and doi_target_pic_width_minus1 syntax elements do not exist.
[0108] If doi_target_pic_width_minus1 exists, the value obtained by adding 1 to doi_target_pic_width_minus1 indicates the width of the target picture.
[0109] If doi_target_pic_height_minus1 exists, the value of doi_target_pic_height_minus1 plus 1 indicates the height of the target picture.
[0110] If doi_nuh_layer_id_present_flag is 1, it indicates that the syntax element doi_nuh_layer_id[i] exists in the SEI message. If doi_nuh_layer_id_present_flag is 0, it indicates that the syntax element doi_nuh_layer_id[i] does not exist in the SEI message.
[0111] If doi_pic_partition_flag is 1, it specifies that the display overlay components are coded as a constituent rectangle or subpicture. If doi_pic_partition_flag is 0, it specifies that the display overlay components are coded as a picture.
[0112] For bitstream suitability, at least one of doi_nuh_layer_id_present_flag or doi_pic_partition_flag is limited to 1.
[0113] If doi_partition_type_flag is 1, it specifies that the display overlay component is coded as a constituent rectangle. If doi_partition_type_flag is 0, it specifies that the display overlay component is coded as a subpicture.
[0114] If doi_partition_type_flag is 1, for bitstream suitability, the current prediction unit (PU) restricts the decoding order to have a constituent rectangle SEI message before the DOI SEI message.
[0115] The value obtained by adding 1 to doi_partition_id_len_minus1 specifies the length of the doi_partition_id[i] syntax element.
[0116] If doi_offset_params_present_flag[i] is 1, it indicates that an offset parameter exists for the i-th display overlay. If doi_offset_params_present_flag[i] is 0, it indicates that an offset parameter does not exist for the i-th display overlay.
[0117] The value obtained by adding 1 to doi_offset_param_length_minus1 specifies the length of the doi_top_left_x[i] and doi_top_left_y[i] syntax elements in bits.
[0118] If doi_resampling_enabled_flag is 1, it specifies that the display overlay component in the target display picture can be resampled. If doi_resampling_enabled_flag is 0, it specifies that the display overlay component in the target display picture is not resampled.
[0119] The value obtained by adding 1 to doi_size_param_length_minus1 specifies the length of the doi_width_minus1[i] and doi_height_minus1[i] syntax elements in bits.
[0120] If doi_nuh_layer_id[i] exists, doi_nuh_layer_id[i] specifies the layer identifier of the texture component of the i-th display overlay. If doi_nuh_layer_id[i] does not exist, the value of doi_nuh_layer_id[i] is inferred to be the same as the layer identifier of the PU containing the DOI SEI message.
[0121] If this SEI message exists in any layer of the current AU, for bitstream conformance, it is restricted that a DOI SEI message with the same doi_id value and the same payload must exist in the layer with layer identifier doi_nuh_layer_id[0].
[0122] If doi_partition_id[i] exists and doi_partition_type_flag is 1, doi_partition_id[i] specifies that the texture component of the i-th display overlay is represented by the j-th constituent rectangle when cr_rect_id[j] is equal to doi_partition_id[i]. If doi_partition_id[i] exists and doi_partition_type_flag is 0, doi_partition_id[i] specifies the subpicture index of the texture component of the i-th display overlay. If doi_partition_id[i] does not exist, the value of doi_partition_id[i] is inferred to be 0.
[0123] If doi_partition_type_flag is 1, the value of doi_partition_id[i] is restricted to the range 0,..., cr_num_rects_minus1[doi_nuh_layer_id[i]] - 1. If doi_partition_type_flag is 0, the value of doi_partition_id[i] is restricted to the range 0,...,NumSubpics[doi_nuh_layer_id[i]] - 1.
[0124] If doi_alpha_present_flag[i] is 1, it specifies that an alpha component is provided for the i-th display overlay. If doi_alpha_present_flag[i] is 0, it specifies that an alpha component is not provided for the i-th display overlay.
[0125] If doi_alpha_nuh_layer_id[i] exists, the layer identifier value of the alpha component of the i-th display overlay is specified. If doi_alpha_nuh_layer_id[i] does not exist, the value of doi_alpha_nuh_layer_id[i] is inferred to be the same as the layer identifier of the PU containing the DOI SEI message.
[0126] If doi_alpha_partition_id[i] exists and doi_partition_type_flag is 1, doi_alpha_partition_id[i] specifies the cr_rect_id[j] of the alpha component of the i-th display overlay. If doi_partition_id[i] exists and doi_partition_type_flag is 0, doi_partition_id[i] specifies the subpicture index of the alpha component of the i-th display overlay. If doi_alpha_partition_id[i] does not exist, the value of doi_alpha_partition_id[i] is inferred to be 0.
[0127] If doi_partition_type_flag is 1, the value of doi_alpha_partition_id[i] is restricted to the range 0,...,cr_num_rects_minus1[doi_nuh_layer_id[i]] - 1. If doi_partition_type_flag is 0, the value of doi_alpha_partition_id[i] is restricted to the range 0,...,NumSubpics[doi_nuh_layer_id[i]] - 1.
[0128] doi_top_left_x[i] and doi_top_left_y[i] specify the horizontal and vertical positions, respectively, of the top-left corner of the i-th display overlay within the target display picture in luminance samples. If doi_top_left_x[i] and doi_top_left_y[i] are not present, their values are inferred as 0. The length of the syntax element is doi_offset_param_length_minus1 + 1 bit.
[0129] If doi_width_minus1[i] and doi_height_minus1[i] exist, the value of doi_width_minus1[i] plus 1 and the value of doi_height_minus1[i] plus 1 specify the width and height of the luminance sample array within the i-th display overlay in the target display picture, respectively. The length of the syntax element is doi_size_param_length_minus1 + 1 bit.
[0130] The variables CodedOverlayTexture[i] and CodedOverlayAlpha[i] are arrays of image samples with a luminance resolution of CodedOverlayWidth[i] × CodedOverlayHeight[i], and are derived as follows:
[0131] 1> If doi_pic_partition_flag is 0, it is applied as follows:
[0132] - CodedOverlayWidth[i] is set to be the same as PicWidthInLumaSamples[doi_nuh_layer_id[i]].
[0133] - CodedOverlayHeight[i] is set to be the same as PicWidthInLumaSamples[doi_nuh_layer_id[i]].
[0134] - If there is a picture in the AU for the layer with layer identifier doi_nuh_layer_id[i], CodedOverlayTexture[i] is set to be the same as the cropped decoded picture from the layer with layer identifier doi_nuh_layer_id[i] in the AU. Otherwise (if there is no picture in the AU for the layer with layer identifier doi_nuh_layer_id[i]), CodedOverlayTexture[i] is set to be the same as the previously cropped decoded picture in the output order within the layer with layer identifier doi_nuh_layer_id[i].
[0135] 1> Otherwise, it applies as follows:
[0136] 2> If doi_alpha_present_flag[i] is 1, it is applied as follows:
[0137] - If a picture exists in the AU for the layer with layer identifier doi_alpha_nuh_layer_id[i], CodedOverlayAlpha[i] is set to be the same as the cropped decoded picture from the layer with layer identifier doi_alpha_nuh_layer_id[i] in the AU.
[0138] - Otherwise (if there is no picture in the AU for the layer with layer identifier doi_nuh_layer_id[i]), CodedOverlayAlpha[i] is set to be the same as the previously cropped decoded picture in the output order within the layer with layer identifier doi_alpha_nuh_layer_id[i].
[0139] 2> Otherwise, if doi_partition_type_flag is 0, it applies as follows:
[0140] - CodedOverlayTexture[i] is set to be the same as the subpicture with subpicture index doi_partition_id[i] in the layer with layer identifier doi_nuh_layer_id[i].
[0141] - CodedOverlayWidth[i] is set to be the same as SubPicWidth[doi_nuh_layer_id[i]][doi_partition_id[i]].
[0142] - CodedOverlayHeight[i] is set to be the same as SubPicHeight[doi_nuh_layer_id[i]][doi_partition_id[i]].
[0143] 2> If doi_alpha_present_flag[i] is 1, it is applied as follows:
[0144] - CodedOverlayAlphaWidth[i] is set to be the same as SubPicWidth[doi_alpha_nuh_layer_id[i]][doi_partition_id[i]].
[0145] - CodedOverlayAlphaHeight[i] is set to be the same as SubPicHeight[doi_alpha_nuh_layer_id[i]][doi_partition_id[i]].
[0146] - CodedOverlayAlpha[i] is set to be the same as the subpicture with subpicture index doi_alpha_partition_id[i] in the layer with layer identifier doi_alpha_nuh_layer_id[i].
[0147] 2> Otherwise (when doi_partition_type_flag is 1), it applies as follows:
[0148] - CodedOverlayWidth[i] is set to be the same as CrRectWidth[doi_nuh_layer_id[i]][doi_partition_id[i]].
[0149] - CodedOverlayHeight[i] is set to be the same as CrRectHeight[doi_nuh_layer_id[i]][doi_partition_id[i]].
[0150] - CodedOverlayTexture[i] is set to be the same as the constituent rectangle with cr_rect_id[j] that is identical to doi_partition_id[i] in the layer with layer identifier doi_nuh_layer_id[i].
[0151] - If doi_alpha_present_flag[i] is 1,
[0152] - CodedOverlayAlphaWidth[i] is set to be the same as CrRectWidth[doi_alpha_nuh_layer_id[i]][doi_partition_id[i]].
[0153] - CodedOverlayAlphaHeight[i] is set to be the same as CrRectHeight[doi_alpha_nuh_layer_id[i]][doi_partition_id[i]].
[0154] - CodedOverlayAlpha is set to be the same as the constituent rectangle with cr_rect_id[j] that is identical to doi_alpha_partition_id[i] in the layer with layer identifier doi_alpha_nuh_layer_id[i].
[0155] The variables DisplayOverlayTexture[i] and DisplayOverlayAlpha[i] are sample arrays with a luminance resolution of DisplayOverlayWidth[i] × DisplayOverlayHeight[i], and are derived as follows:
[0156] 1> If doi_resampling_enabled_flag is 1, it is applied as follows:
[0157] - DisplayOverlayWidth[i] is set to be equal to doi_width_minus1[i] + 1.
[0158] - DisplayOverlayHeight[i] is set to be equal to doi_height_minus1[i] + 1.
[0159] - OverlayTexture[i] is derived by resampling CodedOverlayTexture[i] from a luminance resolution of (CodedOverlayWidth[i] × CodedOverlayHeight[i]) to a luminance resolution of (DisplayOverlayWidth[i] × DisplayOverlayHeight[i]).
[0160] - OverlayAlpha[i] is derived by resampling CodedOverlayAlpha from the luminance resolution of CodedOverlayAlphaWidth[i] × CodedOverlayAlphaHeight[i]] to the luminance resolution of (DisplayOverlayWidth[i] × DisplayOverlayHeight[i]).
[0161] 1> Otherwise (when doi_resampling_enabled_flag is 0), it applies as follows:
[0162] - DisplayOverlayWidth[i] is set to be the same as CodedOverlayWidth[i].
[0163] - DisplayOverlayHeight[i] is set to be the same as CodedOverlayHeight[i].
[0164] - OverlayTexture[i] is set to be the same as CodedOverlayTexture[i].
[0165] - OverlayAlpha[i] is set to be the same as CodedOverlayAlpha[i].
[0166] The TargetPicWidth and TargetPicHeight variables for the target display picture width and height can be derived as follows:
[0167] 1> If doi_target_pic_size_present_flag
[0168] - TargetPicWidth is set to be equal to doi_target_pic_width_minus1 + 1.
[0169] - TargetPicHeight is set to be equal to doi_target_pic_height_minus1 + 1.
[0170] 2> If not
[0171] - TargetPicWidth is set to be the same as DisplayOverlayWidth[0].
[0172] - TargetPicHeight is set to be the same as DisplayOverlayHeight[0].
[0173] If an alpha channel information SEI message exists in CVS or uses a process determined through external means, OverlayWithAlpha(tgtPic[x][y], ovlTex[w][h], ovlAlp[w][h]) is specified as a function that returns sample values derived by applying an alpha channel using tgtPic[x][y] as background samples, ovlTex[w][h] as foreground samples, and ovlAlp[w][h] as alpha channel samples.
[0174] The target display picture is formed as a picture array TargetPicture[cIdx][x][y] as shown in Table 2 below. Here, cIdx = 0..(ChromaFormatIdc = = 0 ) ? 0 : 2, x = 0..( cIdx = = 0 ) ? TargetPicWidth : TargetPicWidth / SubWidthC - 1, and y = 0..( cIdx = = 0 ) ? TargetPicHeight : TargetPicHeight / SubHeightC - 1.
[0175] for( i = 0; i < doi_num_display_overlays_minus2 + 2; i++ ) {for( h = 0, y = doi_top_left_y[ i ]; y < DisplayOverlayHeight[ i ]; h++, y++ )for( w = 0, x = doi_top_left_x[ i ]; x < DisplayOverlayWidth[ i ]; w++, x++ )if( !doi_alpha_present_flag[ i ] )TargetPicture[ 0 ][ x ][ y ] = OverlayTexture[ i ][ 0 ][ w ][ h ]elseTargetPicture[ c ][ x ][ y ] = OverlayWithAlpha( TargetPicture[ 0 ][ x ][ y ], OverlayTexture[ 0 ][ i ][ w ][ h ], OverlayAlpha[ i ][ w ][ h ] )for( ( cIdx = 1; cIdx < ChromaFormatIdc = = 0 ) ? 1 : 3; cIdx++ ++ ) {for( h = 0, y = doi_top_left_y[ i ] / SubHeightC; y < DisplayOverlayHeight[ i ] / SubHeightC; h++, y++ )for( w = 0, x = doi_top_left_x[ i ] / SubWidthC; x < DisplayOverlayWidth[ i ] / SubWidthC; w++, x++ )if( !doi_alpha_present_flag[ i ] )TargetPicture[ cIdx ][ x ][ y ] = OverlayTexture[ i ][ cIdx ][ w ][ h ]elseTargetPicture[ cIdx ][ x ][ y ] = OverlayWithAlpha( TargetPicture[ cIdx ][ x ][ y ],OverlayTexture[ cIdx ][ i ][ w ][ h ], OverlayAlpha[ i ][ w ][ h ] )
[0176] Method for providing a single display overlay within a DOI SEI message
[0177] As mentioned above, display overlay information (DOI) SEI messages (i.e., DOI SEI messages) are currently being considered for future extensions of VSEI (JVET-AI2032) and are included in TuC (Technologies under consideration).
[0178] As mentioned above, the number of existing display overlays is specified by doi_num_display_overlays_minus2. Therefore, a minimum number of display overlays of 2 is required.
[0179] However, when there is only one item (overlay) to be displayed, according to the existing DOI SEI message design, there is an inefficient problem in that two overlays must be defined, with the other one representing an empty area.
[0180] Additionally, for example, the size of the displayed picture can be expanded, and the expanded area can be left empty for future use. However, according to the existing DOI SEI message design, the size information of the target picture is defined. Therefore, although such an empty area can be identified with only a single display overlay, the current DOI SEI message design cannot represent a single overlay.
[0181] In addition, since the coded picture size of the display overlay with index 0 may differ from the target picture size, offset parameters and resampling information for the display overlay need to be signaled.
[0182] However, as mentioned above, the existing DOI SEI message design has a problem in that it can signal offset parameters only when the display overlay index is greater than 0. In addition, the current DOI SEI message design has a problem in that it can signal resampling information only when the display overlay index is greater than 0.
[0183] Accordingly, the present disclosure proposes a method to solve the above-mentioned problem.
[0184] First, the present disclosure proposes a method to change doi_num_display_overlays_minus2, a syntax element for the number of display overlays, to doi_num_display_overlays_minus1.
[0185] In addition, the present disclosure proposes a method for signaling regardless of the number of display overlays when the display overlay information (DOI) SEI message contains offset parameters and resampling information.
[0186] The embodiments of the present disclosure may each be applied individually, or two or more embodiments may be applied in combination.
[0187] FIG. 4 illustrates a method for restoring a video picture performed in a decoding device (300) according to the present disclosure.
[0188] A bitstream containing an encoded video picture can be received (S400).
[0189] The encoded video picture of the bitstream can be restored (S410).
[0190] Video information regarding an encoded video picture can be extracted from the bitstream. Based on the extracted video information, the encoded video picture can be restored.
[0191] Example 1
[0192] This embodiment describes a proposed method based on DOI SEI message design, and the description of this embodiment is based on VSEI and VVC specifications, but this is for convenience of explanation and the present disclosure is not limited thereto.
[0193] According to the present embodiment, the bitstream may include display overlay information (DOI). The DOI may refer to information that enables the formation of a target display picture formed by overlaying the display overlay.
[0194] Additionally, according to the present embodiment, the display overlay information may include information regarding the number of display overlays (doi_num_display_overlays_minus1). Based on the value of the information regarding the number of display overlays, the number of display overlays may be indicated as 1 or more. In other words, a value obtained by adding 1 to the value of the information regarding the number of display overlays may indicate the number of display overlays. The value of the information regarding the number of display overlays may have a value within the range of 0 to 31.
[0195] Hereinafter, the DOI SEI message according to the present disclosure is described in detail.
[0196] Table 3 illustrates a DOI SEI message according to the present disclosure.
[0197] display_overlays_info( payloadSize ) {Descriptordoi_idu(6)doi_cancel_flagu(1)if( !doi_cancel_flag ) {doi_persistence_flagu(1)doi_num_display_overlays_minus1ue(v)doi_target_pic_size_present_flagu(1)if( doi_target_pic_size_present_flag )doi_target_pic_width_minus1u(16)doi_target_pic_height_minus1u(16)}doi_nuh_layer_id_present_flagu(1)doi_pic_partition_flagu(1)if( doi_pic_partition_flag ) { doi_partition_type_flagu(1)doi_partition_id_len_minus1u(4)}doi_offset_params_present_flagu(1)if( doi_offset_params_present_flag )doi_offset_param_length_minus1u(4)doi_resampling_enabled_flagu(1)if( doi_resampling_enabled_flag )doi_size_param_length_minus1u(4)for( i = 0; i < doi_num_display_overlays_minus1 + 1;i++ ) { if ( doi_nuh_layer_id_present_flag ) doi_nuh_layer_id[ i ]u(6)if ( doi_pic_partition_flag ) doi_partition_id[ i ]u(v)doi_alpha_present_flag[ i ]u(1)if ( doi_alpha_present_flag[ i ] ) {if (doi_nuh_layer_id_present_flag) doi_alpha_nuh_layer_id[ i ]u(6)if (doi_pic_partition_flag) doi_alpha_partition_id[ i ]u(v)}if( i > 0 ) {if (doi_offset_params_present_flag )doi_top_left_x[ i ]u(v)doi_top_left_y[ i ]u(v)}if( doi_resampling_enabling_flag ) {doi_width_minus1[ i ]u(v)doi_height_minus1[ i ]u(v)}}}}};
[0198] A display overlay information (DOI) SEI message may provide metadata that enables the formation of a target display picture formed by overlaying multiple aligned display overlays in a specific order. The display overlays include a texture and optionally an alpha channel, each of which may be contained within a cropped decoded picture, a subpicture, and / or a constituent rectangle.
[0199] To use these SEI messages, the following variables can be defined. Here, i is the layer identifier of a layer that may exist in the current CVS (coded video sequence).
[0200] - The arrays of picture width and picture height in units of luma samples can be represented as PicWidthInLumaSamples[i] and PicHeightInLumaSamples[i], respectively.
[0201] - The chroma format specifier can be represented by ChromaFormatIdc[i].
[0202] - An array of subpicture counts can be represented as NumSubpics[i].
[0203] - Arrays of widths and heights of subpictures can be denoted as SubPicWidth[i][j] and SubPicHeight[i][j], respectively, where j is the subpicture index of 0,...,NumSubpics[i] - 1.
[0204] doi_id can identify the identifier of a DOI SEI message.
[0205] If doi_cancel_flag is 1, it may indicate that the persistence of previous DOI SEI messages with the same doi_id in the output order of the SEI message is canceled. If doi_cancel_flag is 0, it may indicate that display overlay information is followed.
[0206] doi_persistence_flag can specify the persistence of DOI SEI messages for CVS. If doi_persistence_flag is 0, it specifies that DOI SEI messages apply only to the current access unit (AU). If doi_persistence_idc is 1, it specifies that DOI SEI messages apply to the current AU and persist to all subsequent AUs in output order until one or more of the following conditions are met.
[0207] - A new CVS begins.
[0208] - The bitstream ends.
[0209] - Pictures of the current AU containing DOI SEI messages with the same doi_id value are output after the current picture in output order.
[0210] The value of doi_num_display_overlays_minus1 plus 1 can specify the number of display overlays to which information is conveyed in SEI messages. The value of doi_num_display_overlays_minus1 can be limited to a range from 0 to 31.
[0211] If doi_target_pic_size_present_flag is 1, it can be specified that doi_target_pic_width_minus1 and doi_target_pic_width_minus1 syntax elements exist.
[0212] If doi_target_pic_size_present_flag is 0, it can be specified that doi_target_pic_width_minus1 and doi_target_pic_width_minus1 syntax elements do not exist.
[0213] If doi_target_pic_width_minus1 exists, the value of doi_target_pic_width_minus1 plus 1 can indicate the width of the target picture.
[0214] If doi_target_pic_height_minus1 exists, the value of doi_target_pic_height_minus1 plus 1 can indicate the height of the target picture.
[0215] If doi_nuh_layer_id_present_flag is 1, it can be determined that the syntax element doi_nuh_layer_id[i] exists in the SEI message. If doi_nuh_layer_id_present_flag is 0, it can be determined that the syntax element doi_nuh_layer_id[i] does not exist in the SEI message.
[0216] If doi_pic_partition_flag is 1, it can be specified that the display overlay components are coded as a constituent rectangle or a subpicture. If doi_pic_partition_flag is 0, it can be specified that the display overlay components are coded as a picture.
[0217] For bitstream suitability, at least one of doi_nuh_layer_id_present_flag or doi_pic_partition_flag may be limited to 1.
[0218] If doi_partition_type_flag is 1, it can be specified that the display overlay component is coded as a constituent rectangle. If doi_partition_type_flag is 0, it can be specified that the display overlay component is coded as a subpicture.
[0219] If doi_partition_type_flag is 1, for bitstream suitability, the current prediction unit (PU) may restrict the presence of a constituent rectangle SEI message before a DOI SEI message in the decoding order.
[0220] The value obtained by adding 1 to doi_partition_id_len_minus1 can specify the length of the doi_partition_id[i] syntax element.
[0221] If doi_offset_params_present_flag[i] is 1, it can be determined that an offset parameter exists for the i-th display overlay. If doi_offset_params_present_flag[i] is 0, it can be determined that an offset parameter does not exist for the i-th display overlay.
[0222] The value obtained by adding 1 to doi_offset_param_length_minus1 can specify the length of the doi_top_left_x[i] and doi_top_left_y[i] syntax elements in bits.
[0223] If doi_resampling_enabled_flag is 1, it specifies that the display overlay component in the target display picture can be resampled. If doi_resampling_enabled_flag is 0, it specifies that the display overlay component in the target display picture is not resampled.
[0224] The value obtained by adding 1 to doi_size_param_length_minus1 can specify the length of the doi_width_minus1[i] and doi_height_minus1[i] syntax elements in bits.
[0225] If doi_nuh_layer_id[i] exists, doi_nuh_layer_id[i] can identify the layer identifier of the texture component of the i-th display overlay. If doi_nuh_layer_id[i] does not exist, the value of doi_nuh_layer_id[i] can be inferred to be the same as the layer identifier of the PU containing the DOI SEI message.
[0226] If this SEI message exists in any layer of the current AU, for bitstream conformance, it may be restricted that a DOI SEI message with the same doi_id value and the same payload exists in the layer with layer identifier doi_nuh_layer_id[0].
[0227] If doi_partition_id[i] exists and doi_partition_type_flag is 1, doi_partition_id[i] can identify that the texture component of the i-th display overlay is represented by the j-th constituent rectangle when cr_rect_id[j] is equal to doi_partition_id[i]. If doi_partition_id[i] exists and doi_partition_type_flag is 0, doi_partition_id[i] can identify the subpicture index of the texture component of the i-th display overlay. If doi_partition_id[i] does not exist, the value of doi_partition_id[i] can be inferred as 0.
[0228] If doi_partition_type_flag is 1, the value of doi_partition_id[i] may be restricted to the range 0,..., cr_num_rects_minus1[doi_nuh_layer_id[i]] - 1. If doi_partition_type_flag is 0, the value of doi_partition_id[i] may be restricted to the range 0,...,NumSubpics[doi_nuh_layer_id[i]] - 1.
[0229] If doi_alpha_present_flag[i] is 1, it can be specified that an alpha component is provided for the i-th display overlay. If doi_alpha_present_flag[i] is 0, it can be specified that an alpha component is not provided for the i-th display overlay.
[0230] If doi_alpha_nuh_layer_id[i] exists, the layer identifier value of the alpha component of the i-th display overlay can be identified. If doi_alpha_nuh_layer_id[i] does not exist, the value of doi_alpha_nuh_layer_id[i] can be inferred to be the same as the layer identifier of the PU containing the DOI SEI message.
[0231] If doi_alpha_partition_id[i] exists and doi_partition_type_flag is 1, doi_alpha_partition_id[i] can identify cr_rect_id[j] of the alpha component of the i-th display overlay. If doi_partition_id[i] exists and doi_partition_type_flag is 0, doi_partition_id[i] can identify the subpicture index of the alpha component of the i-th display overlay. If doi_alpha_partition_id[i] does not exist, the value of doi_alpha_partition_id[i] can be inferred as 0.
[0232] If doi_partition_type_flag is 1, the value of doi_alpha_partition_id[i] may be restricted to the range 0,...,cr_num_rects_minus1[doi_nuh_layer_id[i]] - 1. If doi_partition_type_flag is 0, the value of doi_alpha_partition_id[i] may be restricted to the range 0,...,NumSubpics[doi_nuh_layer_id[i]] - 1.
[0233] doi_top_left_x[i] and doi_top_left_y[i] can specify the horizontal and vertical positions, respectively, of the top-left corner of the i-th display overlay within the target display picture in luminance samples. If doi_top_left_x[i] and doi_top_left_y[i] are missing, their values can be inferred as 0. The length of the syntax element can be doi_offset_param_length_minus1 + 1 bit.
[0234] If doi_width_minus1[i] and doi_height_minus1[i] exist, the value of doi_width_minus1[i] plus 1 and the value of doi_height_minus1[i] plus 1 can specify the width and height of the luminance sample array within the i-th display overlay in the target display picture, respectively. The length of the syntax element can be doi_size_param_length_minus1 + 1 bit.
[0235] The variables CodedOverlayTexture[i] and CodedOverlayAlpha[i] are arrays of image samples with a luminance resolution of CodedOverlayWidth[i] × CodedOverlayHeight[i], and can be derived as follows:
[0236] 1> If doi_pic_partition_flag is 0, it is applied as follows:
[0237] - CodedOverlayWidth[i] is set to be the same as PicWidthInLumaSamples[doi_nuh_layer_id[i]].
[0238] - CodedOverlayHeight[i] is set to be the same as PicWidthInLumaSamples[doi_nuh_layer_id[i]].
[0239] - If there is a picture in the AU for the layer with layer identifier doi_nuh_layer_id[i], CodedOverlayTexture[i] is set to be the same as the cropped decoded picture from the layer with layer identifier doi_nuh_layer_id[i] in the AU. Otherwise (if there is no picture in the AU for the layer with layer identifier doi_nuh_layer_id[i]), CodedOverlayTexture[i] is set to be the same as the previously cropped decoded picture in the output order within the layer with layer identifier doi_nuh_layer_id[i].
[0240] 1> Otherwise, it applies as follows:
[0241] 2> If doi_alpha_present_flag[i] is 1, it is applied as follows:
[0242] - If a picture exists in the AU for the layer with layer identifier doi_alpha_nuh_layer_id[i], CodedOverlayAlpha[i] is set to be the same as the cropped decoded picture from the layer with layer identifier doi_alpha_nuh_layer_id[i] in the AU.
[0243] - Otherwise (if there is no picture in the AU for the layer with layer identifier doi_nuh_layer_id[i]), CodedOverlayAlpha[i] is set to be the same as the previously cropped decoded picture in the output order within the layer with layer identifier doi_alpha_nuh_layer_id[i].
[0244] 2> Otherwise, if doi_partition_type_flag is 0, it applies as follows:
[0245] - CodedOverlayTexture[i] is set to be the same as the subpicture with subpicture index doi_partition_id[i] in the layer with layer identifier doi_nuh_layer_id[i].
[0246] - CodedOverlayWidth[i] is set to be the same as SubPicWidth[doi_nuh_layer_id[i]][doi_partition_id[i]].
[0247] - CodedOverlayHeight[i] is set to be the same as SubPicHeight[doi_nuh_layer_id[i]][doi_partition_id[i]].
[0248] 2> If doi_alpha_present_flag[i] is 1, it is applied as follows:
[0249] - CodedOverlayAlphaWidth[i] is set to be the same as SubPicWidth[doi_alpha_nuh_layer_id[i]][doi_partition_id[i]].
[0250] - CodedOverlayAlphaHeight[i] is set to be the same as SubPicHeight[doi_alpha_nuh_layer_id[i]][doi_partition_id[i]].
[0251] - CodedOverlayAlpha[i] is set to be the same as the subpicture with subpicture index doi_alpha_partition_id[i] in the layer with layer identifier doi_alpha_nuh_layer_id[i].
[0252] 2> Otherwise (when doi_partition_type_flag is 1), it applies as follows:
[0253] - CodedOverlayWidth[i] is set to be the same as CrRectWidth[doi_nuh_layer_id[i]][doi_partition_id[i]].
[0254] - CodedOverlayHeight[i] is set to be the same as CrRectHeight[doi_nuh_layer_id[i]][doi_partition_id[i]].
[0255] - CodedOverlayTexture[i] is set to be the same as the constituent rectangle with cr_rect_id[j] that is identical to doi_partition_id[i] in the layer with layer identifier doi_nuh_layer_id[i].
[0256] - If doi_alpha_present_flag[i] is 1,
[0257] - CodedOverlayAlphaWidth[i] is set to be the same as CrRectWidth[doi_alpha_nuh_layer_id[i]][doi_partition_id[i]].
[0258] - CodedOverlayAlphaHeight[i] is set to be the same as CrRectHeight[doi_alpha_nuh_layer_id[i]][doi_partition_id[i]].
[0259] - CodedOverlayAlpha is set to be the same as the constituent rectangle with cr_rect_id[j] that is identical to doi_alpha_partition_id[i] in the layer with layer identifier doi_alpha_nuh_layer_id[i].
[0260] The variables DisplayOverlayTexture[i] and DisplayOverlayAlpha[i] are sample arrays with a luminance resolution of DisplayOverlayWidth[i] × DisplayOverlayHeight[i], and can be derived as follows:
[0261] 1> If doi_resampling_enabled_flag is 1, it is applied as follows:
[0262] - DisplayOverlayWidth[i] is set to be equal to doi_width_minus1[i] + 1.
[0263] - DisplayOverlayHeight[i] is set to be equal to doi_height_minus1[i] + 1.
[0264] - OverlayTexture[i] is derived by resampling CodedOverlayTexture[i] from a luminance resolution of (CodedOverlayWidth[i] × CodedOverlayHeight[i]) to a luminance resolution of (DisplayOverlayWidth[i] × DisplayOverlayHeight[i]).
[0265] - OverlayAlpha[i] is derived by resampling CodedOverlayAlpha from the luminance resolution of CodedOverlayAlphaWidth[i] × CodedOverlayAlphaHeight[i]] to the luminance resolution of (DisplayOverlayWidth[i] × DisplayOverlayHeight[i]).
[0266] 1> Otherwise (when doi_resampling_enabled_flag is 0), it applies as follows:
[0267] - DisplayOverlayWidth[i] is set to be the same as CodedOverlayWidth[i].
[0268] - DisplayOverlayHeight[i] is set to be the same as CodedOverlayHeight[i].
[0269] - OverlayTexture[i] is set to be the same as CodedOverlayTexture[i].
[0270] - OverlayAlpha[i] is set to be the same as CodedOverlayAlpha[i].
[0271] The TargetPicWidth and TargetPicHeight variables for the target display picture width and height can be derived as follows:
[0272] 1> If doi_target_pic_size_present_flag
[0273] - TargetPicWidth is set to be equal to doi_target_pic_width_minus1 + 1.
[0274] - TargetPicHeight is set to be equal to doi_target_pic_height_minus1 + 1.
[0275] 2> If not
[0276] - TargetPicWidth is set to be the same as DisplayOverlayWidth[0].
[0277] - TargetPicHeight is set to be the same as DisplayOverlayHeight[0].
[0278] If an alpha channel information SEI message exists in CVS or uses a process determined through external means, OverlayWithAlpha(tgtPic[x][y], ovlTex[w][h], ovlAlp[w][h]) can be specified as a function that returns sample values derived by applying an alpha channel using tgtPic[x][y] as background samples, ovlTex[w][h] as foreground samples, and ovlAlp[w][h] as alpha channel samples.
[0279] The target display picture can be formed as a picture array TargetPicture[cIdx][x][y] as shown in Table 4 below. Here, cIdx = 0..(ChromaFormatIdc = = 0 ) ? 0 : 2, x = 0..( cIdx = = 0 ) ? TargetPicWidth : TargetPicWidth / SubWidthC - 1, and y = 0..( cIdx = = 0 ) ? TargetPicHeight : TargetPicHeight / SubHeightC - 1.
[0280] for( i = 0; i < doi_num_display_overlays_minus1 + 1; i++ ) {for( h = 0, y = doi_top_left_y[ i ]; y < DisplayOverlayHeight[ i ]; h++, y++ )for( w = 0, x = doi_top_left_x[ i ]; x < DisplayOverlayWidth[ i ]; w++, x++ )if( !doi_alpha_present_flag[ i ] )TargetPicture[ 0 ][ x ][ y ] = OverlayTexture[ i ][ 0 ][ w ][ h ]elseTargetPicture[ c ][ x ][ y ] = OverlayWithAlpha( TargetPicture[ 0 ][ x ][ y ], OverlayTexture[ 0 ][ i ][ w ][ h ], OverlayAlpha[ i ][ w ][ h ] )for( ( cIdx = 1; cIdx < ChromaFormatIdc = = 0 ) ? 1 : 3; cIdx++ ++ ) {for( h = 0, y = doi_top_left_y[ i ] / SubHeightC; y < DisplayOverlayHeight[ i ] / SubHeightC; h++, y++ )for( w = 0, x = doi_top_left_x[ i ] / SubWidthC; x < DisplayOverlayWidth[ i ] / SubWidthC; w++, x++ )if( !doi_alpha_present_flag[ i ] )TargetPicture[ cIdx ][ x ][ y ] = OverlayTexture[ i ][ cIdx ][ w ][ h ]elseTargetPicture[ cIdx ][ x ][ y ] = OverlayWithAlpha( TargetPicture[ cIdx ][ x ][ y ],OverlayTexture[ cIdx ][ i ][ w ][ h ], OverlayAlpha[ i ][ w ][ h ] )
[0281] Referring to Table 4, when the doi_num_display_overlays_minus1 syntax element is used, a picture array of the target display picture can be formed so that one display overlay is displayed in the target display picture even when the number of display overlays is one.
[0282] Information regarding the number of display overlays may be configured in a supplemental enhancement information (SEI) message of the bitstream. The SEI message may be included in a network abstraction layer (NAL) unit of the bitstream. Alternatively, information regarding the number of display overlays according to the present disclosure may be configured in a high-level syntax of the bitstream. Here, the high-level syntax may be at least one of a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), or a slice header (SH). Alternatively, information regarding the number of display overlays according to the present disclosure may be defined as a separate NAL unit type within the bitstream.
[0283] Accordingly, according to an embodiment of the present disclosure, one display overlay within a display picture can be represented based on information regarding the number of display overlays. Additionally, even when the size of the displayed picture is expanded and the expanded area is left empty for future use, one display overlay may be used to identify the empty space.
[0284] Example 2
[0285] This embodiment describes a proposed method based on the DOI SEI message design described above. The description of this embodiment is based on VSEI and VVC specifications, but this is for convenience of explanation and the present disclosure is not limited thereto.
[0286] According to the present embodiment, the bitstream may include display overlay information (DOI).
[0287] Additionally, according to the present embodiment, display overlay information may include offset parameters (doi_top_left_x[ i ], doi_top_left_y[ i ]) that specify the horizontal and vertical positions for the top-left sample of at least one display overlay, regardless of the number of display overlays (i.e., when the number of display overlays is 1 or more). Here, the bitstream includes an offset parameter present flag (doi_offset_params_present_flag) that indicates whether the offset parameters exist, and the offset parameters may be obtained from the bitstream based on the value of the offset parameter present flag.
[0288] Additionally, according to the present embodiment, the display overlay information may include resampling information (doi_width_minus1[i], doi_height_minus1[i]) that specifies the width and height of a sample array of at least one display overlay regardless of the number of display overlays (i.e., when the number of display overlays is 1 or more). Here, the bitstream includes a resampling enabled flag (doi_resampling_enabling_flag) that indicates whether the display overlay component is capable of resampling, and the resampling information may be obtained from the bitstream based on the value of the offset parameter flag.
[0289] Hereinafter, the DOI SEI message according to the present disclosure is described in detail.
[0290] Table 5 illustrates a DOI SEI message according to the present disclosure.
[0291] display_overlays_info( payloadSize ) {Descriptordoi_idu(6)doi_cancel_flagu(1)if( !doi_cancel_flag ) {doi_persistence_flagu(1)doi_num_display_overlays_minus1ue(v)doi_target_pic_size_present_flagu(1)if( doi_target_pic_size_present_flag )doi_target_pic_width_minus1u(16)doi_target_pic_height_minus1u(16)}doi_nuh_layer_id_present_flagu(1)doi_pic_partition_flagu(1)if( doi_pic_partition_flag ) { doi_partition_type_flagu(1)doi_partition_id_len_minus1u(4)}doi_offset_params_present_flagu(1)if( doi_offset_params_present_flag )doi_offset_param_length_minus1u(4)doi_resampling_enabled_flagu(1)if( doi_resampling_enabled_flag )doi_size_param_length_minus1u(4)for( i = 0; i < doi_num_display_overlays_minus1 + 1;i++ ) { if ( doi_nuh_layer_id_present_flag ) doi_nuh_layer_id[ i ]u(6)if ( doi_pic_partition_flag ) doi_partition_id[ i ]u(v)doi_alpha_present_flag[ i ]u(1)if ( doi_alpha_present_flag[ i ] ) {if (doi_nuh_layer_id_present_flag) doi_alpha_nuh_layer_id[ i ]u(6)if (doi_pic_partition_flag) doi_alpha_partition_id[ i ]u(v)}if (doi_offset_params_present_flag) {doi_top_left_x[ i ]u(v)doi_top_left_y[ i ]u(v)}if(doi_resampling_enabling_flag) {doi_width_minus1[ i ]u(v)doi_height_minus1[ i ]u(v)}}}};
[0292] The description of Example 2 explains only the parts that differ from Example 1 above, and redundant descriptions regarding the same explanation as Example 1 are omitted.
[0293] Referring again to the existing DOI SEI message according to Table 1, the offset parameters doi_top_left_x[i] and doi_top_left_y[i] for the i-th display overlay are signaled only when the display overlay index is greater than 1 ('if(i>0)') (i.e., when the number of display overlays is 2 or more). Also, the resampling information doi_width_minus1[i] and doi_height_minus1[i] for the i-th display overlay is signaled only when the display overlay index is greater than 1 ('if(i>0)').
[0294] On the other hand, according to Table 5 of Example 2, regardless of the number of display overlays, the offset parameters doi_top_left_x[i] and doi_top_left_y[i] for the i-th display overlay can be signaled based on the value of doi_offset_params_present_flag. Additionally, regardless of the number of display overlays, the resampling information doi_width_minus1[i] and doi_height_minus1[i] for the i-th display overlay can be signaled.
[0295] Offset parameters for a display overlay and / or resampling information for a display overlay may be configured in a supplemental enhancement information (SEI) message of the bitstream. The SEI message may be included in a network abstraction layer (NAL) unit of the bitstream. Alternatively, offset parameters for a display overlay and / or resampling information for a display overlay according to the present disclosure may be configured in a high-level syntax of the bitstream. Here, the high-level syntax may be at least one of a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), or a slice header (SH). Alternatively, offset parameters for a display overlay and / or resampling information for a display overlay according to the present disclosure may be defined as a separate NAL unit type within the bitstream.
[0296] Accordingly, according to an embodiment of the present disclosure, if the display overlay information (DOI) SEI message contains offset parameters and resampling information, it can be signaled regardless of the number of display overlays.
[0297] FIG. 5 illustrates a schematic configuration of a decoding device (300) that performs a method for restoring a video picture according to the present disclosure.
[0298] Referring to FIG. 5, the decoding device (300) may include a receiving unit (500), a video information extraction unit (510), and a video restoration unit (520).
[0299] The receiver (500) can receive a bitstream including an encoded video picture.
[0300] The video information extraction unit (510) can extract video information regarding an encoded video picture from the bitstream. Additionally, the video information extraction unit (510) can extract at least one of display overlay information and / or information regarding the number of display overlays and / or offset parameters for the display overlay and / or resampling information for the display overlay from the bitstream, as seen with reference to FIG. 4.
[0301] The video restoration unit (520) can restore an encoded video picture based on extracted video information.
[0302] FIG. 6 illustrates a method for generating a bitstream performed in an encoding device (200) according to the present disclosure.
[0303] A video picture being encoded can be received (S600).
[0304] Video information regarding the video picture can be generated by encoding the received video picture (S610).
[0305] A bitstream containing video information about a video picture can be generated (S620).
[0306] Additionally, at least one of display overlay information applied to the bitstream and / or information regarding the number of display overlays and / or offset parameters for the display overlays and / or resampling information for the display overlays can be generated, as seen with reference to FIG. 4. At least one of the generated display overlay information and / or information regarding the number of display overlays and / or offset parameters for the display overlays and / or resampling information for the display overlays can be included in the bitstream.
[0307] FIG. 7 illustrates a schematic configuration of an encoding device (200) that performs a method for generating a bitstream according to the present disclosure.
[0308] Referring to FIG. 7, the encoding device (200) may include a receiving unit (700), a video compression unit (710), and a bitstream generation unit (720).
[0309] The receiver (700) can receive one or more video pictures that are encoded.
[0310] The video compression unit (710) can generate video information regarding the video picture by encoding one or more received video pictures. The video compression unit (710) can generate at least one of display overlay information applied to the bitstream and / or information about the number of display overlays and / or offset parameters for the display overlays and / or resampling information for the display overlays, as seen with reference to FIG. 4.
[0311] The bitstream generation unit (720) can generate a bitstream including the video information. The bitstream generation unit (720) can generate a bitstream that further includes at least one of the generated display overlay information and / or information on the number of display overlays and / or offset parameters for the display overlay and / or resampling information for the display overlay.
[0312] In the embodiments described above, methods are described based on flowcharts as a series of steps or blocks; however, the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps as described above. Furthermore, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and other steps may be included, or one or more steps of the flowcharts may be omitted without affecting the scope of the embodiments of this document.
[0313] The method according to the embodiments of the present document described above may be implemented in the form of software, and the encoding device and / or decoding device according to the present document may be included in a device that performs image processing, such as a TV, computer, smartphone, set-top box, display device, etc.
[0314] When the embodiments described in this document are implemented in software, the method described above may be implemented as a module (process, function, etc.) that performs the function described above. The module may be stored in memory and executed by a processor. The memory may be located inside or outside the processor and may be connected to the processor by various well-known means. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this document may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each figure may be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for implementation may be stored on a digital storage medium.
[0315] In addition, the decoding device and encoding device to which the embodiment(s) of the present specification are applied may be included in multimedia broadcasting transmission and reception devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video conversation devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, Video on Demand (VoD) service providers, Over-the-top video (OTT) devices, internet streaming service providers, 3D video devices, virtual reality (VR) devices, augmented reality (AR) devices, video phone video devices, transportation terminals (e.g., vehicle terminals (including autonomous vehicles), airplane terminals, ship terminals, etc.), and medical video devices, and may be used to process video signals or data signals. For example, Over-the-top video (OTT) devices may include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, Digital Video Recorders (DVRs), etc.
[0316] Additionally, the processing method to which the embodiment(s) of this specification are applied may be produced in the form of a program that is executed by a computer and may be stored on a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of this specification may also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium may include, for example, a Blu-ray disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. Additionally, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission over the Internet). Additionally, a bitstream generated by an encoding method may be stored on a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0317] Additionally, the embodiments of this specification may be implemented as a computer program product by program code, and said program code may be executed on a computer by the embodiments of this specification. said program code may be stored on a carrier readable by a computer.
[0318] FIG. 8 shows an example of a content streaming system to which embodiments of the present disclosure can be applied.
[0319] Referring to FIG. 8, a content streaming system to which the embodiment(s) of the present specification are applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0320] The above encoding server compresses content input from multimedia input devices, such as smartphones, cameras, and camcorders, into digital data to generate a bitstream and transmits it to the streaming server. As another example, if multimedia input devices, such as smartphones, cameras, and camcorders, generate the bitstream directly, the encoding server may be omitted.
[0321] The bitstream above may be generated by an encoding method or a bitstream generation method to which the embodiment(s) of the present specification are applied, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0322] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as a medium to inform the user of available services. When a user requests a desired service from the web server, the web server transmits it to the streaming server, and the streaming server transmits the multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server plays the role of controlling commands and responses between each device within the content streaming system.
[0323] The streaming server may receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content may be received in real time. In this case, to provide a seamless streaming service, the streaming server may store the bitstream for a certain period of time.
[0324] Examples of the above user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs (head-mounted displays)), digital TVs, desktop computers, digital signage, etc.
[0325] Each server within the above-mentioned content streaming system can be operated as a distributed server, and in this case, data received from each server can be processed in a distributed manner.
[0326] The claims described in this specification may be combined in various ways. For example, the technical features of the method claims in this specification may be combined to be implemented as a device, and the technical features of the device claims in this specification may be combined to be implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims in this specification may be combined to be implemented as a device, and the technical features of the method claims and the technical features of the device claims in this specification may be combined to be implemented as a method.
Claims
1. A step of receiving a bitstream including an encoded video picture; and The method includes the step of restoring an encoded video picture contained in the bitstream, The above bitstream includes information about the number of display overlays, and Based on the value of the information regarding the number of the above display overlays, the number of display overlays is specified to be 1 or more, and A method in which information regarding the number of display overlays is obtained from the NAL (network abstraction layer) unit of the bitstream.
2. In Paragraph 1, A method for specifying the number of display overlays by adding 1 to the value of the information regarding the number of display overlays above.
3. In Paragraph 1, A method in which the value of information regarding the number of the display overlays has a value within the range of 0 to 31.
4. In Paragraph 1, A method based on the number of display overlays being one or more, wherein the bitstream includes offset parameters that specify horizontal and vertical positions for the top-left sample of at least one display overlay.
5. In Paragraph 4, The bitstream above includes an offset parameter existence flag indicating whether the offset parameter exists, and A method in which the offset parameter is signaled from the bitstream based on the value of the offset parameter existence flag.
6. In Paragraph 1, A method in which, based on the number of display overlays being 1 or more, the bitstream includes resampling information that specifies the width and height of a sample array of at least one display overlay.
7. In Paragraph 6, The above bitstream includes a resampling enabled flag indicating whether the display overlay component is capable of resampling, and A method in which resampling information is signaled from the bitstream based on the value of the offset parameter flag.
8. A step of receiving a video picture to be encoded; A step of encoding the received video picture to generate video information regarding the video picture; A step of generating information about the number of display overlays; and The method includes the step of generating a bitstream including the video information and information regarding the number of display overlays, Based on the value of the information regarding the number of the above display overlays, the number of display overlays is specified to be 1 or more, and A method in which information regarding the number of display overlays is encoded in the NAL (network abstraction layer) unit of the bitstream.
9. A computer-readable storage medium for storing a bitstream generated by the method according to paragraph 8.
10. A step of generating a bitstream, wherein the bitstream is generated based on receiving a video picture to be encoded, encoding the received video picture to generate video information regarding the video picture, and generating information regarding the number of display overlays; and The method includes the step of transmitting data including the above bitstream, Based on the value of the information regarding the number of the above display overlays, the number of display overlays is specified to be 1 or more, and A method in which information regarding the number of display overlays is encoded in the NAL (network abstraction layer) unit of the bitstream.