Method and computer-readable storage medium
Patent Information
- Application Number
- PCT/KR2026/095132
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-17
- Filing Date
- 2026-03-17
- Publication Date
- 2026-09-24
Smart Images

Figure KR2026095132_24092026_PF_FP_ABST
Abstract
Description
Method and computer-readable storage medium
[0001] The present disclosure relates to a method for decoding / encoding image information, a computer-readable storage medium for storing image information, and a method for transmitting image information.
[0002] Recently, the demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition), has been increasing across various fields. As video data becomes higher in resolution and quality, the relative amount of information or bits transmitted increases compared to conventional video data. This increase in transmitted information or bits leads to higher transmission and storage costs.
[0003] Accordingly, high-efficiency video compression technology is required to effectively transmit, store, and play back high-resolution, high-quality video information.
[0004] The present disclosure aims to provide an encoding / decoding method and / or apparatus with improved coding efficiency.
[0005] The present disclosure aims to provide an encoding / decoding method and / or apparatus having data transmission efficiency.
[0006] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure belongs from the description below.
[0007] A method according to one aspect comprises: a step of obtaining a Packed Regions Information (PRI) related message from a bitstream that provides information about rectangular regions packed within a picture of one or more layers; and a step of deriving a picture size variable representing information about the size of the picture for each of the one or more layers based on the PRI related message, wherein the picture size variable may be defined for each of the one or more layers.
[0008] The above picture size variable can be indexed by a layer identifier representing a specific layer.
[0009] The above picture size variable can be indexed by the layer identifier of the picture associated with the information regarding the above rectangular areas.
[0010] The above picture size variable may include at least one of a picture width variable representing the width of the picture or a picture height variable representing the height of the picture.
[0011] The above picture width variable or the above picture height variable can be defined in units of luma samples.
[0012] The above picture size variable may include at least one of a picture maximum width variable representing the maximum width of the picture or a picture maximum height variable representing the maximum height of the picture.
[0013] The above picture maximum width variable or the above picture maximum height variable can be defined in units of luma samples.
[0014] The method further includes the step of deriving a rectangular position variable indicating the location of the rectangular area within the picture, wherein the rectangular position variable may be derived based on the picture size variable indexed by the layer identifier of the picture associated with the information regarding the rectangular area.
[0015] The method further includes the step of deriving a square size variable representing the size of the square area within the picture, wherein the square size variable may be derived based on the picture size variable indexed by the layer identifier of the picture associated with the information regarding the square area.
[0016] A method according to one aspect comprises: generating a Packed Regions Information (PRI) related message that provides information regarding rectangular regions packed within a picture of one or more layers; and encoding image information including the generated PRI related message; wherein the step of generating the PRI related message comprises: deriving a picture size variable representing information regarding the size of the picture for each of the one or more layers; and generating information regarding the location or size of the rectangular regions based on the picture size variable; wherein the PRI related message includes information regarding the location or size of the rectangular regions, and the picture size variable may be defined for each of the one or more layers.
[0017] The above picture size variable can be indexed by a layer identifier representing a specific layer.
[0018] The above picture size variable can be indexed by the layer identifier of the picture associated with the information regarding the above rectangular areas.
[0019] The above picture size variable may include at least one of a picture width variable representing the width of the picture, a maximum picture width variable representing the maximum width of the picture, a picture height variable representing the height of the picture, or a picture maximum height variable representing the maximum height of the picture.
[0020] A computer-readable storage medium according to one aspect can non-transiently store a bitstream generated by the above method.
[0021] A method according to one aspect comprises: a step of generating a bitstream; and a step of transmitting data including the bitstream; wherein the step of generating the bitstream comprises: a step of generating a Packed Regions Information (PRI) related message that provides information regarding rectangular regions packed within a picture of one or more layers; and a step of encoding image information including the generated PRI related message; wherein the step of generating the PRI related message comprises: a step of deriving a picture size variable representing information regarding the size of the picture for each of the one or more layers; and a step of generating information regarding the location or size of the rectangular regions based on the picture size variable; wherein the PRI related message includes information regarding the location or size of the rectangular regions, and the picture size variable may be defined for each of the one or more layers.
[0022] The features briefly summarized above regarding the present disclosure are merely exemplary aspects of the detailed description of the present disclosure that follows and do not limit the scope of the present disclosure.
[0023] According to the present disclosure, an encoding / decoding method and / or apparatus with improved coding efficiency can be provided.
[0024] According to the present disclosure, an encoding / decoding method and / or apparatus having improved data transmission efficiency can be provided.
[0025] According to the present disclosure, variables related to the size of a picture can be indexed as layer identifiers so that source layer information for each region can be accurately referenced even in an environment where multiple layers with different resolutions exist. Through this, information of all layers can be sufficiently represented, and variables can be accurately calculated by reflecting the actual resolution of the layer to which each region belongs when restoring a target picture. In addition, coordinate calculation errors that occur when combining image fragments (rectangular regions) of different layers into a single final screen can be prevented.
[0026] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure pertains from the description below.
[0027] FIG. 1 is a schematic diagram illustrating a video coding system to which an embodiment according to the present disclosure can be applied.
[0028] FIG. 2 is a schematic diagram showing an encoding device to which an embodiment according to the present disclosure can be applied.
[0029] FIG. 3 is a schematic diagram showing a decoding device to which an embodiment according to the present disclosure can be applied.
[0030] FIG. 4 shows an example of a video / image decoding method to which one embodiment can be applied.
[0031] FIG. 5 shows an example of a video / image encoding method to which an embodiment of the present disclosure can be applied.
[0032] Figure 6 illustrates an exemplary hierarchical structure for a coded video / image.
[0033] FIGS. 7 to 9 are drawings illustrating a method for decoding image information according to one embodiment.
[0034] FIGS. 10 and FIGS. 11 are drawings illustrating a method for encoding image information according to one embodiment.
[0035] FIG. 12 is a drawing illustrating an exemplary content streaming system to which an embodiment according to the present disclosure can be applied.
[0036] Hereinafter, embodiments of the present disclosure are described in detail with reference to the attached drawings so that those skilled in the art can easily implement them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein.
[0037] In describing the embodiments of the present disclosure, detailed descriptions of known configurations or functions are omitted if it is determined that such descriptions could obscure the essence of the present disclosure. Furthermore, parts of the drawings unrelated to the description of the present disclosure have been omitted, and similar parts are denoted by similar reference numerals.
[0038] In the present disclosure, when a component is described as being "connected," "combined," or "joined" with another component, this may include not only a direct connection but also an indirect connection in which another component exists in between. Furthermore, when a component is described as "comprising" or "having" another component, this means that, unless specifically stated otherwise, it does not exclude the other component but may include an additional component.
[0039] In the present disclosure, terms such as first, second, etc. are used solely for the purpose of distinguishing one component from another and do not limit the order or importance of the components unless specifically stated otherwise. Accordingly, within the scope of the present disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and likewise, a second component in one embodiment may be referred to as a first component in another embodiment.
[0040] In this disclosure, distinct components are intended to clearly describe their respective features and do not imply that the components are separate. That is, multiple components may be integrated to form a single hardware or software unit, or a single component may be distributed to form multiple hardware or software units. Accordingly, such integrated or distributed embodiments are included within the scope of this disclosure, unless otherwise noted.
[0041] In the present disclosure, the components described in various embodiments do not necessarily mean essential components, and some may be optional components. Accordingly, embodiments consisting of a subset of the components described in one embodiment are also included within the scope of the present disclosure. Furthermore, embodiments including additional components in addition to the components described in various embodiments are also included within the scope of the present disclosure.
[0042] The present disclosure relates to the encoding and decoding of images. For example, the methods and embodiments disclosed in this document may be applied to methods disclosed in the VVC (Versatile Video Coding) standard, VSEI (Versatile Supplemental Enhancement Information), HEVC (High Efficiency Video Coding) standard, EVC (Essential Video Coding) standard, AV1 (AOMedia Video 1) standard, AVC (Advanced Video Coding) standard, AVS2 (2nd generation of audio-video coding standard), or next-generation video / image coding standards (e.g., H.267 or H.268, etc.). However, the aforementioned standards are merely examples to which the methods according to the present disclosure are applicable, and it is understood that they may also be applied to other video coding technologies.
[0043] The present disclosure presents various embodiments relating to video / image coding, and unless otherwise stated, said embodiments may be performed in combination with one another.
[0044] Unless newly defined in this disclosure, the terms used herein may have the ordinary meanings commonly used in the technical field to which this disclosure belongs.
[0045] In this disclosure, "video" may refer to a set of images over time. In this disclosure, "picture" generally refers to a unit representing a single image at a specific time, and a slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles. A brick may represent a rectangular area of rows of CTUs within a tile in a picture. In this document, tile groups and slices may be used interchangeably. For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.
[0046] In the present disclosure, "pixel" or "pel" may refer to the smallest unit constituting a picture (or image). Additionally, "sample" may be used as a term corresponding to pixel. A sample may generally represent a pixel or a pixel value, may represent only the pixel / pixel value of the luminance component, or may represent only the pixel / pixel value of the chroma component.
[0047] In this disclosure, "unit" may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to that area. A unit may include one luminance block and two chroma (e.g., cb, cr) blocks. Depending on the case, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows.
[0048] In the present disclosure, "current block" may mean one of "current coding block," "current coding unit," "block to be encoded," "block to be decoded," or "block to be processed." When prediction is performed, "current block" may mean "current prediction block" or "block to be predicted." When transformation (inverse transformation) / quantization (inverse quantization) is performed, "current block" may mean "current transformation block" or "block to be transformed." When filtering is performed, "current block" may mean "block to be filtered."
[0049] In the present disclosure, "current block" may mean a block comprising both a luminous component block and a chroma component block, or "luma block of the current block," unless explicitly stated as a chroma block. The luminous component block of the current block may be expressed by including an explicit description of a luminous component block, such as "luma block" or "current luminous block." Additionally, the chroma component block of the current block may be expressed by including an explicit description of a chroma component block, such as "chroma block" or "current chroma block."
[0050] In the present disclosure, " / " and "," may be interpreted as "and / or." For example, "A / B" and "A, B" may be interpreted as "A and / or B." Additionally, "A / B / C" and "A, B, C" may mean "at least one of A, B and / or C."
[0051] In the present disclosure, "or" may be interpreted as "and / or". For example, "A or B" may mean 1) "A" only, 2) "B" only, or 3) "A and B". Alternatively, in the present disclosure, "or" may mean "additionally or alternatively".
[0052] In the present disclosure, "at least one A, B, and C" may mean "only A," "only B," "only C," or "any combination of two or more selected from the group consisting of A, B, and C." Additionally, "at least one A, B, or C" or "at least one A, B, and / or C" may mean "at least one A, B, and C."
[0053] FIG. 1 is a schematic diagram illustrating a video / image coding system to which an embodiment according to the present disclosure can be applied.
[0054] Referring to FIG. 1, a video / image coding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image or data in the form of a file or streaming to the receiving device via a digital storage medium or a network.
[0055] The source device may include a video source, an encoding device, and a transmission unit. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.
[0056] A video source may acquire video / images through processes such as video / image capture, synthesis, or generation. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include, for example, a computer, a tablet, and a smartphone, etc., and may generate video / images (electronically). For example, virtual video / images may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process in which related data is generated.
[0057] The encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0058] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device in the form of a file or streaming via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit may include elements for creating a media file through a predetermined file format and elements for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0059] The decoding device can decode video / images by performing a series of procedures such as inverse quantization, inverse transform, and prediction corresponding to the operation of the encoding device.
[0060] The renderer can render the decoded video / image. The rendered video / image can be displayed through the display unit.
[0061] FIG. 2 is a schematic diagram illustrating an encoding device to which an embodiment according to the present disclosure can be applied.
[0062] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a predictor (220), a residual processor (230), an entropy encoder (240), an adder (250), a filter (260), and a memory (270). The predictor (220) may include an inter-predictor (221) and an intra-predictor (222). The residual processor (230) may include a transformer (232), a quantizer (233), a dequantizer (234), and an inverse transformer (235). The residual processor (230) may further include a subtractor (231). The addition unit (250) may be referred to as a reconstructor or a reconstructed block generator. The aforementioned image segmentation unit (210), prediction unit (220), residual processing unit (230), entropy encoding unit (240), addition unit (250), and filtering unit (260) may be configured by one or more hardware components (e.g., an encoder chipset or processor) according to the embodiment. Additionally, the memory (270) may include a DPB (Decoded Picture Buffer) and may be configured by a digital storage medium. The hardware components may further include the memory (270) as an internal / external component.
[0063] The image segmentation unit (210) can divide an input image (or picture, frame) input to an encoding device (200) into one or more processing units. For example, the processing unit may be called a coding unit (CU). A coding unit may be recursively divided into a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, a single coding unit may be divided into multiple coding units of a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. For example, a quad-tree structure may be applied first, and a binary-tree structure and / or a ternary-tree structure may be applied later. Alternatively, a binary-tree structure may be applied first. A coding procedure according to the present disclosure may be performed based on the final coding unit that is no longer divided. In this case, based on coding efficiency according to image characteristics, the maximum coding unit may be used directly as the final coding unit, or, if necessary, the maximum coding unit may be recursively divided into lower-depth coding units so that a coding unit of the optimal size is used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). The prediction unit and the transformation unit may each be divided or partitioned from the final coding unit.The above prediction unit may be a unit of sample prediction, and the above transformation unit may be a unit that derives transformation coefficients and / or a unit that derives a residual signal from transformation coefficients.
[0064] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the context. In general, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chroma component. A sample may be used to refer to a single picture (or image) as a term corresponding to a pixel or pel.
[0065] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (predicted block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input image signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, as illustrated, the unit that subtracts the prediction signal (predicted block, prediction sample array) from the input image signal (original block, original sample array) within the encoding device (200) may be called a subtraction unit (231). The prediction unit (220) can perform a prediction for a block to be processed (hereinafter, current block) and generate a predicted block (predicted block) containing prediction samples for said current block. The prediction unit (220) can determine whether intra prediction is applied or inter prediction is applied in units of the current block or CU. The prediction unit (220) can generate various information regarding prediction, such as prediction mode information, as described below in the description of each prediction mode, and transmit it to the entropy encoding unit (240). The information regarding prediction can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.
[0066] The intra prediction unit (222) can predict the current block by referring to samples within the current picture. The referenced samples may be located near the current block or away from it, depending on the prediction mode. In intra prediction, the prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a Planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is merely an example, and depending on the settings, more or fewer directional prediction modes may be used. The intra prediction unit (222) may also determine the prediction mode applied to the current block by using the prediction mode applied to the surrounding blocks.
[0067] The inter prediction unit (221) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture containing the reference blocks and the reference picture containing the temporal neighboring blocks may be the same or different from each other. The temporal neighboring blocks may be referred to by names such as collocated reference block, collocated CU (colCU), etc. A reference picture containing the aforementioned temporal surrounding blocks may be called a collocated picture (colPic). For example, the inter prediction unit (221) may construct a list of motion information candidates based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of surrounding blocks as motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0068] The prediction unit (220) may generate a prediction signal based on various prediction methods and / or prediction techniques described below. For example, the prediction unit (220) may apply intra prediction or inter prediction for the prediction of the current block, as well as apply intra prediction and inter prediction simultaneously. A prediction method that applies intra prediction and inter prediction simultaneously for the prediction of the current block may be called combined inter and intra prediction (CIIP). Additionally, the prediction unit (220) may be based on an intra block copy (IBC) prediction mode or a palette mode for the prediction of the block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, for example, screen content coding (SCC). IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. Palette mode can be viewed as an example of intra-coding or intra-prediction. When palette mode is applied, sample values within a picture can be signaled based on information regarding palette tables and palette indices.
[0069] The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or to generate a residual signal. The subtraction unit (231) can generate a residual signal (residual signal, residual block, residual sample array) by subtracting the prediction signal (predicted block, prediction sample array) output from the prediction unit (220) from the input image signal (original block, original sample array). The generated residual signal can be transmitted to the conversion unit (232).
[0070] The transformation unit (232) can generate transform coefficients by applying a transformation technique to a residual signal. For example, the transformation technique may include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transformation obtained based on a prediction signal generated using all previously reconstructed pixels. The transformation process may be applied to a block of pixels of the same size in a square, or to a block of variable size that is not square.
[0071] The quantization unit (233) can quantize the transformation coefficients and transmit them to the entropy encoding unit (240). The entropy encoding unit (240) can encode the quantized signal (information regarding the quantized transformation coefficients) and output it as a bitstream. The information regarding the quantized transformation coefficients may be called residual information. The quantization unit (233) can rearrange the block-shaped quantized transformation coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information regarding the quantized transformation coefficients based on the one-dimensional vector-shaped quantized transformation coefficients.
[0072] The entropy encoding unit (240) can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit (190) may encode information required for video / image restoration (e.g., values of syntax elements) together or separately, in addition to quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may further include general constraint information. The signaling information, transmitted information, and / or syntax elements mentioned in the present disclosure may be included in the video / image information. The video / image information may be encoded through the encoding procedure described above and included in the bitstream.
[0073] The above bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmitting a signal output from the entropy encoding unit (240) and / or a storage unit (not shown) for storing it may be provided as an internal / external element of the encoding device (200), or the transmission unit may be provided as a component of the entropy encoding unit (240).
[0074] The quantized transformation coefficients output from the quantization unit (233) can be used to generate a residual signal. For example, a residual signal (residual block or residual samples) can be restored by applying inverse quantization and inverse transformation to the quantized transformation coefficients through the inverse quantization unit (234) and the inverse transformation unit (235).
[0075] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.
[0076] The adder (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit (221) or the intra prediction unit (222). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The adder (250) may be called a reconstructed unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after undergoing filtering as described below.
[0077] The filtering unit (260) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (260) can generate a modified restored picture by applying various filtering methods to the restored picture, and can store the modified restored picture in memory (270), specifically in the DPB of memory (170). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (240), as described below in the description of each filtering method. The information regarding filtering can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.
[0078] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter-prediction unit (221). Through this, the encoding device (200) can avoid prediction mismatches between the encoding device (200) and the decoding device when inter-prediction is applied, and can also improve encoding efficiency.
[0079] The DPB in memory (270) can store a modified restored picture to be used as a reference picture in the inter prediction unit (221). Memory (270) can store motion information of blocks from which motion information is derived (or encoded) in the current picture and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter prediction unit (221) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory (270) can store restoration samples of restored blocks in the current picture and transmit them to the intra prediction unit (222).
[0080] FIG. 3 is a schematic diagram illustrating a decoding device to which an embodiment according to the present disclosure can be applied.
[0081] As illustrated in FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-predictor (332) and an intra-predictor (331). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321). The aforementioned entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) may be configured by a single hardware component (e.g., a decoder chipset or a processor) according to an embodiment. Additionally, the memory (360) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.
[0082] When a bitstream containing video / image information is input, the decoding device (300) can restore the image by performing a process corresponding to the process performed by the encoding device (200) of FIG. 2. For example, the decoding device (300) can perform decoding using a processing unit applied in the encoding device (200). Thus, the processing unit for decoding may be, for example, a coding unit. The coding unit may be a coding tree unit, or a maximum coding unit may be obtained by dividing it according to a quad tree structure, a binary tree structure, and / or a binary tree structure. And, the restored image signal decoded and output through the decoding device (300) can be played back through a playback device (not shown).
[0083] The decoding device (300) can receive a signal output from the encoding device (200) of FIG. 2 in the form of a bitstream. The received signal can be decoded through an entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information necessary for image restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may further include general constraint information. The decoding device (300) can decode the picture based on the information regarding the parameter sets and / or the general constraint information. The signaling / received information and / or syntax elements described below can be obtained from the bitstream by decoding through the decoding procedure. For example, the entropy decoding unit (310) can decode information within the bitstream based on coding methods such as exponential chord coding, CAVLC, or CABAC, and output values of syntax elements required for image restoration and quantized values of transformation coefficients regarding residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using information on the syntax element to be decoded and decoding information of surrounding and decoding target blocks or information on symbols / bins decoded in the previous step, predict the probability of occurrence of the bin according to the determined context model, and perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin after determining the context model. Among the information decoded in the entropy decoding unit (310), information regarding prediction is provided to the prediction unit (330), and residual values for which entropy decoding was performed in the entropy decoding unit (310), i.e., quantized transformation coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive residual signals (residual blocks, residual samples, residual sample array). Additionally, among the information decoded in the entropy decoding unit (310), information regarding filtering can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (300), or the receiving unit may be a component of the entropy decoding unit (310). Meanwhile, the decoding device according to the present document may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit (310), and the sample decoder may include at least one of the inverse quantization unit (321), inverse transform unit (322), adder (340), filtering unit (350), memory (360), inter prediction unit (332), and intra prediction unit (331).
[0084] In the inverse quantization unit (321), the quantized transformation coefficients can be inversely quantized to output transformation coefficients. The inverse quantization unit (321) can rearrange the quantized transformation coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed in the encoding device (200). The inverse quantization unit (321) can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain transformation coefficients.
[0085] In the inverse conversion unit (322), the conversion coefficients can be inversely converted to obtain a residual signal (residual block, residual sample array).
[0086] The prediction unit (330) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit may apply intra prediction or inter prediction for a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called combined inter and intra prediction (CIIP). Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, for example, screen content coding (SCC). IBC basically performs prediction within the current picture, but it can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. The palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and palette index can be included in the above video / image information and signaled.
[0087] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The description of the intra prediction unit (222) may be applied equally to the intra prediction unit (331). The referenced samples may be located in the neighborhood of the current block or located away from it, depending on the prediction mode. In intra prediction, the prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit (331) may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0088] The inter prediction unit (332) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) may construct a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-prediction can be performed based on various prediction modes (techniques), and information regarding the prediction may include information indicating the mode (technique) of inter-prediction for the current block.
[0089] The adder (340) can generate a restoration signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the prediction signal (predicted block, predicted sample array) output from the prediction unit (330) (including the inter prediction unit (332) and / or intra prediction unit (331)). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the restoration block. The description of the adder (250) can be applied equally to the adder (340). The adder (340) may be called a restoration unit or a restoration block generation unit. The generated restoration signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after undergoing filtering as described below.
[0090] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0091] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can generate a modified restored picture by applying various filtering methods to the restored picture, and can store the modified restored picture in memory (360), specifically in the DPB of memory (360). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0092] The (modified) restored picture stored in the DPB of the memory (360) can be used as a reference picture in the inter-prediction unit (332). The memory (360) can store motion information of blocks from which motion information within the current picture has been derived (or decoded) and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (332) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (360) can store restoration samples of blocks restored within the current picture and transmit them to the intra-prediction unit (331).
[0093] In this specification, the embodiments described in the filtering unit (260), inter prediction unit (221), and intra prediction unit (222) of the encoding device (200) may be applied to the filtering unit (350), inter prediction unit (332), and intra prediction unit (331) of the decoding device (300) in the same or corresponding manner.
[0094] A video / image coding method according to the present disclosure may be performed based on the following partitioning structure. Specifically, the procedures described below, such as prediction, residual processing ((inverse)transform, (inverse)quantization, etc.), syntax element coding, and filtering, may be performed based on CTU and CU (and / or TU, PU) derived based on the partitioning structure. The block partitioning procedure may be performed in the image segmentation unit (210) of the encoding device described above, and the partitioning-related information may be processed (encoded) in the entropy encoding unit (240) and transmitted to the decoding device in the form of a bitstream. The entropy decoding unit (310) of the decoding device may derive the block partitioning structure of the current picture based on the partitioning-related information obtained from the bitstream, and perform a series of procedures for image decoding (e.g., prediction, residual processing, block / picture restoration, in-loop filtering, etc.) based thereon. The CU size and the TU size may be the same, or multiple TUs may exist within the CU area. Meanwhile, the term CU size generally refers to the CB size of the luminous component (sample). The term TU size generally refers to the TB size of the luminous component (sample).
[0095] The chroma component (sample) CB or TB size can be derived based on the luminance component (sample) CB or TB size according to the component ratio based on the color format (chroma format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.) of the picture / image. The TU size can be derived based on maxTbSize. For example, if the CU size is larger than the maxTbSize, multiple TUs (TBs) of the maxTbSize are derived from the CU, and conversion / inverse conversion can be performed in units of the TU (TB). Additionally, for example, when intra prediction is applied, the intra prediction mode / type is derived in units of the CU (or CB), and the procedure for deriving surrounding reference samples and generating prediction samples can be performed in units of the TU (or TB). In this case, one or more TUs (or TBs) may exist within a single CU (or CB) region, and in this case, the multiple TUs (or TBs) may share the same intra prediction mode / type.
[0096] Additionally, in the coding of video / image according to the present disclosure, the image processing unit may have a hierarchical structure. A picture may be divided into one or more tiles, bricks, slices, and / or tile groups. A slice may include one or more bricks. A brick may include one or more CTU rows within the tile. A slice may include an integer number of bricks in the picture. A tile group may include one or more tiles. A tile may include one or more CTUs. The CTU may be divided into one or more CUs. A tile is a rectangular area within a picture that includes CTUs within a specific tile row and a specific tile column. A tile group may include an integer number of tiles according to a tile raster scan within the picture. A slice header may carry information / parameters that can be applied to the corresponding slice (blocks within the slice). If the encoding / decoding device has a multi-core processor, the encoding / decoding procedure for the tile, slice, brick, and / or tile group may be processed in parallel.
[0097] In the present disclosure, slices or tile groups may be used interchangeably. That is, a tile group header may be referred to as a slice header. Here, a slice may have one of the slice types including an intra (I) slice, a predictive (P) slice, and a bi-predictive (B) slice. For blocks within an I slice, only intra prediction may be used for prediction, and no inter prediction may be used. Of course, even in this case, the original sample value may be coded and signaled without prediction. For blocks within a P slice, intra prediction or inter prediction may be used, and if inter prediction is used, only uni prediction may be used. Meanwhile, for blocks within a B slice, intra prediction or inter prediction may be used, and if inter prediction is used, up to bi-prediction may be used.
[0098] In an encoding device, tile / tile group, brick, slice, and maximum and minimum coding unit sizes are determined based on the characteristics of the video image (e.g., resolution) or by considering coding efficiency or parallel processing, and information regarding this or information that can derive it may be included in the bitstream.
[0099] The decoding device can obtain information indicating whether the tile / tile group, brick, slice, or CTU within the tile of the current picture has been divided into multiple coding units. Efficiency can be increased by obtaining (transmitting) this information only under specific conditions.
[0100] The slice header (slice header syntax) may include information / parameters that can be commonly applied to the slice. The APS (APS syntax) or PPS (PPS syntax) may include information / parameters that can be commonly applied to one or more pictures. The SPS (SPS syntax) may include information / parameters that can be commonly applied to one or more sequences. The VPS (VPS syntax) may include information / parameters that can be commonly applied to multiple layers. The DPS (DPS syntax) may include information / parameters that can be commonly applied to the entire video. The DPS may include information / parameters related to the concatenation of the CVS (coded video sequence).
[0101] In the present disclosure, the term "higher-level syntax" may include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, and slice header syntax.
[0102] In addition, for example, information regarding the division and configuration of the tile / tile group / brick / slice can be configured at the encoding stage through the upper-level syntax and transmitted to the decoding device in the form of a bitstream.
[0103] Pictures can be divided into sequences of Coding Tree Units (CTUs). A CTU may correspond to a Coding Tree Block (CTB). Alternatively, a CTU may include a Coding Tree Block of Luma Samples and two Coding Tree Blocks of corresponding Chroma Samples. In other words, for a picture containing three sample arrays, a CTU may include an NxN block of Luma Samples and two corresponding blocks of Chroma Samples.
[0104] The maximum allowable size of a CTU for coding and prediction, etc., may differ from the maximum allowable size of a CTU for transformation. For example, the maximum allowable size of a luminance block within a CTU may be 128x128 (even though the maximum size of luminance ring blocks is 64x64).
[0105] A picture is divided into one or more tile rows and one or more tile columns. A tile is a sequence of CTUs covering a rectangular area of the picture. The CTUs within a tile are scanned in raster scan order within that tile.
[0106] A slice consists of an integer number of complete tiles within a picture or an integer number of consecutive complete CTU rows within a single tile. Two modes are supported for slicing: raster-scan slice mode and rectangular slice mode.
[0107] In raster-scan slice mode, a slice comprises a sequence of complete tiles according to the tile raster scan order of the picture. In rectangular slice mode, a slice comprises a number of complete tiles collectively configured to form a rectangular area of the picture, or a number of consecutive complete CTU rows that collectively form a rectangular area within a single tile. The tiles within a rectangular slice are scanned in the tile raster scan order within the rectangular area corresponding to that slice.
[0108] A subpicture includes one or more slices that collectively cover a rectangular area of a picture. For example, it is possible for a picture to be divided into 28 subpictures of different sizes.
[0109] When a picture is encoded into three separate color planes (where separate_colour_plane_flag is 1), the slice contains only CTUs of a single color component identified by the corresponding value of colour_plane_id, and each array of color components of the picture consists of slices having the same colour_plane_id value.
[0110] Encoded slice NAL units having different colour_plane_id values within a picture can be interleaved with respect to each colour_plane_id value, provided that for each colour_plane_id value, the encoded slice NAL units having that colour_plane_id value are arranged in an order of increasing CTU addresses in the tile scan order for the first CTU of each slice NAL unit.
[0111] Meanwhile, when separate_colour_plane_flag is 0, each CTU of the picture is contained in exactly one slice. When separate_colour_plane_flag is 1, the CTU of each color component is contained in exactly one slice (i.e., information for each CTU of the picture is contained in exactly three slices, and these three slices have different colour_plane_id values).
[0112] Tiles change the order of CTUs within a picture. If a picture is divided into two or more tiles, the order of CTUs becomes the raster-scan order within each tile, which can be exemplified by a case where the picture is divided into two tiles and each tile has 8 CTUs. Note that the CTUs are arranged in raster-scan order within each tile.
[0113] FIG. 4 shows an example of a video / image decoding method to which one embodiment can be applied.
[0114] In video coding, the pictures constituting the video can be decoded according to a series of decoding orders. The picture order corresponding to the output order of the decoded pictures can be set differently from the decoding order, and based on this, not only forward prediction but also reverse prediction can be performed during inter-prediction.
[0115] In FIG. 4, S400 may be performed in the entropy decoding unit (310) of the aforementioned decoding device (300), S410 may be performed in the prediction unit (330), S420 may be performed in the residual processing unit (320), S430 may be performed in the addition unit (340), and S440 may be performed in the filtering unit (350). S400 may include a decoding procedure according to the present disclosure, S410 may include an inter / intra prediction procedure according to the present disclosure, S420 may include a residual processing procedure according to the present disclosure, S430 may include a block / picture restoration procedure according to the present disclosure, and S440 may include an in-loop filtering procedure according to the present disclosure.
[0116] Referring to FIG. 4, the decoding device acquires image / video information from a bitstream (S400), performs a prediction based on the acquired image / video information (S410), and can restore a picture through residual processing (S420, inverse quantization and inverse transformation of quantized transformation coefficients) (S430).
[0117] A modified restored picture can be generated by applying an in-loop filtering procedure (S440) to the restored picture generated through the above restoration procedure, and the modified restored picture can be output as a decoded picture and also stored in the buffer or memory of the decoding device to be used as a reference picture in the inter-prediction procedure when decoding the next picture. In some cases, the above in-loop filtering procedure may be omitted, in which case the restored picture can be output as a decoded picture and also stored in the buffer or memory of the decoding device to be used as a reference picture in the inter-prediction procedure when decoding a subsequent picture.
[0118] The in-loop filtering procedure (S440) may include a deblocking filtering procedure, a sample adaptive offset (SAO) procedure, an adaptive loop filter (ALF) procedure, and / or a bilateral filter procedure, and some or all of these may be omitted. Additionally, one or some of the deblocking filtering procedure, the sample adaptive offset (SAO) procedure, the adaptive loop filter (ALF) procedure, and the bilateral filter procedure may be applied sequentially, or all of them may be applied sequentially. For example, the SAO procedure may be performed after the deblocking filtering procedure is applied to the restored picture. Alternatively, for example, the ALF procedure may be performed after the deblocking filtering procedure is applied to the restored picture. This may be performed in the same manner in the encoding device.
[0119] FIG. 5 shows an example of a video / image encoding method to which an embodiment of the present disclosure can be applied.
[0120] In FIG. 5, the prediction step (S500) may be performed in the prediction unit (220) of the aforementioned encoding device (200), residual processing (S510) based on the prediction result may be performed in the residual processing unit (230), and the step (S520) of encoding image information including prediction information and residual information may be performed in the entropy encoding unit (240). S500 may include an inter / intra prediction procedure according to the present disclosure, S510 may include a residual processing procedure according to the present disclosure, and S520 may include an encoding procedure according to the present disclosure.
[0121] The encoding procedure may optionally include not only a procedure for encoding information for picture restoration (e.g., prediction information, residual information, partitioning information, etc.) and outputting it in the form of a bitstream, but also a procedure for generating a restored picture for the current picture and a procedure for applying in-loop filtering to the restored picture.
[0122] The encoding device (200) can derive (modified) residual samples from quantized transform coefficients through the inverse quantization unit (234) and the inverse transform unit (235), and can generate a restored picture based on the (modified) residual samples and the predicted samples which are the outputs of S500. The restored picture thus generated may be identical to the restored picture generated by the decoding device (300) described above. A modified restored picture may be generated through an in-loop filtering procedure on the restored picture, which may be stored in a buffer or memory, and, as in the case of the decoding device, may be used as a reference picture in the inter-prediction procedure during the subsequent encoding of the picture.
[0123] As described above, depending on the case, part or all of the in-loop filtering procedure may be omitted. When the in-loop filtering procedure is performed, (in-loop) filtering-related information (parameters) may be encoded in the entropy encoding unit (240) and output in the form of a bitstream, and the decoding device (300) may perform the in-loop filtering procedure in the same way as the encoding device based on the filtering-related information.
[0124] Through this in-loop filtering procedure, noise generated during video / image coding, such as blocking artifacts and ringing artifacts, can be reduced, and subjective / objective image quality can be improved. In addition, by performing the in-loop filtering procedure in both the encoding device (200) and the decoding device (300), the same prediction results can be derived in both the encoding device (200) and the decoding device (300), the reliability of picture coding can be increased, and the amount of data that must be transmitted for picture coding can be reduced.
[0125] As described above, the picture restoration procedure can be performed in the encoding device (200) as well as the decoding device (300). Restoration blocks can be generated based on intra prediction / inter prediction for each block unit, and a restored picture containing the restoration blocks can be generated. If the current picture / slice / tile group is an I picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based solely on intra prediction. Meanwhile, if the current picture / slice / tile group is a P or B picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based on intra prediction or inter prediction. In this case, inter prediction may be applied to some blocks within the current picture / slice / tile group, and intra prediction may be applied to the remaining blocks.
[0126] The color components of the picture may include a luminance component and a chroma component, and unless explicitly limited in the present disclosure, embodiments according to the present disclosure may be applied to the luminance component and the chroma component.
[0127] Figure 6 illustrates an exemplary hierarchical structure for a coded video / image.
[0128] Referring to Fig. 6, the coded image is divided into a Video Coding Layer (VCL) that handles the decoding processing of the image and the image itself, a subsystem that transmits and stores the encoded information, and a Network Abstraction Layer (NAL) that exists between the VCL and the subsystem and is responsible for network adaptation functions.
[0129] In VCL, VCL data containing compressed image data (slice data) can be generated, or parameter sets containing information such as Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS), or SEI (Supplemental Enhancement Information) messages that are additionally required in the decoding process of the image can be generated.
[0130] In NAL, a NAL unit can be created by adding header information (NAL unit header) to the Raw Byte Sequence Payload (RBSP) generated in VCL. In this case, the RBSP refers to slice data, parameter sets, SEI messages, etc. generated in VCL. The NAL unit header may include NAL unit type information specified according to the RBSP data included in the NAL unit.
[0131] As shown in FIG. 6, NAL units can be classified into VCL NAL units and Non-VCL NAL units depending on the RBSP generated in VCL. A VCL NAL unit may refer to a NAL unit containing information about an image (slice data), and a Non-VCL NAL unit may refer to a NAL unit containing information necessary to decode an image (parameter set or SEI message).
[0132] The aforementioned VCL NAL unit and Non-VCL NAL unit can be transmitted over a network by attaching header information according to the data specifications of the underlying system. For example, the NAL unit can be transformed into a data format of a specified specification, such as H.266 / VVC file format, RTP (Real-time Transport Protocol), TS (Transport Stream), etc., and transmitted over various networks.
[0133] As described above, the NAL unit type can be determined according to the RBSP data structure included in the NAL unit, and information about this NAL unit type can be stored in the NAL unit header and signaled.
[0134] For example, NAL units can be broadly classified into VCL NAL unit types and Non-VCL NAL unit types depending on whether they contain information about the image (slice data). VCL NAL unit types can be classified according to the properties and types of the picture included in the VCL NAL unit, while Non-VCL NAL unit types can be classified according to the types of parameter sets.
[0135] The following is an example of a NAL unit type specified according to the type of parameter set included in the Non-VCL NAL unit type.
[0136] - APS (Adaptation Parameter Set) NAL unit: Type for the NAL unit containing the APS
[0137] - DPS(Decoding Parameter Set) NAL unit: Type for the NAL unit containing the DPS
[0138] - VPS (Video Parameter Set) NAL unit: Type for the NAL unit containing the VPS
[0139] - SPS (Sequence Parameter Set) NAL unit: Type for the NAL unit containing the SPS
[0140] - PPS(Picture Parameter Set) NAL unit: Type for the NAL unit containing the PPS
[0141] The above-described NAL unit types have syntax information for the NAL unit type, and said syntax information can be stored in the NAL unit header and signaled. For example, said syntax information may be nal_unit_type, and NAL unit types may be specified by the nal_unit_type value.
[0142] A slice header (slice header syntax, slice header information) may include information / parameters that can be commonly applied to the slice. The APS (APS syntax) or PPS (PPS syntax) may include information / parameters that can be commonly applied to one or more slices or pictures. The SPS (SPS syntax) may include information / parameters that can be commonly applied to one or more sequences. The VPS (VPS syntax) may include information / parameters that can be commonly applied to multiple layers. The DPS (DPS syntax) may include information / parameters that can be commonly applied to the entire video. The DPS may include information / parameters related to the concatenation of a CVS (coded video sequence). In the present disclosure, High Level Syntax (HLS) may include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, or slice header syntax.
[0143] In the present disclosure, image / video information encoded by an encoding device and signaled in the form of a bitstream includes not only information related to picture partitioning, intra / inter prediction information, residual information, in-loop filtering information, etc., but may also include information included in the slice header, information included in the APS, information included in the PPS, information included in the SPS, information included in the VPS, and / or information included in the DPS.
[0144] A coded picture may consist of one or more slices. Parameters describing the coded picture are signaled within the picture header (PH), and parameters describing the slices are signaled within the slice header (SH). The PH is transmitted as its own NAL unit type. The SH is located at the beginning of the NAL unit containing the slice payload (i.e., slice data).
[0145] Hereinafter, SEI messages related to embodiments of the present disclosure will be described.
[0146] An embodiment according to the present disclosure relates to a Packed Regions Information (PRI) SEI message. A PRI SEI message provides information regarding rectangular regions packed into a coded picture. A specific Region of Interest (ROI) within an image is of greater interest than the rest of the image in many use cases. A PRI SEI message enables the packing of rectangular ROIs from an original picture into a smaller resolution picture for image coding, thereby reducing the pixel rate and bitrate. Such SEI messages signal metadata describing the size and location of the ROIs in the coded picture and the original picture. A decoder may use said metadata to reconstruct target pictures of the original resolution from decoded pictures containing the packed regions.
[0147] The PRI SEI message provides information about rectangular regions packed into the coded picture. This information can be optionally used to reconstruct the target picture from samples of cropped decoded pictures corresponding to the regions described in the SEI message.
[0148] An example of a syntax table for a PRI SEI message is shown in Table 1 below.
[0149] [Table 1]
[0150]
[0151] The use of this SEI message requires the definition of the following variables:
[0152] - The width and height of the picture expressed in luma samples, denoted as PicWidthInLumaSamples and PicHeightInLumaSamples, respectively.
[0153] - Maximum picture width and maximum picture height expressed in luma samples, indicated as MaxPicWidth and MaxPicHeight, respectively.
[0154] - As a chroma format indicator, it is displayed as ChromaFormatIdc.
[0155] - BitDepthY is indicated as the bit depth for samples of the chroma component, and BitDepthC is indicated as the bit depth for samples of two associated chroma components if ChromaFormatIdc is not equal to 0.
[0156] pri_cancel_flag may indicate whether the SEI message cancels the persistence of the previous PRI SEI message in the output order applied to the current layer. A pri_cancel_flag equal to 1 may indicate that the SEI message cancels the persistence of the previous PRI SEI message in the output order applied to the current layer. A pri_cancel_flag equal to 0 may indicate that packed regions information follows.
[0157] The pri_persistence_flag specifies the persistence of PRI SEI messages for the current layer. A pri_persistence_flag equal to 0 specifies that packed region information applies only to the currently decoded picture. A pri_persistence_flag equal to 1 specifies that PRI SEI messages apply to the currently decoded picture and persist to all subsequent pictures of the current layer in the output order until one or more of the following conditions are true:
[0158] - When a new CLVS of the current layer starts.
[0159] - When the bitstream ends.
[0160] - When the picture of the current layer included in the AU (Access Unit) associated with the PRI SEI message is output, and that picture is positioned after the current picture in the output order.
[0161] pri_num_regions_minus1 may represent information related to the number of regions where information is signaled. A value of pri_num_regions_minus1 plus 1 may specify the number of regions where the information is signaled.
[0162] pri_multilayer_flag can indicate whether the syntax element pri_region_layer_id[i] exists. pri_multilayer_flag equal to 1 can indicate that pri_region_layer_id[i] exists. pri_multilayer_flag equal to 0 can indicate that pri_region_layer_id[i] does not exist.
[0163] pri_use_max_dimensions_flag can specify whether MaxPicWidth, MaxPicHeight, PicWidthInLumaSamples, and PicHeightInLumaSamples are used in variable calculations. pri_use_max_dimensions_flag equal to 1 can specify that MaxPicWidth, MaxPicHeight, PicWidthInLumaSamples, and PicHeightInLumaSamples are used in variable calculations. pri_use_max_dimensions_flag equal to 0 can specify that MaxPicWidth, MaxPicHeight, PicWidthInLumaSamples, and PicHeightInLumaSamples are not used in variable calculations for area parameters.
[0164] pri_log2_unit_size can specify the unit size used for variable calculations for area parameters.
[0165] The variable priUnitSize can be set to 1 << pri_log2_unit_size.
[0166] pri_region_size_len_minus1 can represent information regarding the number of bits used by various syntax elements that define the size and location of a region within a PRI SEI message. pri_region_size_len_minus1 plus 1 specifies the number of bits used to signal pri_region_top_left_in_units_x[i], pri_region_top_left_in_units_y[i], pri_region_width_in_units_minus1[i], pri_region_height_in_units_minus1[i], pri_target_region_top_left_x[i], and pri_target_region_top_left_y[i].
[0167] pri_region_id_present_flag can indicate whether the syntax element pri_region_id[i] exists. pri_region_id_present_flag equal to 1 can indicate that the syntax element pri_region_id[i] exists. pri_region_id_present_flag equal to 0 can indicate that the syntax element pri_region_id[i] does not exist.
[0168] pri_target_pic_params_present_flag may indicate whether syntax elements representing parameters required to restore the target picture (e.g., the total size of the target picture, area-by-area layout, etc.) exist in the PRI SEI message. pri_target_pic_params_present_flag identical to 1 may indicate that the syntax elements pri_target_region_top_left_x[i], pri_target_region_top_left_y[i], pri_target_pic_width_minus1, and pri_target_pic_height_minus1 exist. pri_target_pic_params_present_flag equal to 0 may indicate that the syntax elements pri_target_region_top_left_x[i], pri_target_region_top_left_y[i], pri_target_pic_width_minus1, and pri_target_pic_height_minus1 do not exist.
[0169] pri_target_pic_width_minus1 may represent information regarding the width of the target picture in luminance samples that can be reconstructed from samples of the cropped decoded picture corresponding to the regions described in the PRI SEI message. If pri_target_pic_width_minus1 is present, the value of pri_target_pic_width_minus1 plus 1 may represent the width of the target picture in luminance samples.
[0170] pri_target_pic_height_minus1 may represent information regarding the height of the target picture in luminance samples that can be reconstructed from samples of the cropped decoded picture corresponding to the regions described in the PRI SEI message. If pri_target_pic_height_minus1 is present, the value of pri_target_pic_height_minus1 plus 1 may represent the height of the target picture in luminance samples.
[0171] pri_num_resampling_ratios_minus1 can represent information related to the number of signaled resampling ratios. A value of pri_num_resampling_ratios_minus1 plus 1 can specify the number of signaled resampling ratios.
[0172] The value obtained by adding 1 to pri_resampling_width_num_minus1[i] and the value obtained by adding 1 to pri_resampling_width_denom_minus1[i] can respectively specify the numerator and denominator for width resampling of the i-th resampling rate. Both pri_resampling_width_num_minus1[i] and pri_resampling_width_denom_minus1[i] may be within the range of 0 to 65535 (inclusive).
[0173] If pri_resampling_width_num_minus1[i] and pri_resampling_width_denom_minus1[i] do not exist, the values of pri_resampling_ratio_width_num_minus1[0] and pri_resampling_ratio_width_denom_minus1[0] can be estimated to be 0.
[0174] pri_fixed_aspect_ratio_flag[i] can specify whether the syntax elements pri_resampling_height_num_minus1[i] and pri_resampling_height_denom_minus1[i] exist. pri_fixed_aspect_ratio_flag[i] equal to 1 can specify that the syntax elements pri_resampling_height_num_minus1[i] and pri_resampling_height_denom_minus1[i] do not exist. pri_fixed_aspect_ratio_flag[i] equal to 0 can specify that pri_resampling_height_num_minus1[i] and pri_resampling_height_denom_minus1[i] exist.
[0175] The value obtained by adding 1 to pri_resampling_height_num_minus1[i] and the value obtained by adding 1 to pri_resampling_height_denom_minus1[i] can respectively specify the numerator and denominator of height resampling for the i-th resampling rate. Both pri_resampling_height_num_minus1[i] and pri_resampling_height_denom_minus1[i] may be within the range of 0 to 65,535 (inclusive).
[0176] If pri_resampling_height_num_minus1[i] and pri_resampling_height_denom_minus1[i] do not exist, the values of pri_resampling_height_num_minus1[i] and pri_resampling_height_denom_minus1[i] can be assumed to be the same as pri_resampling_width_num_minus1[i] and pri_resampling_width_denom_minus1[i], respectively.
[0177] pri_region_id[i] can represent the ID of the i-th region. If it does not exist, the value of pri_region_id[i] can be assumed to be the same as i.
[0178] pri_region_layer_id[i] can specify the layer ID of the picture to which area information is associated, such as pri_region_top_left_in_units_x[i], pri_region_top_left_in_units_y[i], pri_region_width_in_units_minus1[i], and pri_region_height_in_units_minus1[i]. If it does not exist, the value of pri_region_layer_id[i] can be assumed to be equal to 0.
[0179] pri_region_is_a_layer_flag[i] can specify whether the picture width and height in the layer with ID pri_region_layer_id[i] are equal to the width and height of the region at index i. pri_region_is_a_layer_flag[i] equal to 1 can specify that the picture width and height in the layer with ID pri_region_layer_id[i] are equal to the width and height of the region at index i, in which case pri_region_top_left_in_units_x[i], pri_region_top_left_in_units_y[i], pri_region_width_in_units_minus1[i], and pri_region_height_in_units_minus1[i] are not signaled. If pri_region_is_a_layer_flag[i] does not exist, the value of pri_region_is_a_layer_flag[i] can be assumed to be equal to 0.
[0180] pri_region_top_left_in_units_x[i] and pri_region_top_left_in_units_y[i] can specify the horizontal and vertical positions of the top-left sample of the i-th region, respectively, in units. The length of these syntax elements can be pri_region_size_len_minus1 + 1.
[0181] Variables priRegionTopLeftX[i] and priRegionTopLeftY[i], representing the horizontal and vertical positions in luminance sample units of the i-th region respectively in a cropped decoded picture where the layer identifier is the same as pri_region_layer_id[i], can be derived as shown in the following Equation 1:
[0182] [Equation 1]
[0183]
[0184] pri_region_width_in_units_minus1[i] and pri_region_height_in_units_minus1[i] can respectively represent information regarding the width and height of the i-th region expressed in units. Values obtained by adding 1 to pri_region_width_in_units_minus1[i] and values obtained by adding 1 to pri_region_height_in_units_minus1[i] can respectively specify the width and height of the i-th region expressed in units. The length of these syntax elements can be pri_region_size_len_minus1 + 1.
[0185] The variables priRegionWidth[i] and priRegionHeight[i], representing the width and height in luma sample units of the i-th region in the cropped decoded picture, respectively, can be derived as shown in the following Equation 2:
[0186] [Equation 2]
[0187]
[0188] As a requirement for bitstream conformance, the value of priRegionWidth[i] % SubWidthC must be equal to 0 and the value of priRegionHeight[i] % SubHeightC must be equal to 0. In other words, the remainder when priRegionWidth[i] is divided by SubWidthC must be 0, and the remainder when priRegionHeight[i] is divided by SubHeightC must be 0. That is, priRegionWidth[i] must be a multiple of priRegionWidth[i] and priRegionHeight[i] must be a multiple of SubHeightC.
[0189] The variables SubWidthC and SubHeightC represent the horizontal and vertical subsampling factors of the chroma component for the luminance component and can be derived from ChromaFormatIdc as specified in Table 2 below.
[0190] [Table 2]
[0191]
[0192] pri_resampling_ratio_idx[i] can specify the index of the resampling ratio used for the i-th region. The length of the corresponding syntax element can be Ceil( Log2( pri_num_resampling_ratios_minus1 + 1 ) )
[0193] The variables priResampleWidthNum[i], priResampleWidthDenom[i], priResampleHeightNum[i], and priResampleHeightDenom[i] can be derived as shown in Equation 3 below.
[0194] [Equation 3]
[0195]
[0196] If present, pri_target_region_top_left_x[i] and pri_target_region_top_left_y[i] may represent the horizontal and vertical positions, respectively, of the top-left sample position in luminance sample units of the i-th region in the restored target picture.
[0197] The variables priTargetRegionWidth and priTargetRegionHeight, representing the luminance sample unit width and height of the resampled region in the restored target picture, respectively, can be derived as shown in the following Equation 4:
[0198] [Equation 4]
[0199]
[0200] When restoring a target picture with a luminance sample array size of (pri_target_pic_width_minus1 + 1) Х (pri_target_pic_height_minus1 + 1), all luminance sample values can be initialized to 1 << (BitDepthY - 1), and if chroma samples exist, the chroma sample values can be initialized to 1 << (BitDepthC - 1).
[0201] For any sample location (x, y) and regions j and k, if all of the following conditions are satisfied, the restored target picture sample at location (x, y) can be determined by the parameters signaled for the j-th region:
[0202] - pri_region_id [ j ] > pri_region_id [ k ]
[0203] - x is within the range in (priRegionTopLeftX[ j ] .. priRegionTopLeftX[ j ] + priRegionWidth[ j ])
[0204] - y is within the range in (priRegionTopLeftY[ j ] .. priRegionTopLeftY[ j ] + priRegionHeight[ j ]
[0205] - x is within the range in (priRegionTopLeftX[ k ] .. priRegionTopLeftX[ k ] + priRegionWidth[ k])
[0206] - y is within the range in (priRegionTopLeftY[ k ] .. priRegionTopLeftY[ k ] + priRegionHeight[ k ]
[0207] The PRI SEI message is included in VSEI version 4. To use the PRI SEI message, some variables need to be assigned by the decoder based on the codec of the bitstream containing the SEI message. To use the PRI SEI message, the following information needs to be provided by the codec:
[0208] The use of this SEI message requires the definition of the following variables:
[0209] - Picture width and picture height in luma sample units, referred to as PicWidthInLumaSamples and PicHeightInLumaSamples, respectively
[0210] - Maximum picture width and maximum picture height in luma sample units, referred to as MaxPicWidth and MaxPicHeight, respectively
[0211] - Chroma format specifier referred to as ChromaFormatIdc
[0212] - Bit depth for the chroma component sample, referred to as BitDepthY. Bit depth for the two associated chroma component samples, referred to as BitDepthC, if ChromaFormatIdc is not equal to 0.
[0213] In the case of VVC, the assignment of necessary variables is as follows:
[0214] The following variables may be specified for the interpretation of PRI SEI messages:
[0215] - PicWidthInLumaSamples can be set to the same value as pps_pic_width_in_luma_samples - SubWidthC * (pps_conf_win_left_offset + pps_conf_win_right_offset).
[0216] - PicHeightInLumaSamples can be set to the value of pps_pic_height_in_luma_samples - SubHeightC * (pps_conf_win_top_offset + pps_conf_win_bottom_offset).
[0217] - MaxPicWidth can be set to the value of sps_pic_width_in_luma_samples - SubWidthC * (sps_conf_win_left_offset + sps_conf_win_right_offset).
[0218] - MaxPicHeight can be set to the value of sps_pic_height_in_luma_samples - SubHeightC * (sps_conf_win_top_offset + sps_conf_win_bottom_offset).
[0219] - ChromaFormatIdc can be set to be the same as sps_chroma_format_idc.
[0220] - BitDepthY and BitDepthC can both be set to be the same as BitDepth.
[0221] When PRI SEI messages exist in a multi-layer bitstream, there may be cases where the variables required to process the PRI SEI messages are insufficient. In such environments, the bitstream may contain multiple layers, each having a different picture resolution and a different sample format. Therefore, the following variables may cause problems:
[0222] Picture width and picture height in luma samples, referred to as -PicWidthInLumaSamples and PicHeightInLumaSamples. These variables need to be provided as a list of picture widths and heights from each layer associated with the SEI message.
[0223] - The maximum picture width and maximum picture height in luma samples, referred to as MaxPicWidth and MaxPicHeight. These variables need to be provided as a list of the maximum picture width and height from each layer associated with the SEI message.
[0224] - A chroma format specifier referred to as ChromaFormatIdc. Each layer may have a different chroma format, but if regions from different layers need to be restored in the target picture for the use of the corresponding SEI message, the chroma formats of these regions must be the same.
[0225] Bit depth for the chroma component sample, designated as -BitDepthY. Bit depth for the two associated chroma component samples, designated as BitDepthC, if ChromaFormatIdc is not equal to 0. Each layer may have a different bit depth, but if regions from different layers need to be restored from the target picture for use in the SEI message, the bit depths of these regions must be the same.
[0226] Embodiments according to the present disclosure may apply the following methods to solve the aforementioned problem. Each method may be applied individually, or two or more methods may be applied in combination:
[0227] The picture width and picture height variables required by the -PRI SEI message can be provided as lists of picture widths and lists of picture heights. That is, the variables PicWidthInLumaSamples and PicHeightInLumaSamples can be provided by indexing them by layer.
[0228] - References to the picture width and picture height variables can be updated according to the layers corresponding to those variables.
[0229] The maximum picture width and maximum picture height variables required by the -PRI SEI message can be provided as a list of maximum picture widths and a list of maximum picture heights. That is, the variables PicWidthInLumaSamples and PicHeightInLumaSamples can be provided by indexing them by layer.
[0230] - References to the maximum picture width and maximum picture height variables can be updated according to the layers corresponding to those variables.
[0231] Hereinafter, embodiments according to the present disclosure will be described in detail.
[0232] For example, one embodiment may be based on the VVC standard and VSEI.
[0233] As described above, the PRI SEI message may provide information regarding rectangular regions packed within the coded picture. This information may optionally be used to restore the target picture from samples of the cropped decoded picture corresponding to the regions described in the SEI message.
[0234] The use of this SEI message may require the definition of the following variables:
[0235] - Lists of picture widths and picture heights in luma samples, referred to as PicWidthInLumaSamples[ ] and PicHeightInLumaSamples[ ], respectively.
[0236] - Lists of maximum picture widths and maximum picture heights in luma samples, referred to as MaxPicWidth[ ] and MaxPicHeight[ ], respectively.
[0237] - The chroma format specifier referred to as ChromaFormatIdc in this document
[0238] - Bit depth for samples of the chroma component, referred to as BitDepthY. Bit depth for samples of the two related chroma components, referred to as BitDepthC, if ChromaFormatIdc is not equal to 0.
[0239] As mentioned above, in this embodiment, the picture height and picture width of the luma sample unit can be indexed by layer and provided in the form of a list. Therefore, when multiple layers exist, information regarding the picture width and height of each layer associated with the PRI SEI message can be accurately provided.
[0240] pri_cancel_flag may indicate whether the SEI message cancels the persistence of a previous PRI SEI message in the output order applied to the current layer. A pri_cancel_flag equal to 1 may indicate that the SEI message cancels the persistence of a previous PRI SEI message in the output order applied to the current layer. A pri_cancel_flag equal to 0 may indicate that packed regions information follows.
[0241] The pri_persistence_flag specifies the persistence of PRI SEI messages for the current layer. A pri_persistence_flag equal to 0 specifies that packed regions information applies only to the currently decoded picture. A pri_persistence_flag equal to 1 specifies that PRI SEI messages apply to the currently decoded picture and persist to all subsequent pictures of the current layer in the output order until one or more of the following conditions are true:
[0242] - When a new CLVS of the current layer starts
[0243] - When the bitstream ends
[0244] - When the picture of the current layer included in the AU associated with the PRI SEI message is output, and that picture is positioned after the current picture in the output order.
[0245] pri_num_regions_minus1 can represent information related to the number of regions where information (e.g., packed region information) is signaled. A value of pri_num_regions_minus1 plus 1 can specify the number of regions where information is signaled.
[0246] pri_multilayer_flag can indicate whether the syntax element pri_region_layer_id[i] exists. pri_multilayer_flag equal to 1 can specify that pri_region_layer_id[i] exists. pri_multilayer_flag equal to 0 can specify that pri_region_layer_id[i] does not exist.
[0247] pri_use_max_dimensions_flag can specify whether MaxPicWidth, MaxPicHeight, PicWidthInLumaSamples, and PicHeightInLumaSamples are used in variable calculations. pri_use_max_dimensions_flag equal to 1 can specify that MaxPicWidth, MaxPicHeight, PicWidthInLumaSamples, and PicHeightInLumaSamples are used in variable calculations. pri_use_max_dimensions_flag equal to 0 can specify that MaxPicWidth, MaxPicHeight, PicWidthInLumaSamples, and PicHeightInLumaSamples are not used in variable calculations for area parameters.
[0248] pri_log2_unit_size can specify the unit size used for variable calculations for area parameters.
[0249] The variable priUnitSize can be set to 1 << pri_log2_unit_size.
[0250] pri_region_size_len_minus1 can represent information regarding the number of bits used by various syntax elements that define the size and location of a region within a PRI SEI message. A value of pri_region_size_len_minus1 plus 1 can specify the number of bits used to signal pri_region_top_left_in_units_x[i], pri_region_top_left_in_units_y[i], pri_region_width_in_units_minus1[i], pri_region_height_in_units_minus1[i], pri_target_region_top_left_x[i], and pri_target_region_top_left_y[i].
[0251] pri_region_id_present_flag can indicate whether the syntax element pri_region_id[i] exists. pri_region_id_present_flag equal to 1 can indicate that pri_region_id[i] exists. pri_region_id_present_flag equal to 0 can indicate that pri_region_id[i] does not exist.
[0252] pri_target_pic_params_present_flag may indicate whether syntax elements representing parameters required to restore the target picture (e.g., the total size of the target picture, area-by-area layout, etc.) exist in the PRI SEI message. pri_target_pic_params_present_flag identical to 1 may indicate that the syntax elements pri_target_region_top_left_x[i], pri_target_region_top_left_y[i], pri_target_pic_width_minus1, and pri_target_pic_height_minus1 exist. pri_target_pic_params_present_flag equal to 0 may indicate that the syntax elements pri_target_region_top_left_x[i], pri_target_region_top_left_y[i], pri_target_pic_width_minus1, and pri_target_pic_height_minus1 do not exist.
[0253] pri_target_pic_width_minus1 may represent information regarding the width of the target picture in luminance samples that can be reconstructed from samples of the cropped decoded picture corresponding to the regions described in the PRI SEI message. If pri_target_pic_width_minus1 is present, the value of pri_target_pic_width_minus1 plus 1 may represent the width of the target picture in luminance samples.
[0254] pri_target_pic_height_minus1 may represent information regarding the height of the target picture in luminance samples that can be reconstructed from samples of the cropped decoded picture corresponding to the regions described in the PRI SEI message. If pri_target_pic_height_minus1 is present, the value of pri_target_pic_height_minus1 plus 1 may represent the height of the target picture in luminance samples.
[0255] pri_num_resampling_ratios_minus1 can represent information related to the number of signaled resampling ratios. A value of pri_num_resampling_ratios_minus1 plus 1 can specify the number of signaled resampling ratios.
[0256] The value obtained by adding 1 to pri_resampling_width_num_minus1[i] and the value obtained by adding 1 to pri_resampling_width_denom_minus1[i] can respectively specify the numerator and denominator for width resampling of the i-th resampling rate. Both pri_resampling_width_num_minus1[i] and pri_resampling_width_denom_minus1[i] may be within the range of 0 to 65535 (inclusive).
[0257] If pri_resampling_width_num_minus1[i] and pri_resampling_width_denom_minus1[i] do not exist, the values of pri_resampling_ratio_width_num_minus1[0] and pri_resampling_ratio_width_denom_minus1[0] can be estimated to be equal to 0.
[0258] pri_fixed_aspect_ratio_flag[i] can specify whether the syntax elements pri_resampling_height_num_minus1[i] and pri_resampling_height_denom_minus1[i] exist. pri_fixed_aspect_ratio_flag[i] equal to 1 can specify that the syntax elements pri_resampling_height_num_minus1[i] and pri_resampling_height_denom_minus1[i] do not exist. pri_fixed_aspect_ratio_flag[i] equal to 0 can specify that pri_resampling_height_num_minus1[i] and pri_resampling_height_denom_minus1[i] exist.
[0259] The value obtained by adding 1 to pri_resampling_height_num_minus1[i] and the value obtained by adding 1 to pri_resampling_height_denom_minus1[i] can respectively specify the numerator and denominator for the height resampling of the i-th resampling rate. Both pri_resampling_height_num_minus1[i] and pri_resampling_height_denom_minus1[i] may be within the range of 0 to 65535 (inclusive).
[0260] If pri_resampling_height_num_minus1[i] and pri_resampling_height_denom_minus1[i] do not exist, the values of pri_resampling_height_num_minus1[i] and pri_resampling_height_denom_minus1[i] can be assumed to be the same as pri_resampling_width_num_minus1[i] and pri_resampling_width_denom_minus1[i], respectively.
[0261] pri_region_id[i] can represent the ID of the i-th region. If it does not exist, the value of pri_region_id[i] can be assumed to be the same as i.
[0262] pri_region_layer_id[i] may specify the layer identifier of the picture to which region information is associated, such as pri_region_top_left_in_units_x[i], pri_region_top_left_in_units_y[i], pri_region_width_in_units_minus1[i], and pri_region_height_in_units_minus1[i]. If pri_region_layer_id[i] does not exist, the value of pri_region_layer_id[i] may be assumed to be equal to 0.
[0263] pri_region_is_a_layer_flag[i] can specify whether the picture width and height in the layer with ID pri_region_layer_id[i] are equal to the width and height of the region at index i. pri_region_is_a_layer_flag[i] equal to 1 can specify that the picture width and height in the layer with the same layer identifier as pri_region_layer_id[i] are equal to the width and height of the region at index i, and pri_region_top_left_in_units_x[i], pri_region_top_left_in_units_y[i], pri_region_width_in_units_minus1[i], and pri_region_height_in_units_minus1[i] may not be signaled. If pri_region_is_a_layer_flag[i] does not exist, the value of pri_region_is_a_layer_flag[i] can be assumed to be equal to 0.
[0264] pri_region_top_left_in_units_x[i] and pri_region_top_left_in_units_y[i] can specify the horizontal and vertical positions, respectively, of the top-left sample of the i-th region expressed in units. The length of these syntax elements can be pri_region_size_len_minus1 + 1.
[0265] When i ranges from 0 to the maximum number of different layers associated with the PRI SEI message, the variable priLayerIdx[i] can represent the layer index of the layer with layer identifier i.
[0266] Variables priRegionTopLeftX[i] and priRegionTopLeftY[i], representing the horizontal and vertical positions in luminance sample units of the i-th region in a cropped decoded picture where the layer identifier is the same as pri_region_layer_id[i], respectively, can be derived as shown in the following Equation 5.
[0267] [Equation 5]
[0268]
[0269] pri_region_width_in_units_minus1[i] and pri_region_height_in_units_minus1[i] can respectively represent information regarding the width and height of the i-th region expressed in units. The value obtained by adding 1 to pri_region_width_in_units_minus1[i] and the value obtained by adding 1 to pri_region_height_in_units_minus1[i] can respectively specify the width and height of the i-th region expressed in units. The length of the above syntax elements can be pri_region_size_len_minus1 + 1.
[0270] Variables priRegionWidth[i] and priRegionHeight[i], representing the width and height of the i-th region in luminance sample units respectively in the cropped decoded picture, can be derived as shown in the following Equation 6.
[0271] [Equation 6]
[0272]
[0273] As a requirement for bitstream conformance, the value of priRegionWidth[i] % SubWidthC must be equal to 0 and the value of priRegionHeight[i] % SubHeightC must be equal to 0. In other words, the remainder when priRegionWidth[i] is divided by SubWidthC must be 0, and the remainder when priRegionHeight[i] is divided by SubHeightC must be 0. That is, priRegionWidth[i] must be a multiple of priRegionWidth[i] and priRegionHeight[i] must be a multiple of SubHeightC.
[0274] The variables SubWidthC and SubHeightC can be derived from ChromaFormatIdc as specified in Table 2 above.
[0275] pri_resampling_ratio_idx[i] may specify the index of the resampling ratio used for the i-th region. The value of pri_resampling_ratio_idx[i] must be within the range (inclusive) of 0 to pri_num_resampling_ratios_minus1. The length of the syntax element may be Ceil(Log2(pri_num_resampling_ratios_minus1 + 1)) bits.
[0276] The variables priResampleWidthNum[i], priResampleWidthDenom[i], priResampleHeightNum[i], and priResampleHeightDenom[i] can be derived as shown in Equation 7 below.
[0277] [Equation 7]
[0278]
[0279] If present, pri_target_region_top_left_x[i] and pri_target_region_top_left_y[i] may represent the horizontal and vertical positions, respectively, of the top-left sample position in luminance sample units of the i-th region in the restored target picture.
[0280] Variables priTargetRegionWidth and priTargetRegionHeight, representing the luminance sample unit width and height of the resampled region in the restored target picture, respectively, can be derived as shown in the following Equation 8.
[0281] [Equation 8]
[0282]
[0283] When restoring a target picture having a luminance sample array of size ( pri_target_pic_width_minus1 + 1 ) Х ( pri_target_pic_height_minus1 + 1 ) all luminance sample values can be initialized to 1 << ( BitDepthY - 1 ) values, and if present, chroma samples can be initialized to 1 << ( BitDepthC - 1 ) values.
[0284] For any sample location (x, y) and regions j and k, if all of the following conditions are true, the restored target picture sample at location (x, y) can be determined by the parameters signaled for the j-th region:
[0285] - When pri_region_id[j] is greater than pri_region_id[k]
[0286] - If x is within the range priRegionTopLeftX[j] .. priRegionTopLeftX[j] + priRegionWidth[j]
[0287] - If y is within the range priRegionTopLeftY[j] .. priRegionTopLeftY[j] + priRegionHeight[j]
[0288] - If x is within the range priRegionTopLeftX[k] .. priRegionTopLeftX[k] + priRegionWidth[k]
[0289] - If y is within the range priRegionTopLeftY[k] .. priRegionTopLeftY[k] + priRegionHeight[k]
[0290] For example, the following changes can be applied to the VVC interface for the VSEI specification.
[0291] The following variables may be specified for the interpretation of PRI SEI messages:
[0292] - PicWidthInLumaSamples[ ] is from all layers containing regions
[0293] pps_pic_width_in_luma_samples - SubWidthC * ( pps_conf_win_left_offset + pps_conf_win_right_offset ) can be set as a list of values, and the list can be organized in ascending order of layer ID.
[0294] - PicHeightInLumaSamples[ ] is from all layers containing regions
[0295] pps_pic_height_in_luma_samples - SubHeightC * ( pps_conf_win_top_offset + pps_conf_win_bottom_offset ) can be set as a list of values, and the list can be organized in ascending order of layer ID.
[0296] - MaxPicWidth[ ] is from all layers containing regions
[0297] sps_pic_width_in_luma_samples - SubWidthC * ( sps_conf_win_left_offset + sps_conf_win_right_offset ) can be set as a list of values, and the list can be organized in ascending order of layer ID.
[0298] - MaxPicHeight[ ] is from all layers containing regions
[0299] sps_pic_height_in_luma_samples - SubHeightC * ( sps_conf_win_top_offset + sps_conf_win_bottom_offset ) can be set as a list of values, and the list can be organized in ascending order of layer ID.
[0300] - ChromaFormatIdc can be set to the same value as sps_chroma_format_idc.
[0301] - BitDepthY and BitDepthC can both be set to the same value as BitDepth.
[0302] Meanwhile, the aforementioned embodiment may also be applied to HEVC. For the VSEI specification, the following changes may be applied to the HEVC interface.
[0303] The following variables may be specified for the interpretation of PRI SEI messages:
[0304] - PicWidthInLumaSamples[ ] is from all layers containing regions
[0305] pic_width_in_luma_samples - SubWidthC * ( conf_win_left_offset + conf_win_right_offset ) can be set as a list of values, and the list can be organized in ascending order of layer ID.
[0306] - PicHeightInLumaSamples[ ] is from all layers containing regions
[0307] pic_height_in_luma_samples - SubHeightC * ( conf_win_top_offset + conf_win_bottom_offset ) can be set as a list of values, and the list can be configured in ascending order of layer ID.
[0308] - MaxPicWidth[ ] is from all layers containing regions
[0309] pic_width_in_luma_samples - SubWidthC * ( conf_win_left_offset + conf_win_right_offset ) can be set as a list of values, and the list can be organized in ascending order of layer ID.
[0310] - MaxPicHeight[ ] is from all layers containing regions
[0311] The values of pic_height_in_luma_samples - SubHeightC * ( conf_win_top_offset + sps_conf_win_bottom_offset ) can be set, and the above list can be organized in ascending order of layer ID.
[0312] - ChromaFormatIdc can be set to the same value as chroma_format_idc.
[0313] - BitDepthY and BitDepthC can both be set to the same value as BitDepth.
[0314] The embodiments described below may also be based on VSEI and VVC standards.
[0315] As described above, the PRI SEI message may provide information regarding rectangular regions packed within the coded picture. This information may be optionally used to restore the target picture from samples of the cropped decoded picture corresponding to the regions described in the SEI message.
[0316] The use of this SEI message may require the definition of the following variables:
[0317] - Lists of picture widths and picture heights in luma samples, referred to as PicWidthInLumaSamples[ ] and PicHeightInLumaSamples[ ], respectively.
[0318] - Lists of maximum picture widths and maximum picture heights in luma samples, referred to as MaxPicWidth[ ] and MaxPicHeight[ ], respectively.
[0319] - The chroma format specifier referred to as ChromaFormatIdc in this document
[0320] - Bit depth for samples of the chroma component, referred to as BitDepthY. Bit depth for samples of the two related chroma components, referred to as BitDepthC, if ChromaFormatIdc is not equal to 0.
[0321] As mentioned above, in this embodiment, the picture height and picture width of the luma sample unit can be indexed by layer and provided in the form of a list. Therefore, when multiple layers exist, information regarding the picture width and height of each layer associated with the PRI SEI message can be accurately provided.
[0322] pri_cancel_flag may indicate whether the SEI message cancels the persistence of a previous PRI SEI message in the output order applied to the current layer. A pri_cancel_flag equal to 1 may indicate that the SEI message cancels the persistence of a previous PRI SEI message in the output order applied to the current layer. A pri_cancel_flag equal to 0 may indicate that packed regions information follows.
[0323] The pri_persistence_flag specifies the persistence of PRI SEI messages for the current layer. A pri_persistence_flag equal to 0 specifies that packed regions information applies only to the currently decoded picture. A pri_persistence_flag equal to 1 specifies that PRI SEI messages apply to the currently decoded picture and persist to all subsequent pictures of the current layer in the output order until one or more of the following conditions are true:
[0324] - When a new CLVS of the current layer starts
[0325] - When the bitstream ends
[0326] - When the picture of the current layer included in the AU associated with the PRI SEI message is output, and that picture is positioned after the current picture in the output order.
[0327] pri_num_regions_minus1 can represent information related to the number of regions where information (e.g., packed region information) is signaled. A value of pri_num_regions_minus1 plus 1 can specify the number of regions where information is signaled.
[0328] pri_multilayer_flag can indicate whether the syntax element pri_region_layer_id[i] exists. pri_multilayer_flag equal to 1 can specify that pri_region_layer_id[i] exists. pri_multilayer_flag equal to 0 can specify that pri_region_layer_id[i] does not exist.
[0329] pri_use_max_dimensions_flag can specify whether MaxPicWidth, MaxPicHeight, PicWidthInLumaSamples, and PicHeightInLumaSamples are used in variable calculations. pri_use_max_dimensions_flag equal to 1 can specify that MaxPicWidth, MaxPicHeight, PicWidthInLumaSamples, and PicHeightInLumaSamples are used in variable calculations. pri_use_max_dimensions_flag equal to 0 can specify that MaxPicWidth, MaxPicHeight, PicWidthInLumaSamples, and PicHeightInLumaSamples are not used in variable calculations for area parameters.
[0330] pri_log2_unit_size can specify the unit size used for variable calculations for area parameters.
[0331] The variable priUnitSize can be set to 1 << pri_log2_unit_size.
[0332] pri_region_size_len_minus1 can represent information regarding the number of bits used by various syntax elements that define the size and location of a region within a PRI SEI message. A value of pri_region_size_len_minus1 plus 1 can specify the number of bits used to signal pri_region_top_left_in_units_x[i], pri_region_top_left_in_units_y[i], pri_region_width_in_units_minus1[i], pri_region_height_in_units_minus1[i], pri_target_region_top_left_x[i], and pri_target_region_top_left_y[i].
[0333] pri_region_id_present_flag can indicate whether the syntax element pri_region_id[i] exists. pri_region_id_present_flag equal to 1 can indicate that pri_region_id[i] exists. pri_region_id_present_flag equal to 0 can indicate that pri_region_id[i] does not exist.
[0334] pri_target_pic_params_present_flag may indicate whether syntax elements representing parameters required to restore the target picture (e.g., the total size of the target picture, area-by-area layout, etc.) exist in the PRI SEI message. pri_target_pic_params_present_flag identical to 1 may indicate that the syntax elements pri_target_region_top_left_x[i], pri_target_region_top_left_y[i], pri_target_pic_width_minus1, and pri_target_pic_height_minus1 exist. pri_target_pic_params_present_flag equal to 0 may indicate that the syntax elements pri_target_region_top_left_x[i], pri_target_region_top_left_y[i], pri_target_pic_width_minus1, and pri_target_pic_height_minus1 do not exist.
[0335] pri_target_pic_width_minus1 may represent information regarding the width of the target picture in luminance samples that can be reconstructed from samples of the cropped decoded picture corresponding to the regions described in the PRI SEI message. If pri_target_pic_width_minus1 is present, the value of pri_target_pic_width_minus1 plus 1 may represent the width of the target picture in luminance samples.
[0336] pri_target_pic_height_minus1 may represent information regarding the height of the target picture in luminance samples that can be reconstructed from samples of the cropped decoded picture corresponding to the regions described in the PRI SEI message. If pri_target_pic_height_minus1 is present, the value of pri_target_pic_height_minus1 plus 1 may represent the height of the target picture in luminance samples.
[0337] pri_num_resampling_ratios_minus1 can represent information related to the number of signaled resampling ratios. A value of pri_num_resampling_ratios_minus1 plus 1 can specify the number of signaled resampling ratios.
[0338] The value obtained by adding 1 to pri_resampling_width_num_minus1[i] and the value obtained by adding 1 to pri_resampling_width_denom_minus1[i] can respectively specify the numerator and denominator for width resampling of the i-th resampling rate. Both pri_resampling_width_num_minus1[i] and pri_resampling_width_denom_minus1[i] may be within the range of 0 to 65535 (inclusive).
[0339] If pri_resampling_width_num_minus1[i] and pri_resampling_width_denom_minus1[i] do not exist, the values of pri_resampling_ratio_width_num_minus1[0] and pri_resampling_ratio_width_denom_minus1[0] can be estimated to be equal to 0.
[0340] pri_fixed_aspect_ratio_flag[i] can specify whether the syntax elements pri_resampling_height_num_minus1[i] and pri_resampling_height_denom_minus1[i] exist. pri_fixed_aspect_ratio_flag[i] equal to 1 can specify that the syntax elements pri_resampling_height_num_minus1[i] and pri_resampling_height_denom_minus1[i] do not exist. pri_fixed_aspect_ratio_flag[i] equal to 0 can specify that pri_resampling_height_num_minus1[i] and pri_resampling_height_denom_minus1[i] exist.
[0341] The value obtained by adding 1 to pri_resampling_height_num_minus1[i] and the value obtained by adding 1 to pri_resampling_height_denom_minus1[i] can respectively specify the numerator and denominator for the height resampling of the i-th resampling rate. Both pri_resampling_height_num_minus1[i] and pri_resampling_height_denom_minus1[i] may be within the range of 0 to 65535 (inclusive).
[0342] If pri_resampling_height_num_minus1[i] and pri_resampling_height_denom_minus1[i] do not exist, the values of pri_resampling_height_num_minus1[i] and pri_resampling_height_denom_minus1[i] can be assumed to be the same as pri_resampling_width_num_minus1[i] and pri_resampling_width_denom_minus1[i], respectively.
[0343] pri_region_id[i] can represent the ID of the i-th region. If it does not exist, the value of pri_region_id[i] can be assumed to be the same as i.
[0344] pri_region_layer_id[i] may specify the layer identifier of the picture to which region information is associated, such as pri_region_top_left_in_units_x[i], pri_region_top_left_in_units_y[i], pri_region_width_in_units_minus1[i], and pri_region_height_in_units_minus1[i]. If pri_region_layer_id[i] does not exist, the value of pri_region_layer_id[i] may be assumed to be equal to 0.
[0345] pri_region_is_a_layer_flag[i] can specify whether the picture width and height in the layer with ID pri_region_layer_id[i] are equal to the width and height of the region at index i. pri_region_is_a_layer_flag[i] equal to 1 can specify that the picture width and height in the layer with the same layer identifier as pri_region_layer_id[i] are equal to the width and height of the region at index i, and pri_region_top_left_in_units_x[i], pri_region_top_left_in_units_y[i], pri_region_width_in_units_minus1[i], and pri_region_height_in_units_minus1[i] may not be signaled. If pri_region_is_a_layer_flag[i] does not exist, the value of pri_region_is_a_layer_flag[i] can be assumed to be equal to 0.
[0346] pri_region_top_left_in_units_x[i] and pri_region_top_left_in_units_y[i] can specify the horizontal and vertical positions, respectively, of the top-left sample of the i-th region expressed in units. The length of these syntax elements can be pri_region_size_len_minus1 + 1.
[0347] Variables priRegionTopLeftX[i] and priRegionTopLeftY[i], representing the horizontal and vertical positions in luminance sample units of the i-th region in a cropped decoded picture with the same layer identifier as pri_region_layer_id[i], respectively, can be derived as shown in the following Equation 9.
[0348] [Equation 9]
[0349]
[0350] Referring to Equation 9 above, it can be seen that picture size variables representing information related to the size of the picture are indexed by the picture layer identifier pri_region_layer_id[ i ] related to the information of the rectangular area as follows:
[0351] -PicWidthInLumaSamples[pri_region_layer_id[i]]
[0352] -PicHeightInLumaSamples[pri_region_layer_id[i]]
[0353] -MaxPicWidth[pri_region_layer_id[i]]
[0354] -MaxPicHeight[pri_region_layer_id[i]]
[0355] pri_region_width_in_units_minus1[i] and pri_region_height_in_units_minus1[i] can respectively represent information regarding the width and height of the i-th region expressed in units. The value obtained by adding 1 to pri_region_width_in_units_minus1[i] and the value obtained by adding 1 to pri_region_height_in_units_minus1[i] can respectively specify the width and height of the i-th region expressed in units. The length of the above syntax elements can be pri_region_size_len_minus1 + 1.
[0356] Variables priRegionWidth[i] and priRegionHeight[i], representing the width and height of the i-th region in luma sample units respectively in the cropped decoded picture, can be derived as shown in the following Equation 10.
[0357] [Equation 10]
[0358]
[0359] As a requirement for bitstream conformance, the value of priRegionWidth[i] % SubWidthC must be equal to 0 and the value of priRegionHeight[i] % SubHeightC must be equal to 0. In other words, when priRegionWidth[i] is divided by SubWidthC, the remainder must be 0, and when priRegionHeight[i] is divided by SubHeightC, the remainder must be 0. That is, priRegionWidth[i] must be a multiple of SubWidthC and priRegionHeight[i] must be a multiple of SubHeightC.
[0360] The variables SubWidthC and SubHeightC can be derived from ChromaFormatIdc as specified in Table 2 above.
[0361] pri_resampling_ratio_idx[i] may specify the index of the resampling ratio used for the i-th region. The value of pri_resampling_ratio_idx[i] must be within the range (inclusive) of 0 to pri_num_resampling_ratios_minus1. The length of the syntax element may be Ceil(Log2(pri_num_resampling_ratios_minus1 + 1)) bits.
[0362] The variables priResampleWidthNum[i], priResampleWidthDenom[i], priResampleHeightNum[i], and priResampleHeightDenom[i] can be derived as shown in the following Equation 11.
[0363] [Equation 11]
[0364]
[0365] If present, pri_target_region_top_left_x[i] and pri_target_region_top_left_y[i] may represent the horizontal and vertical positions, respectively, of the top-left sample position in luminance sample units of the i-th region in the restored target picture.
[0366] Variables priTargetRegionWidth and priTargetRegionHeight, representing the luminance sample unit width and height of the resampled region in the restored target picture, respectively, can be derived as shown in the following Equation 12.
[0367] [Equation 12]
[0368]
[0369] When restoring a target picture having a luminance sample array of size ( pri_target_pic_width_minus1 + 1 ) Х ( pri_target_pic_height_minus1 + 1 ) all luminance sample values can be initialized to 1 << ( BitDepthY - 1 ) values, and if present, chroma samples can be initialized to 1 << ( BitDepthC - 1 ) values.
[0370] For any sample location (x, y) and regions j and k, if all of the following conditions are true, the restored target picture sample at location (x, y) can be determined by the parameters signaled for the j-th region:
[0371] - When pri_region_id[j] is greater than pri_region_id[k]
[0372] - If x is within the range priRegionTopLeftX[j] .. priRegionTopLeftX[j] + priRegionWidth[j]
[0373] - If y is within the range priRegionTopLeftY[j] .. priRegionTopLeftY[j] + priRegionHeight[j]
[0374] - If x is within the range priRegionTopLeftX[k] .. priRegionTopLeftX[k] + priRegionWidth[k]
[0375] - If y is within the range priRegionTopLeftY[k] .. priRegionTopLeftY[k] + priRegionHeight[k]
[0376] Meanwhile, the following changes can be applied to the VVC interface for the VSEI standard.
[0377] The following variables may be specified for the interpretation of PRI SEI messages:
[0378] - PicWidthInLumaSamples[ lId ] can be set to a list of values of pps_pic_width_in_luma_samples - SubWidthC * ( pps_conf_win_left_offset + pps_conf_win_right_offset ) for pictures where nuh_layer_id is the same as lId.
[0379] - PicHeightInLumaSamples[ lId ] can be a list of pps_pic_height_in_luma_samples - SubHeightC * ( pps_conf_win_top_offset + pps_conf_win_bottom_offset ) values for pictures where nuh_layer_id is the same as lId.
[0380] - MaxPicWidth[ lId ] can be set to a list of sps_pic_width_in_luma_samples - SubWidthC * ( sps_conf_win_left_offset + sps_conf_win_right_offset ) values for pictures where nuh_layer_id is the same as lId.
[0381] - MaxPicHeight[ lId ] can be set to a list of sps_pic_height_in_luma_samples - SubHeightC * ( sps_conf_win_top_offset + sps_conf_win_bottom_offset ) values for pictures where nuh_layer_id is the same as lId.
[0382] - ChromaFormatIdc can be set to be the same as sps_chroma_format_idc.
[0383] - BitDepthY and BitDepthC can both be set to be the same as BitDepth.
[0384] In the description of the above variables, nuh_layer_id may represent the NAL unit header layer ID. This information may be included in the header of the NAL unit and can identify which layer the NAL unit belongs to in a multi-layer bitstream environment.
[0385] Meanwhile, the aforementioned embodiment may also be applied to HEVC. For the VSEI specification, the following changes may be applied to the HEVC interface.
[0386] The following variables may be specified for the interpretation of PRI SEI messages:
[0387] - PicWidthInLumaSamples[ lId ] can be set to a list of pic_width_in_luma_samples - SubWidthC * ( conf_win_left_offset + conf_win_right_offset ) values for pictures where nuh_layer_id is the same as lId.
[0388] - PicHeightInLumaSamples[ lId ] can be a list of pic_height_in_luma_samples - SubHeightC * ( conf_win_top_offset + conf_win_bottom_offset ) values for pictures where nuh_layer_id is the same as lId.
[0389] - MaxPicWidth[ lId ] can be set to a list of pic_width_in_luma_samples - SubWidthC * ( conf_win_left_offset + conf_win_right_offset ) values for pictures where nuh_layer_id is the same as lId.
[0390] - MaxPicHeight[ lId ] can be set to a list of pic_height_in_luma_samples - SubHeightC * ( conf_win_top_offset + conf_win_bottom_offset ) values for pictures where nuh_layer_id is the same as lId.
[0391] - ChromaFormatIdc can be set to be the same as chroma_format_idc.
[0392] - BitDepthY and BitDepthC can both be set to be the same as BitDepth.
[0393] In the above-described embodiment, a variable related to the size of the picture was indexed as the layer identifier pri_region_layer_id[ i ] so that source layer information for each region could be accurately referenced even in an environment where multiple layers with different resolutions exist. Through this, when restoring the target picture, the actual resolution and crop offset information of the layer to which each region belongs can be individually mapped and calculated. In addition, coordinate calculation errors that occur when combining image fragments (rectangular regions) of different layers into a single final screen can be prevented.
[0394] FIGS. 7 to 9 are drawings illustrating a method for decoding image information according to one embodiment.
[0395] A method according to one embodiment may be performed by a decoding device (300) according to one embodiment. Accordingly, all or part of the description of the decoding device (300) described above and the description of the decoding method described with reference to FIG. 4 may also be applied to a method according to one embodiment. That is, the descriptions described above may also be applied to a method according to one embodiment to the extent that they do not conflict with the descriptions described below.
[0396] The decoding device (300) may include a memory and a processor electrically connected to the memory, and the operation of the decoding device (300) described above, the decoding method, or the method described below may be executed by the processor of the decoding device (300).
[0397] Terms or names used in this disclosure (e.g., names of syntax elements or names of variables, etc.) are merely examples, and the scope of the embodiments is not limited to these terms. Even if a term is not used in this disclosure, if substantial features such as the function performed, the definition thereof, or the method by which it is derived are identical or similar to those of this disclosure, it may be considered to be included within the scope of the embodiments described in this disclosure.
[0398] In addition, the method according to one embodiment may include other operations in addition to the operations described below, and some of the operations described below may be omitted depending on the example.
[0399] Referring to FIG. 7, a method according to one embodiment may include the step (S1010) of obtaining a Packed Regions Information (PRI) related message from a bitstream that provides information about rectangular regions packed within a picture of one or more layers, and the step (S1020) of deriving a picture size variable representing information about the size of the picture for each of the one or more layers based on the PRI related message.
[0400] The above picture size variable can be defined for each of the one or more layers.
[0401] The PRI-related message obtained in step S1010 may be the aforementioned PRI SEI message. Accordingly, all or part of the description regarding the aforementioned PRI SEI message may be applied to the method according to one embodiment. Syntax elements included in the previously described PRI SEI message and related variables may be obtained or derived by the method according to one embodiment, and the description regarding the syntax elements included in the previously described PRI SEI message and related variables may be applied to the method according to one embodiment without separate mention.
[0402] The above picture size variable can be indexed by a layer identifier representing a specific layer. For example, the above picture size variable can be indexed by a layer identifier of a picture associated with information regarding the rectangular areas.
[0403] The above picture size variable may include at least one of a picture width variable representing the width of the picture or a picture height variable representing the height of the picture. The picture width variable or the picture height variable may be defined in units of luma samples.
[0404] Additionally, the picture size variable may include at least one of a picture maximum width variable representing the maximum width of the picture or a picture maximum height variable representing the maximum height of the picture. The picture maximum width variable or the picture maximum height variable may be defined in units of luma samples.
[0405] That is, the above picture size variable may include at least one of the following variables:
[0406] -PicWidthInLumaSamples[pri_region_layer_id[i]]
[0407] -PicHeightInLumaSamples[pri_region_layer_id[i]]
[0408] -MaxPicWidth[pri_region_layer_id[i]]
[0409] -MaxPicHeight[pri_region_layer_id[i]]
[0410] PicWidthInLumaSamples[pri_region_layer_id[i]] represents the picture width in luma samples, and since it is indexed as pri_region_layer_id[i], it can be seen that it is the width of the cropped decoded picture with the same layer identifier as pri_region_layer_id[i].
[0411] For example, PicWidthInLumaSamples[pri_region_layer_id[i]] can be derived based on Equation 13 below.
[0412] [Equation 13]
[0413] PicWidthInLumaSamples[] = pps_pic_width_in_luma_samples - SubWidthC * ( pps_conf_win_left_offset + pps_conf_win_right_offset )
[0414] Here, pps_pic_width_in_luma_samples represents the width of each decoded picture referencing PPS in units of luma samples, and pps_conf_win_left_offset and pps_conf_win_right_offset specify samples of pictures within CLVS output from the decoding process in terms of a rectangular area specified by picture coordinates for the output. That is, by the above Equation 13, the width of the actual effective image area (cropped image area) with padding or unnecessary borders removed can be derived by removing the width of the left and right margins from the total width of the picture.
[0415] The above Equation 13 is merely an example of a method for deriving PicWidthInLumaSamples[pri_region_layer_id[i]], and the scope of this embodiment is not limited thereto. It is also possible to derive it by a method other than Equation 13.
[0416] PicHeightInLumaSamples[pri_region_layer_id[i]] represents the picture height in luma samples, and since it is indexed as pri_region_layer_id[i], it can be seen that it is the height of the cropped decoded picture with the same layer identifier as pri_region_layer_id[i].
[0417] For example, PicHeightInLumaSamples[pri_region_layer_id[i]] can be derived based on Equation 14 below.
[0418] [Equation 14]
[0419] PicHeightInLumaSamples[] = pps_pic_height_in_luma_samples - SubHeightC * (pps_conf_win_top_offset + pps_conf_win_bottom_offset)
[0420] Here, pps_pic_height_in_luma_samples represents the height of each decoded picture referencing PPS in units of luma samples, and pps_conf_win_top_offset and pps_conf_win_bottom_offset specify samples of pictures within CLVS output from the decoding process in terms of a rectangular area specified by picture coordinates for the output. That is, by the above Equation 14, the height of the actual effective image area (cropped image area) with padding or unnecessary borders removed can be derived by removing the height of the top and bottom margins from the total height of the picture.
[0421] The above Equation 14 is merely an example of a method for deriving PicHeightInLumaSamples[pri_region_layer_id[i]], and the scope of this embodiment is not limited thereto. It is also possible to derive it by a method other than Equation 14.
[0422] MaxPicWidth[pri_region_layer_id[i]] represents the maximum picture width in luma samples, and since it is indexed as pri_region_layer_id[i], it can be seen that it is the maximum width of the cropped decoded picture with the same layer identifier as pri_region_layer_id[i].
[0423] For example, MaxPicWidth[pri_region_layer_id[i]] can be derived based on Equation 15 below.
[0424] [Equation 15]
[0425] MaxPicWidth[] = sps_pic_width_in_luma_samples - SubWidthC * (sps_conf_win_left_offset + sps_conf_win_right_offset)
[0426] Here, sps_pic_width_in_luma_samples can specify the maximum width in luma samples for each decoded picture referencing the SPS. sps_conf_win_left_offset and sps_conf_win_right_offset specify the cropping window applied to pictures where pps_pic_width_in_luma_samples is equal to sps_pic_width_max_in_luma_samples and pps_pic_height_in_luma_samples is equal to sps_pic_height_max_in_luma_samples. If sps_conformance_window_flag is equal to 0, the values of sps_conf_win_left_offset, sps_conf_win_right_offset, sps_conf_win_top_offset, and sps_conf_win_bottom_offset are inferred to be equal to 0.
[0427] The above Equation 15 is merely an example of a method for deriving MaxPicWidth[pri_region_layer_id[i]], and the scope of this embodiment is not limited thereto. It is also possible to derive it by a method other than Equation 15.
[0428] MaxPicHeight[pri_region_layer_id[i]] represents the maximum picture height in luma samples, and since it is indexed as pri_region_layer_id[i], it can be seen that it is the maximum height of the cropped decoded picture with the same layer identifier as pri_region_layer_id[i].
[0429] For example, MaxPicHeight[pri_region_layer_id[i]] can be derived based on Equation 16 below.
[0430] [Equation 16]
[0431] MaxPicHeight[] = sps_pic_height_in_luma_samples - SubHeightC * ( sps_conf_win_top_offset + sps_conf_win_bottom_offset )
[0432] Here, sps_pic_height_in_luma_samples can specify the maximum height in luma samples for each decoded picture referencing the SPS. sps_conf_win_top_offset and sps_conf_win_bottom_offset specify the cropping window applied to pictures where pps_pic_width_in_luma_samples is equal to sps_pic_width_max_in_luma_samples and pps_pic_height_in_luma_samples is equal to sps_pic_height_max_in_luma_samples. If sps_conformance_window_flag is equal to 0, the values of sps_conf_win_left_offset, sps_conf_win_right_offset, sps_conf_win_top_offset, and sps_conf_win_bottom_offset are inferred to be equal to 0.
[0433] Referring to FIG. 8, a method according to one embodiment may further include the step (S1030) of deriving a rectangular position variable representing the position of the rectangular area within the picture.
[0434] The above rectangle position variable can be derived based on the above picture size variable indexed by the layer identifier of the picture associated with the information regarding the above rectangle area.
[0435] The above rectangular position variable may include at least one of the variables priRegionTopLeftX[i] or priRegionTopLeftY[i], which respectively represent the horizontal position and vertical position in luminance sample units of the i-th region in a cropped decoded picture where the layer identifier is the same as pri_region_layer_id[i]. For example, the variables priRegionTopLeftX[i] or priRegionTopLeftY[i] can be derived based on the aforementioned Equation 9.
[0436] Referring again to Equation 9, if the value of the syntax element pri_use_max_dimensions_flag, which specifies whether the picture size variable is used in the calculation of the variables, is false (e.g., 0), the rectangle position variable priRegionTopLeftX[ i ] can be derived by multiplying the syntax element pri_region_top_left_in_units_x[i], which specifies the horizontal position of the top-left sample of the i-th region expressed in units, by the variable priUnitSize, which represents the basic unit size. The rectangle position variable priRegionTopLeftY[ i ] can be derived by multiplying the syntax element pri_region_top_left_in_units_y[i], which specifies the vertical position of the top-left sample of the i-th region expressed in units, by the variable priUnitSize, which represents the basic unit size.
[0437] If the value of the above syntax element pri_use_max_dimensions_flag is true (e.g., 1), the rectangle position variable can be derived based on the picture size variables and the syntax elements pri_region_top_left_in_units_x[i] and pri_region_top_left_in_units_x[i], which respectively specify the horizontal and vertical positions of the top-left sample of the i-th region expressed in units of the above picture size variable.
[0438] Referring to FIG. 9, a method according to one embodiment may further include the step (S1040) of deriving a rectangular size variable representing the size of the rectangular area within the picture.
[0439] The above rectangle size variable can be derived based on the picture size variable indexed by the layer identifier of the picture associated with the information regarding the above rectangle area.
[0440] The above rectangular size variable may include at least one of the variables priRegionWidth[i] or priRegionHeight[i], which respectively represent the width and vertical height in luminance sample units of the i-th region in a cropped decoded picture where the layer identifier is the same as pri_region_layer_id[i]. For example, the variables priRegionWidth[i] or priRegionHeight[i] may be derived based on the aforementioned Equation 10.
[0441] Referring again to Equation 10, if the value of the syntax element pri_use_max_dimensions_flag, which specifies whether the picture size variable is used in the calculation of the variables, is false (e.g., 0), the rectangle size variable priRegionWidth[ i ] can be derived by adding 1 to the syntax element pri_region_width_in_units_minus1[i], which is related to the width of the i-th region in units, and multiplying it by the variable priUnitSize, which represents the basic unit size. The rectangle size variable priRegionHeight[ i ] can be derived by adding 1 to the syntax element pri_region_height_in_units_minus1[i], which is related to the height of the i-th region in units, and multiplying it by the variable priUnitSize, which represents the basic unit size.
[0442] If the value of the above syntax element pri_use_max_dimensions_flag is true (e.g., 1), it can be derived based on picture size variables and syntax elements representing information about the width and height of the i-th area expressed in units.
[0443] FIGS. 10 and FIGS. 11 are drawings illustrating a method for encoding image information according to one embodiment.
[0444] The method according to one embodiment may be performed by an encoding device (200) according to one embodiment. Accordingly, all or part of the description of the encoding device (200) described above and the description of the encoding method described with reference to FIG. 5 may also be applied to the method according to one embodiment. That is, the descriptions described above may also be applied to the method according to one embodiment to the extent that they do not conflict with the descriptions described below.
[0445] The encoding device (200) may include a memory and a processor electrically connected to the memory, and the operation of the aforementioned encoding device (200), the encoding method, or the method described below may be executed by the processor of the encoding device (200).
[0446] Terms or names used in this disclosure (e.g., names of syntax elements or names of variables, etc.) are merely examples, and the scope of the embodiments is not limited to these terms. Even if a term is not used in this disclosure, if substantial features such as the function performed, the definition thereof, or the method by which it is derived are identical or similar to those of this disclosure, it may be considered to be included within the scope of the embodiments described in this disclosure.
[0447] In addition, the method according to one embodiment may include other operations in addition to the operations described below, and some of the operations described below may be omitted depending on the example.
[0448] Referring to FIG. 10, a method for encoding image information according to one embodiment may include the step of generating a Packed Regions Information (PRI) related message that provides information about rectangular regions packed within a picture of one or more layers (S1110), and the step of encoding image information including the generated PRI related message (S1120). The PRI related message may be a PRI SEI message.
[0449] The PRI-related message generated in step S1110 may be the aforementioned PRI SEI message. Accordingly, all or part of the description regarding the aforementioned PRI SEI message may be applied to the method according to one embodiment. Syntax elements included in the previously described PRI SEI message and related variables may be generated or derived by the method according to one embodiment, and the description regarding the syntax elements included in the previously described PRI SEI message and related variables may be applied to the method according to one embodiment without separate mention.
[0450] When image information is encoded in the form of a bitstream by a method according to one embodiment, this bitstream can be transmitted to a decoding device by a transmission unit or a transmission device. That is, the syntax elements used in the method described above based on FIGS. 7 to 9 can be generated by the method according to the embodiment.
[0451] In addition, as previously explained, the encoding device (200) and the decoding device (300) can perform corresponding operations. Accordingly, all or part of the contents described above with reference to FIGS. 7 to 9 may also be applied to the image encoding method according to the embodiment.
[0452] Referring to FIG. 11, the step of generating the PRI-related message (S1110) may include the step of deriving a picture size variable representing information related to the size of the picture for each of the one or more layers (S1111), and the step of generating information related to the location or size of the rectangular areas based on the picture size variable (S1112).
[0453] The above PRI-related message includes information related to the location or size of the above rectangular areas, and
[0454] The above picture size variable can be defined for each of the one or more layers.
[0455] The above picture size variable can be indexed by a layer identifier representing a specific layer. For example, the above picture size variable can be indexed by a layer identifier of a picture associated with information regarding the rectangular areas.
[0456] The above picture size variable may include at least one of a picture width variable representing the width of the picture or a picture height variable representing the height of the picture. The picture width variable or the picture height variable may be defined in units of luma samples.
[0457] Additionally, the picture size variable may include at least one of a picture maximum width variable representing the maximum width of the picture or a picture maximum height variable representing the maximum height of the picture. The picture maximum width variable or the picture maximum height variable may be defined in units of luma samples.
[0458] That is, the above picture size variable may include at least one of the following variables:
[0459] -PicWidthInLumaSamples[pri_region_layer_id[i]]
[0460] -PicHeightInLumaSamples[pri_region_layer_id[i]]
[0461] -MaxPicWidth[pri_region_layer_id[i]]
[0462] -MaxPicHeight[pri_region_layer_id[i]]
[0463] PicWidthInLumaSamples[pri_region_layer_id[i]] represents the picture width in luma samples, and since it is indexed as pri_region_layer_id[i], it can be seen that it is the width of the cropped decoded picture with the same layer identifier as pri_region_layer_id[i].
[0464] PicHeightInLumaSamples[pri_region_layer_id[i]] represents the picture height in luma samples, and since it is indexed as pri_region_layer_id[i], it can be seen that it is the height of the cropped decoded picture with the same layer identifier as pri_region_layer_id[i].
[0465] MaxPicWidth[pri_region_layer_id[i]] represents the maximum picture width in luma samples, and since it is indexed as pri_region_layer_id[i], it can be seen that it is the maximum width of the cropped decoded picture with the same layer identifier as pri_region_layer_id[i].
[0466] MaxPicHeight[pri_region_layer_id[i]] represents the maximum picture height in luma samples, and since it is indexed as pri_region_layer_id[i], it can be seen that it is the maximum height of the cropped decoded picture with the same layer identifier as pri_region_layer_id[i].
[0467] In the decoding device (300), a rectangle position variable and a rectangle size variable can be derived based on the aforementioned Equations 9 and 10. Correspondingly, the encoding device (200) can derive or generate information (or syntax elements) related to the position of the rectangle region, pri_region_top_left_in_units_x[ i ] and pri_region_top_left_in_units_y[ i ], based on the picture size variable, and can derive or generate information (or syntax elements) related to the size of the rectangle region, pri_region_width_in_units_minus1[ i ] and pri_region_height_in_units_minus1[ i ].
[0468] Information (or syntax elements) related to the location of the rectangular area, pri_region_top_left_in_units_x[ i ] and pri_region_top_left_in_units_y[ i ], and information (or syntax elements) related to the size of the rectangular area, pri_region_width_in_units_minus1[ i ] and pri_region_height_in_units_minus1[ i ], may be included in the PRI SEI message. The encoding device (200) may encode image information containing the PRI SEI message into a bitstream form.
[0469] That is, a bitstream can be generated by the aforementioned method, and the bitstream can be stored non-transiently on a computer-readable storage medium.
[0470] Additionally, a method for transmitting a bitstream according to one embodiment may include the step of generating a bitstream and the step of transmitting data including said bitstream. Here, the bitstream may be generated based on the method described above. A method for transmitting a bitstream according to one embodiment may be performed by a transmission device, and the transmission device may include at least one processor for generating a bitstream and at least one memory for transmitting the generated bitstream. At least one processor and at least one memory may be electrically connected.
[0471] According to the aforementioned embodiment, variables related to the size of a picture can be indexed as layer identifiers so that source layer information for each region can be accurately referenced even in an environment where multiple layers with different resolutions exist. Through this, information of all layers can be sufficiently represented, and variables can be accurately calculated by reflecting the actual resolution of the layer to which each region belongs when restoring the target picture. In addition, coordinate calculation errors that occur when combining image fragments (rectangular regions) of different layers into a single final screen can be prevented.
[0472] FIG. 12 is a drawing illustrating an exemplary content streaming system to which an embodiment according to the present disclosure can be applied.
[0473] As illustrated in FIG. 12, a content streaming system to which an embodiment of the present disclosure is applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0474] The above encoding server compresses content input from multimedia input devices, such as smartphones, cameras, and camcorders, into digital data to generate a bitstream and transmits it to the streaming server. As another example, if multimedia input devices, such as smartphones, cameras, and camcorders, generate the bitstream directly, the encoding server may be omitted.
[0475] The bitstream may be generated by a video encoding method and / or encoding device to which an embodiment of the present disclosure is applied, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0476] The streaming server transmits multimedia data to a user device based on a user request through a web server, and the web server can act as a medium to inform the user of available services. When a user requests a desired service from the web server, the web server transmits it to the streaming server, and the streaming server can transmit multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server can perform the role of controlling commands and responses between each device within the content streaming system.
[0477] The streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, to provide a seamless streaming service, the streaming server can store the bitstream for a certain period of time.
[0478] Examples of the above user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs (head-mounted displays)), digital TVs, desktop computers, digital signage, etc.
[0479] Each server within the above-mentioned content streaming system can be operated as a distributed server, and in this case, data received from each server can be processed in a distributed manner.
[0480] The scope of the present disclosure includes software or machine-executable instructions (e.g., operating system, application, firmware, program, etc.) that enable an operation according to a method of various embodiments to be executed on a device or computer, and a non-transitory computer-readable medium on which such software or instructions, etc. are stored and executable on a device or computer.
[0481] An embodiment according to the present disclosure can be used to encode / decode images.
Claims
1. A step of obtaining a Packed Regions Information (PRI) related message from a bitstream that provides information about rectangular regions packed within a picture of one or more layers; and Based on the above PRI-related message, the method includes the step of deriving a picture size variable representing information related to the size of the picture for each of the one or more layers; The above picture size variable is defined for each of the above one or more layers, in a method.
2. In Paragraph 1, A method in which the above picture size variable is indexed by a layer identifier representing a specific layer.
3. In Paragraph 1, The above picture size variable is, A method indexed by a layer identifier of a picture associated with information regarding the above rectangular areas.
4. In Paragraph 1, The above picture size variable is, A method comprising at least one of a picture width variable representing the width of the picture or a picture height variable representing the height of the picture.
5. In Paragraph 4, The above picture width variable or the above picture height variable is defined in units of luma samples, a method.
6. In Paragraph 1, The above picture size variable is, A method comprising at least one of a picture maximum width variable representing the maximum width of the picture or a picture maximum height variable representing the maximum height of the picture.
7. In Paragraph 6, The above picture maximum width variable or the above picture maximum height variable is a method defined in units of luma samples.
8. In Paragraph 1, The method further includes the step of deriving a rectangular position variable representing the position of the rectangular area within the picture above, and A method in which the above-mentioned rectangle position variable is derived based on the above-mentioned picture size variable indexed by the layer identifier of the picture associated with the information regarding the above-mentioned rectangle area.
9. In Paragraph 1, The method further includes the step of deriving a rectangular size variable representing the size of the rectangular area within the picture above, and A method in which the above-mentioned rectangle size variable is derived based on the above-mentioned picture size variable indexed by the layer identifier of the picture associated with the information regarding the above-mentioned rectangle area.
10. A step of generating a Packed Regions Information (PRI) message that provides information about rectangular regions packed within a picture of one or more layers; and The method includes the step of encoding image information containing the generated PRI-related message; and The step of generating the above PRI-related message is, A step of deriving a picture size variable representing information related to the size of the picture for each of the above one or more layers; and The method includes the step of generating information related to the position or size of the rectangular areas based on the above picture size variable, and The above PRI-related message includes information related to the location or size of the above rectangular areas, and The above picture size variable is defined for each of the above one or more layers, in a method.
11. In Paragraph 10, A method in which the above picture size variable is indexed by a layer identifier representing a specific layer.
12. In Paragraph 10, The above picture size variable is, A method indexed by a layer identifier of a picture associated with information regarding the above rectangular areas.
13. In Paragraph 10, The above picture size variable is, A method comprising at least one of a picture width variable representing the width of the picture, a maximum picture width variable representing the maximum width of the picture, a picture height variable representing the height of the picture, or a picture maximum height variable representing the maximum height of the picture.
14. A computer-readable storage medium that non-transiently stores a bitstream generated by the method of claim 10 above.
15. Step of generating a bitstream; and The method includes the step of transmitting data including the bitstream above; and The step of generating the above bitstream is, A step of generating a Packed Regions Information (PRI) message that provides information about rectangular regions packed within a picture of one or more layers; and The method includes the step of encoding image information containing the generated PRI-related message; and The step of generating the above PRI-related message is, A step of deriving a picture size variable representing information related to the size of the picture for each of the above one or more layers; and The method includes the step of generating information related to the position or size of the rectangular areas based on the above picture size variable, and The above PRI-related message includes information related to the location or size of the above rectangular areas, and The above picture size variable is defined for each of the above one or more layers, in a method.