Image encoding / decoding method and device based on wrap-around motion compensation, and recording medium storing bitstream

The image encoding/decoding method and apparatus address the challenge of efficiently compressing high-resolution images by employing wrap-around motion compensation, resulting in improved encoding/decoding efficiency and reduced costs.

JP2025092737AInactive Publication Date: 2025-06-19LG ELECTRONICS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025061796
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-04-14
Filing Date
2025-04-03
Publication Date
2025-06-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The increasing demand for high-resolution, high-quality images has led to a need for highly efficient image compression techniques to reduce transmission and storage costs.

Method used

An image encoding/decoding method and apparatus based on wrap-around motion compensation, which includes determining whether to apply wrap-around motion compensation to a current block, generating a predicted block using inter-prediction, and encoding the inter-prediction information and wrap-around information.

Benefits of technology

This approach improves encoding/decoding efficiency and enables efficient transmission and storage of high-resolution images by effectively utilizing wrap-around motion compensation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025092737000001_ABST
    Figure 2025092737000001_ABST
Patent Text Reader

Abstract

To provide an image encoding / decoding method and device.SOLUTION: An image decoding method is provided, comprising the steps of: obtaining inter prediction information of a current block and wraparound information from a bitstream; and generating a prediction block of the current block on the basis of the inter prediction information and the wraparound information. The wraparound information may comprise a first flag indicating whether wraparound motion compensation is enabled for a current picture including the current block. On the basis of a fact that the first flag having a first value indicating that the wraparound motion compensation is enabled, the prediction block of the current block may be generated by the wraparound motion compensation, and the wraparound motion compensation may be performed on the basis of either boundaries of a current subpicture or boundaries of a reference picture of the current block, on the basis of whether the current subpicture including the current block is independently coded.SELECTED DRAWING: Figure 20
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly, to an image encoding / decoding method and apparatus based on wrap-around motion compensation, and a recording medium storing a bitstream generated by the image encoding method / apparatus of the present disclosure.

Background Art

[0002] Recently, the demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, has been increasing in various fields. As the image data becomes higher in resolution and quality, the amount of information or bits to be transmitted increases relatively compared to conventional image data. The increase in the amount of information or bits to be transmitted results in an increase in transmission costs and storage costs.

[0003] Accordingly, there is a need for a highly efficient image compression technique for effectively transmitting, storing, and reproducing information of high-resolution, high-quality images.

Summary of the Invention

Problems to be Solved by the Invention

[0004] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0005] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus based on wrap-around motion compensation.

[0006] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus based on wrap-around motion compensation for independently coded sub-pictures.

[0007] Another object of the present disclosure is to provide a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0008] Also, an object of the present disclosure is to provide a recording medium storing a bitstream generated by an image encoding method or apparatus according to the present disclosure.

[0009] Also, an object of the present disclosure is to provide a recording medium storing a bitstream received by an image decoding apparatus according to the present disclosure, decoded, and used for restoring an image.

[0010] The technical problems to be solved in the present disclosure are not limited to the above-described technical problems, and other technical problems not described above will be clearly understood by those having ordinary knowledge in the technical field to which the present disclosure pertains from the following description.

Means for Solving the Problems

[0011] An image decoding method according to an aspect of the present disclosure includes: obtaining, from a bitstream, inter-prediction information and wrap-around information of a current block; and generating a predicted block of the current block based on the inter-prediction information and the wrap-around information, wherein the wrap-around information includes a first flag indicating whether wrap-around motion compensation is available for a current picture including the current block, and the predicted block of the current block is generated by the wrap-around motion compensation based on the first flag having a first value indicating that the wrap-around motion compensation is available, and the wrap-around motion compensation can be performed based on either a boundary of the current sub-picture including the current block or a boundary of a reference picture of the current block based on whether the current sub-picture including the current block is independently coded.

[0012] An image decoding apparatus according to another aspect of the present disclosure includes a memory and at least one processor. The at least one processor acquires inter-prediction information and wrap-around information of a current block from a bitstream, generates a predicted block of the current block based on the inter-prediction information and the wrap-around information. The wrap-around information includes a first flag indicating whether wrap-around motion compensation is available for a current picture including the current block. The predicted block of the current block is generated by the wrap-around motion compensation based on the first flag having a first value indicating that the wrap-around motion compensation is available. The wrap-around motion compensation can be performed based on either the boundary of the current sub-picture including the current block or the boundary of the reference picture of the current block, based on whether the current sub-picture including the current block is independently coded.

[0013] An image encoding method according to another aspect of the present disclosure includes a step of determining whether to apply wrap-around motion compensation to a current block, a step of generating a predicted block of the current block by performing inter-prediction based on the determination, and a step of encoding the inter-prediction information of the current block and the wrap-around information regarding the wrap-around motion compensation. The wrap-around information includes a first flag indicating whether wrap-around motion compensation is available for a current picture including the current block. The wrap-around motion compensation can be performed based on either the boundary of the current sub-picture including the current block or the boundary of the reference picture of the current block, based on whether the current sub-picture including the current block is independently coded.

[0014] A computer-readable recording medium according to another aspect of the present disclosure can store a bitstream generated by the image encoding method or the image encoding apparatus of the present disclosure.

[0015] According to another aspect of the present disclosure, a transmission method can transmit a bitstream generated by the image encoding apparatus or the image encoding method of the present disclosure.

[0016] The features briefly summarized and described above about the present disclosure are merely exemplary aspects of the detailed description of the present disclosure to be described later, and do not limit the scope of the present disclosure.

Effects of the Invention

[0017] According to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0018] Also, according to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus based on wrap-around motion compensation.

[0019] Also, according to the present disclosure, it is possible to provide an image encoding / decoding method and apparatus based on wrap-around motion compensation for independently coded sub-pictures.

[0020] Also, according to the present disclosure, it is possible to provide a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0021] Also, according to the present disclosure, it is possible to provide a recording medium storing a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0022] Also, the present disclosure can provide a recording medium storing a bitstream received by the image decoding apparatus according to the present disclosure, decoded, and used for image restoration.

[0023] The effects obtained in the present disclosure are not limited to the effects described above, and other effects not described above will be clearly understood by those of ordinary skill in the technical field to which the present disclosure pertains from the following description.

Brief Description of the Drawings

[0024]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14a

Figure 14b

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Mode for Carrying Out the Invention

[0025] Hereinafter, with reference to the accompanying drawings, embodiments of the present disclosure will be described in detail so that those having ordinary knowledge in the technical field to which the present disclosure pertains can easily implement them. However, the present disclosure can be realized in various different forms and is not limited to the embodiments described herein.

[0026] When it is determined that a specific description of a known configuration or function may obscure the gist of the present disclosure in describing an embodiment of the present disclosure, the detailed description thereof will be omitted. In the drawings, parts not related to the description of the present disclosure are omitted, and the same reference numerals are given to the same parts.

[0027] In the present disclosure, when a component is "connected", "coupled", or "joined" to another component, this can include not only a direct connection relationship but also an indirect connection relationship in which another component exists between them. Also, when a component "includes" or "has" another component, this means that, unless otherwise stated to the contrary, it does not exclude other components but can further include other components.

[0028] In the present disclosure, terms such as "first", "second", etc. are used only for the purpose of distinguishing one component from another and do not limit the order or importance, etc. between components unless otherwise specifically mentioned. Therefore, within the scope of the present disclosure, the first component of one embodiment may be referred to as the second component in another embodiment, and similarly, the second component of one embodiment may be referred to as the first component in another embodiment.

[0029] In the present disclosure, components that are distinguished from each other are for clearly explaining their respective features and do not necessarily mean that the components are separated. That is, a plurality of components may be integrated and configured as one hardware or software unit, or one component may be distributed and configured as a plurality of hardware or software units. Therefore, even without separate mention, such integrated or distributed embodiments are also included within the scope of the present disclosure.

[0030] In the present disclosure, the components described in various embodiments do not necessarily mean essential components, and some may be optional components. Therefore, embodiments constituted by a subset of the components described in one embodiment are also included within the scope of the present disclosure. Also, embodiments that further include other components in addition to the components described in various embodiments are included within the scope of the present disclosure.

[0031] The present disclosure relates to image encoding and decoding, and the terms used in the present disclosure can have the ordinary meanings in the technical field to which the present disclosure belongs unless newly defined in the present disclosure.

[0032] In the present disclosure, "picture" generally means a unit indicating any one image in a specific time period, and a slice / tile is an encoding unit constituting a part of a picture, and one picture can be composed of one or more slices / tiles. Also, a slice / tile can include one or more CTUs (coding tree units).

[0033] In the present disclosure, "pixel" or "pel" can mean the smallest unit constituting one picture (or image). Also, the term "sample" can be used as a term corresponding to a pixel. A sample can generally indicate a pixel or a pixel value, and can also indicate only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component.

[0034] In the present disclosure, "unit" can indicate a basic unit of image processing. A unit can include at least one of a specific region of a picture and information related to the region. A unit can be used interchangeably with terms such as "sample array", "block", or "area" as the case may be. In general, an M×N block can include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.

[0035] In the present disclosure, "current block" can mean any one of "current coding block", "current coding unit", "block to be encoded", "block to be decoded", or "block to be processed". When prediction is performed, "current block" can mean "current prediction block" or "block to be predicted". When transformation (inverse transformation) / quantization (inverse quantization) is performed, "current block" can mean "current transformation block" or "block to be transformed". When filtering is performed, "current block" can mean "block to be filtered".

[0036] Also, in the present disclosure, "current block" can mean a block that includes all luma component blocks and chroma component blocks or "luma block of the current block", unless explicitly stated as a chroma block. The luma component block of the current block can be expressed explicitly including an explicit description of the luma component block, such as "luma block" or "current luma block". Also, the chroma component block of the current block can be expressed explicitly including an explicit description of the chroma component block, such as "chroma block" or "current chroma block".

[0037] In the present disclosure, " / " and "," can be interpreted as "and / or". For example, "A / B" and "A, B" can be interpreted as "A and / or B". Also, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C".

[0038] In the present disclosure, "or" can be interpreted as "and / or". For example, "A or B" can mean 1) only "A", 2) only "B", or 3) "A and B". Alternatively, in the present disclosure, "or" can mean "additionally or alternatively".

[0039] Overview of the Video Coding System

[0040] FIG. 1 is a diagram schematically showing a video coding system to which an embodiment according to the present disclosure can be applied.

[0041] A video coding system according to an embodiment can include an encoding device 10 and a decoding device 20. The encoding device 10 can transmit encoded video and / or image information or data to the decoding device 20 in a file or streaming format via a digital storage medium or a network.

[0042] The encoding device 10 according to an embodiment can include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. The decoding device 20 according to an embodiment can include a reception unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 can be referred to as a video / image encoding unit, and the decoding unit 22 can be referred to as a video / image decoding unit. The transmission unit 13 can be included in the encoding unit 12. The reception unit 21 can be included in the decoding unit 22. The rendering unit 23 can also include a display unit, and the display unit can be configured as a separate device or an external component.

[0043] The video source generation unit 11 can obtain a video / image through a process such as capture, synthesis, or generation of the video / image. The video source generation unit 11 can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / images, and the like. The video / image generation device can include, for example, a computer, a tablet, a smartphone, and the like, and can generate a video / image (electronically). For example, a virtual video / image can be generated via a computer or the like, and in this case, the video / image capture process can be replaced by a process in which related data is generated.

[0044] The symbolization unit 12 can encode the input video / image. For compression and encoding efficiency, the symbolization unit 12 can perform a series of procedures such as prediction, transformation, quantization, etc. The symbolization unit 12 can output the encoded data (encoded video / image information) in the form of a bitstream.

[0045] The transmission unit 13 can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit 21 of the decoding device 20 via a digital storage medium or network in a file or streaming format. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray (registered trademark), HDD, SSD, etc. The transmission unit 13 can include elements for generating a media file via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit 21 can extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit 22.

[0046] The decoding unit 22 can perform a series of procedures such as inverse quantization, inverse transformation, prediction, etc., corresponding to the operations of the symbolization unit 12 to decode the video / image.

[0047] The rendering unit 23 can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0048] Overview of the Image Encoding Device

[0049] FIG. 2 is a diagram schematically showing an image encoding apparatus to which an embodiment according to the present disclosure can be applied.

[0050] As shown in FIG. 2, the image encoding apparatus 100 can include an image dividing unit 110, a subtraction unit 115, a conversion unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse conversion unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190. The inter prediction unit 180 and the intra prediction unit 185 can be collectively referred to as a "prediction unit". The conversion unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse conversion unit 150 can be included in a residual processing unit. The residual processing unit can further include the subtraction unit 115.

[0051] All or at least a part of the plurality of components constituting the image encoding apparatus 100 can be realized by one hardware component (for example, an encoder or a processor) according to an embodiment. Further, the memory 170 can include a DPB (decoded picture buffer) and can be realized by a digital storage medium.

[0052] The image segmentation unit 110 can divide an input image (or picture, frame) input to the image encoding apparatus 100 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). The coding unit can be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) according to a QT / BT / TT (Quad-tree / Binary-tree / Ternary-tree) structure. For example, one coding unit can be divided into a plurality of coding units at a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the division of the coding unit, the quadtree structure can be applied first, and the binary tree structure and / or the ternary tree structure can be applied later. Based on the final coding unit that cannot be divided any further, the coding procedure according to the present disclosure can be performed. The largest coding unit can be used as the final coding unit, and the coding units at a lower depth obtained by dividing the largest coding unit can also be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and / or restoration, which will be described later. As another example, the processing unit of the coding procedure can be a prediction unit (PU: Prediction Unit) or a transformation unit (TU: Transform Unit). The prediction unit and the transformation unit can be divided or partitioned from the final coding unit, respectively. The prediction unit can be a unit of sample prediction, and the transformation unit can be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from the transformation coefficients.

[0053] The prediction unit (inter prediction unit 180 or intra prediction unit 185) can perform a prediction on a processing target block (current block) and generate a predicted block including prediction samples for the current block. The prediction unit can determine whether intra prediction is applied in units of the current block or CU, or whether inter prediction is applied. The prediction unit can generate various information related to the prediction of the current block and transmit it to the entropy encoding unit 190. The information related to the prediction can be encoded by the entropy encoding unit 190 and output in the form of a bit stream.

[0054] The intra prediction unit 185 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located in the neighborhood of the current block or at a distance according to the intra prediction mode and / or intra prediction technique. The intra prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the Planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes according to the degree of detail of the prediction direction. However, this is only an example, and more or fewer directional prediction modes can be used based on the setting. The intra prediction unit 185 can also determine the prediction mode to be applied to the current block using the prediction mode applied to the neighboring blocks.

[0055] The inter prediction unit 180 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the peripheral blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different from each other. The temporal neighboring block can be called by names such as a collocated reference block, a collocated CU (colCU), etc. The reference picture including the temporal neighboring block can be called a collocated picture (colPic). For example, the inter prediction unit 180 can construct a motion information candidate list based on the peripheral blocks, and generate information indicating which candidate is used to derive the motion vector and / or the reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the inter prediction unit 180 can use the motion information of the peripheral blocks as the motion information of the current block. In the case of the skip mode, unlike the merge mode, the residual signal cannot be transmitted.In the case of the motion information prediction (MVP) mode, the motion vectors of neighboring blocks are used as motion vector predictors, and the motion vector difference and the indicator for the motion vector predictor are encoded to signal the motion vector of the current block. The motion vector difference can mean the difference between the motion vector of the current block and the motion vector predictor.

[0056] The prediction unit can generate a prediction signal based on various prediction methods and / or prediction techniques described later. For example, the prediction unit can apply intra prediction or inter prediction for predicting the current block, and can also apply intra prediction and inter prediction simultaneously. The prediction method of applying intra prediction and inter prediction simultaneously for predicting the current block can be called CIIP (combined inter and intra prediction). In addition, the prediction unit can also perform intra block copy (IBC) for predicting the current block. Intra block copy can be used for content image / video coding such as games, for example, like SCC (screen content coding). IBC is a method of predicting the current block using a restored reference block in the current picture at a position separated from the current block by a predetermined distance. When IBC is applied, the position of the reference block in the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but can be performed in the same way as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in the present disclosure.

[0057] The prediction signal generated by the prediction unit can be used to generate a restored signal or can be used to generate a residual signal. The subtraction unit 115 can subtract the prediction signal (predicted block, predicted sample array) output from the prediction unit from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array). The generated residual signal can be transmitted to the conversion unit 120.

[0058] The conversion unit 120 can apply a conversion technique to the residual signal to generate conversion coefficients (transform coefficients). For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the conversion obtained from this graph when the relationship information between pixels is represented by a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. The conversion process can also be applied to a pixel block having the same size of a square, or can be applied to a non-square, variable-size block.

[0059] The quantization unit 130 can quantize the transform coefficients and transmit them to the entropy encoding unit 190. The entropy encoding unit 190 can encode the quantized signal (information regarding the quantized transform coefficients) and output it in the form of a bitstream. The information regarding the quantized transform coefficients can be called residual information. The quantization unit 130 can reorder the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.

[0060] The entropy encoding unit 190 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding). The entropy encoding unit 190 can also encode, together or separately, information necessary for video / image restoration (for example, values of syntax elements) in addition to the quantized transform coefficients. The encoded information (for example, encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information can further include information regarding various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information can further include general constraint information. The signaling information, the transmitted information, and / or the syntax elements mentioned in the present disclosure can be encoded through the above-described encoding procedure and included in the bitstream.

[0061] The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) for transmitting and / or a storage unit (not shown) for storing the signal output from the entropy encoding unit 190 can be provided as internal / external elements of the image encoding apparatus 100, or the transmission unit can also be provided as a component of the entropy encoding unit 190.

[0062] The quantized transform coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 140 and the inverse transformation unit 150, a residual signal (residual block or residual sample) can be restored.

[0063] The addition unit 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual for the block to be processed, such as when the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 155 can be called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed in the current picture and can also be used for inter prediction of the next picture after passing through filtering as described later.

[0064] The filtering unit 160 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 160 can apply various filtering methods to the restored picture to generate a modified restored picture, and can save the modified restored picture in the memory 170, specifically in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit 160 can generate various information related to filtering as described later in the description of each filtering method and transmit it to the entropy encoding unit 190. The information related to filtering can be encoded by the entropy encoding unit 190 and output in the form of a bitstream.

[0065] The modified restored picture transmitted to the memory 170 can be used as a reference picture in the inter prediction unit 180. When inter prediction is applied through this, the image encoding apparatus 100 can avoid prediction mismatches between the image encoding apparatus 100 and the image decoding apparatus, and can also improve the encoding efficiency.

[0066] The DPB in the memory 170 can save the modified restored picture for use as a reference picture in the inter prediction unit 180. The memory 170 can save the motion information of the block where the motion information in the current picture was derived (or encoded) and / or the motion information of the block in the already restored picture. The saved motion information can be transmitted to the inter prediction unit 180 for utilization as the motion information of the spatial neighboring blocks or the motion information of the temporal neighboring blocks. The memory 170 can save the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 185.

[0067] Overview of the Image Decoding Device

[0068] Figure 3 is a diagram schematically showing an image decoding apparatus to which an embodiment according to the present disclosure can be applied.

[0069] As shown in FIG. 3, the image decoding apparatus 200 can be configured to include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an addition unit 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 can be collectively referred to as a "prediction unit". The inverse quantization unit 220 and the inverse transform unit 230 can be included in a residual processing unit.

[0070] All or at least a part of the plurality of components constituting the image decoding apparatus 200 can be realized by one hardware component (for example, a decoder or a processor) according to an embodiment. Further, the memory 170 can include a DPB and can be realized by a digital storage medium.

[0071] The image decoding apparatus 200 that has received a bitstream including video / image information can execute a process corresponding to the process performed by the image encoding apparatus 100 in FIG. 2 to restore an image. For example, the image decoding apparatus 200 can perform decoding using the processing unit applied in the image encoding apparatus. Therefore, the decoding processing unit can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. Then, the restored image signal decoded and output via the image decoding apparatus 200 can be played back via a playback apparatus (not shown).

[0072] The image decoding device 200 can receive the signal output from the image encoding device in FIG. 2 in bitstream format. The received signal can be decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information can further include information regarding various parameter sets such as an Adaptive Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Also, the video / image information can further include general constraint information. The image decoding device can further use the information regarding the parameter set and / or the general constraint information for decoding the image. The signaling information, the received information, and / or the syntax elements referred to in the present disclosure can be obtained from the bitstream by being decoded through the decoding procedure. For example, the entropy decoding unit 210 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for image restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element from the bitstream, determines a context model using the syntax element information to be decoded, the decoding information of the surrounding blocks and the block to be decoded, or the information of the symbol / bin decoded in the previous step, predicts the occurrence probability of the bin based on the determined context model, and performs arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model.Of the information decoded by the entropy decoding unit 210, the information related to prediction is provided to the prediction units (inter prediction unit 260 and intra prediction unit 265), and the residual values that have undergone entropy decoding in the entropy decoding unit 210, that is, the quantized transform coefficients and related parameter information, can be input to the inverse quantization unit 220. Also, of the information decoded by the entropy decoding unit 210, the information related to filtering can be provided to the filtering unit 240. On the other hand, a receiving unit (not shown) that receives a signal output from the image encoding device can be further provided as an internal / external element of the image decoding device 200, or the receiving unit can also be provided as a component of the entropy decoding unit 210.

[0073] On the other hand, the image decoding device according to the present disclosure can be called a video / image / picture decoding device. The image decoding device can also include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoding unit 210, and the sample decoder can include at least one of the inverse quantization unit 220, the inverse transform unit 230, the addition unit 235, the filtering unit 240, the memory 250, the inter prediction unit 260, and the intra prediction unit 265.

[0074] In the inverse quantization unit 220, the quantized transform coefficients can be inverse quantized to output transform coefficients. The inverse quantization unit 220 can reorder the quantized transform coefficients in a two-dimensional block format. In this case, the reordering can be performed based on the coefficient scan order performed in the image encoding device. The inverse quantization unit 220 can perform inverse quantization on the quantized transform coefficients using a quantization parameter (for example, quantization step size information) to obtain transform coefficients.

[0075] In the inverse conversion unit 230, the conversion coefficients can be inversely converted to obtain a residual signal (residual block, residual sample array).

[0076] The prediction unit can perform prediction on the current block and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 210, and can determine a specific intra / inter prediction mode (prediction technique).

[0077] The prediction unit can generate a prediction signal based on various prediction methods (techniques) described later, which is the same as described in the explanation of the prediction unit of the image encoding device 100.

[0078] The intra prediction unit 265 can predict the current block by referring to samples within the current picture. The explanation of the intra prediction unit 185 can be similarly applied to the intra prediction unit 265.

[0079] The inter prediction unit 260 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, neighboring blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 260 can construct a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes (techniques), and the information regarding the prediction can include information indicating the mode (technique) of inter prediction for the current block.

[0080] The addition unit 235 can generate a restored signal (restored picture, restored block, restored sample array) by adding the obtained residual signal to a predicted signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 260 and / or the intra prediction unit 265). When there is no residual for the processing target block as in the case where the skip mode is applied, the predicted block can be used as the restored block. The description of the addition unit 155 can be similarly applied to the addition unit 235. The addition unit 235 may also be referred to as a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next processing target block within the current picture and can also be used for inter prediction of the next picture through filtering as described later.

[0081] The filtering unit 240 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 240 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 250, specifically, in the DPB of the memory 250. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0082] The (modified) restored picture stored in the DPB of the memory 250 can be used as a reference picture by the inter prediction unit 260. The memory 250 can store the motion information of the block where the motion information in the current picture has been derived (or decoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 260 for utilization as the motion information of the spatial neighboring blocks or the motion information of the temporal neighboring blocks. The memory 250 can store the restored samples of the restored blocks in the current picture and transmit them to the intra prediction unit 265.

[0083] In this specification, the embodiments described in the filtering unit 160, the inter prediction unit 180, and the intra prediction unit 185 of the image encoding apparatus 100 can be similarly or correspondingly applied to the filtering unit 240, the inter prediction unit 260, and the intra prediction unit 265 of the image decoding apparatus 200.

[0084] Overview of Inter Prediction

[0085] Hereinafter, the inter prediction according to the present disclosure will be described.

[0086] The prediction unit of the image encoding device / image decoding device according to the present disclosure can derive a prediction sample by performing inter prediction in units of blocks. The inter prediction can indicate a prediction derived in a method that depends on data elements (e.g., sample values, motion information, etc.) of pictures other than the current picture. When inter prediction is applied to the current block, a predicted block (prediction block or prediction sample array) for the current block can be induced based on a reference block (reference sample array) specified by a motion vector on a reference picture indicated by a reference picture index. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information of the current block can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the peripheral block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. When inter prediction is applied, the peripheral block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block can be called by names such as a collocated reference block, a collocated CU (colCU), a col block (colBlock), etc., and the reference picture including the temporal neighboring block can be called by names such as a collocated picture (colPic), a col picture (col Picture), etc. For example, a motion information candidate list can be configured based on the peripheral blocks of the current block, and flag or index information indicating which candidate is selected (used) can be signaled to derive the motion vector and / or reference picture index of the current block.

[0087] Inter prediction can be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the motion information of the current block may be the same as the motion information of the selected neighboring block. In the case of the skip mode, different from the merge mode, the residual signal cannot be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of the selected neighboring block is used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived by using the sum of the motion vector predictor and the motion vector difference. In the present disclosure, the MVP mode can be used in the same meaning as AMVP (Advanced Motion Vector Prediction).

[0088] The motion information can include L0 motion information and / or L1 motion information based on an inter prediction type (such as L0 prediction, L1 prediction, Bi prediction, etc.). The motion vector in the L0 direction can be called the L0 motion vector or MVL0, and the motion vector in the L1 direction can be called the L1 motion vector or MVL1. The prediction based on the L0 motion vector can be called L0 prediction, the prediction based on the L1 motion vector can be called L1 prediction, and the prediction based on both the L0 motion vector and the L1 motion vector can be called bi (Bi) prediction. Here, the L0 motion vector can indicate a motion vector related to the reference picture list L0 (L0), and the L1 motion vector can indicate a motion vector related to the reference picture list L1 (L1). The reference picture list L0 can include pictures previous to the current picture in output order as reference pictures, and the reference picture list L1 can include pictures subsequent to the current picture in output order. The previous picture can be called a forward (reference) picture, and the subsequent picture can be called a backward (reference picture). The reference picture list L0 can further include pictures subsequent to the current picture in output order as reference pictures. In this case, the previous picture can be indexed first within the reference picture list L0, and the subsequent pictures can be indexed next. The reference picture list L1 can further include pictures previous to the current picture in output order as reference pictures. In this case, the subsequent pictures can be indexed first within the reference picture list L1, and the previous pictures can be indexed next. Here, the output order can correspond to the POC (picture order count) order (order).

[0089] FIG. 4 is a flowchart showing an inter prediction-based video / image encoding method.

[0090] FIG. 5 is a diagram exemplarily showing the configuration of the inter prediction unit 180 according to the present disclosure.

[0091] The encoding method of FIG. 4 can be performed by the image encoding apparatus of FIG. 2. Specifically, step S410 can be performed by the inter prediction unit 180, and step S420 can be performed by the residual processing unit. Specifically, step S420 can be performed by the subtraction unit 115. Step S430 can be performed by the entropy encoding unit 190. The prediction information of step S430 is derived by the inter prediction unit 180, and the residual information of step S430 can be derived by the residual processing unit. The residual information is information regarding the residual sample. The residual information can include information regarding the quantized transform coefficients for the residual sample. As described above, the residual sample is derived as a transform coefficient via the transform unit 120 of the image encoding apparatus, and the transform coefficient can be derived as a quantized transform coefficient via the quantization unit 130. The information regarding the quantized transform coefficient can be encoded by the entropy encoding unit 190 via the residual coding procedure.

[0092] Referring to FIGS. 4 and 5 together, the image encoding device can perform inter prediction on the current block (S410). The image encoding device can derive an inter prediction mode and motion information for the current block and generate a prediction sample for the current block. Here, the inter prediction mode determination, motion information derivation, and prediction sample generation procedures may be performed simultaneously, or any one of the procedures may be performed prior to the other procedures. For example, as shown in FIG. 5, the inter prediction unit 180 of the image encoding device may include a prediction mode determination unit 181, a motion information derivation unit 182, and a prediction sample derivation unit 183. The prediction mode determination unit 181 can determine a prediction mode for the current block, the motion information derivation unit 182 can derive the motion information of the current block, and the prediction sample derivation unit 183 can derive the prediction sample of the current block. For example, the inter prediction unit 180 of the image encoding device can search for a block similar to the current block within a certain region (search region) of the reference picture through motion estimation, and derive a reference block whose difference from the current block is the smallest or below a certain criterion. Based on this, a reference picture index indicating the reference picture where the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The image encoding device can determine a mode to be applied to the current block among various prediction modes. The image encoding device can compare the rate-distortion (RD) costs for the various prediction modes and determine an optimal prediction mode for the current block. However, the method by which the image encoding device determines a prediction mode for the current block is not limited to the above example, and various methods can be used.

[0093] For example, when the skip mode or the merge mode is applied to the current block, the image encoding device can derive merge candidates from the surrounding blocks of the current block and configure a merge candidate list using the derived merge candidates. Further, the image encoding device can derive a reference block among the reference blocks pointed to by the merge candidates included in the merge candidate list, where the difference from the current block is the smallest or below a certain criterion. In this case, the merge candidate related to the derived reference block can be selected, and merge index information indicating the selected merge candidate can be generated and signaled to the image decoding device. The motion information of the current block can be derived using the motion information of the selected merge candidate.

[0094] As another example, when the MVP mode is applied to the current block, the image encoding device can derive mvp (motion vector predictor) candidates from the surrounding blocks of the current block and configure an mvp candidate list using the derived mvp candidates. Further, among the mvp candidates included in the mvp candidate list, the image encoding device can use the motion vector of the selected mvp candidate as the mvp of the current block. In this case, for example, the motion vector pointing to the reference block derived by the above-described motion estimation can be used as the motion vector of the current block, and among the mvp candidates, the mvp candidate having the motion vector with the smallest difference from the motion vector of the current block can be the selected mvp candidate. An MVD (motion vector difference), which is the difference obtained by subtracting the mvp from the motion vector of the current block, can be derived. In this case, the index information indicating the selected mvp candidate and the information regarding the MVD can be signaled to the image decoding device. Further, when the MVP mode is applied, the value of the reference picture index can be configured with reference picture index information and signaled separately to the image decoding device.

[0095] The image encoding device can derive a residual sample based on the prediction sample (S420). The image encoding device can derive the residual sample by comparing the original sample of the current block with the prediction sample. For example, the residual sample can be derived by subtracting the corresponding prediction sample from the original sample.

[0096] The image encoding device can encode image information including prediction information and residual information (S430). The image encoding device can output the encoded image information in bitstream format. The prediction information is information related to the prediction procedure and can include prediction mode information (e.g., skip flag, merge flag, or mode index, etc.) and information related to motion information. Among the prediction mode information, the skip flag is information indicating whether the skip mode is applied to the current block, and the merge flag is information indicating whether the merge mode is applied to the current block. Alternatively, the prediction mode information may be information indicating any one of a plurality of prediction modes, such as a mode index. When the skip flag and the merge flag are both 0, it can be determined that the MVP mode is applied to the current block. The information related to the motion information can include candidate selection information (e.g., merge index, mvp flag, or mvp index) which is information for deriving a motion vector. Among the candidate selection information, the merge index can be signaled when the merge mode is applied to the current block and can be information for selecting any one of the merge candidates included in the merge candidate list. Among the candidate selection information, the mvp flag or the mvp index can be signaled when the MVP mode is applied to the current block and can be information for selecting any one of the mvp candidates included in the mvp candidate list. Also, the information related to the motion information can include information related to the above-mentioned MVD and / or reference picture index information. Also, the information related to the motion information can include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information related to the residual samples. The residual information can include information related to the quantized transform coefficients for the residual samples.

[0097] The output bitstream can be stored in a (digital) storage medium and transmitted to the image decoding device, or can also be transmitted to the image decoding device via a network.

[0098] On the other hand, as described above, the image encoding device can generate a reconstructed picture (a picture including reconstructed samples and reconstructed blocks) based on the reference samples and the residual samples. This is because the same prediction result as that performed by the image decoding device is derived by the image encoding device, and thus the coding efficiency can be improved. Therefore, the image encoding device can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in the memory and utilize it as a picture for inter prediction. As described above, in-loop filtering procedures and the like can be further applied to the reconstructed picture.

[0099] FIG. 6 is a flowchart showing an inter prediction-based video / image decoding method, and FIG. 7 is a diagram exemplarily showing the configuration of the inter prediction unit 260 according to the present disclosure.

[0100] The image decoding device can perform operations corresponding to the operations performed by the image encoding device. The image decoding device can perform prediction on the current block based on the received prediction information and derive prediction samples.

[0101] The decoding method of FIG. 6 can be performed by the image decoding apparatus of FIG. 3. Steps S610 to S630 can be performed by the inter prediction unit 260, and the prediction information of step S610 and the residual information of step S640 can be obtained from the bitstream by the entropy decoding unit 210. The residual processing unit of the image decoding apparatus can derive a residual sample for the current block based on the residual information (S640). Specifically, the inverse quantization unit 220 of the residual processing unit performs inverse quantization based on the quantized transform coefficients derived based on the residual information to derive transform coefficients, and the inverse transform unit 230 of the residual processing unit can perform an inverse transform on the transform coefficients to derive a residual sample for the current block. Step S650 can be performed by the addition unit 235 or the restoration unit.

[0102] Referring to FIGS. 6 and 7 together, the image decoding apparatus can determine a prediction mode for the current block based on the received prediction information (S610). The image decoding apparatus can determine which inter prediction mode is applied to the current block based on the prediction mode information in the prediction information.

[0103] For example, based on the skip flag, it can be determined whether the skip mode is applied to the current block. Also, based on the merge flag, it can be determined whether the merge mode is applied to the current block or the MVP mode is determined. Alternatively, based on the mode index, any one of various inter prediction mode candidates can be selected. The inter prediction mode candidates can include the skip mode, the merge mode, and / or the MVP mode, or can include various inter prediction modes described later.

[0104] The image decoding device can derive the motion information of the current block based on the determined inter prediction mode (S620). For example, when the skip mode or the merge mode is applied to the current block, the image decoding device can construct a merge candidate list described later and select any one of the merge candidates included in the merge candidate list. The selection can be performed based on the candidate selection information (merge index) described above. The motion information of the current block can be derived using the motion information of the selected merge candidate. For example, the motion information of the selected merge candidate can be used as the motion information of the current block.

[0105] As another example, when the MVP mode is applied to the current block, the image decoding device can construct an mvp candidate list and use the motion vector of the mvp candidate selected from among the mvp candidates included in the mvp candidate list as the mvp of the current block. The selection can be performed based on the candidate selection information (mvp flag or mvp index) described above. In this case, based on the information regarding the MVD, the MVD of the current block can be derived, and based on the mvp of the current block and the MVD, the motion vector of the current block can be derived. Also, based on the reference picture index information, the reference picture index of the current block can be derived. The picture pointed to by the reference picture index within the related reference picture list regarding the current block can be derived as the reference picture to be referred to for the inter prediction of the current block.

[0106] The image decoding apparatus can generate a prediction sample for the current block based on the motion information of the current block (S630). In this case, the reference picture can be derived based on the reference picture index of the current block, and the prediction sample of the current block can be derived using the samples of the reference block pointed to by the motion vector of the current block on the reference picture. Optionally, a prediction sample filtering procedure can be further performed on all or part of the prediction samples of the current block.

[0107] For example, as shown in FIG. 7, the inter prediction unit 260 of the image decoding apparatus can include a prediction mode determination unit 261, a motion information derivation unit 262, and a prediction sample derivation unit 263. The inter prediction unit 260 of the image decoding apparatus determines a prediction mode for the current block based on the prediction mode information received from the prediction mode determination unit 261, and derives the motion information (such as a motion vector and / or a reference picture index) of the current block based on the information related to the motion information received from the motion information derivation unit 262, and the prediction sample derivation unit 263 can derive the prediction sample of the current block.

[0108] The image decoding apparatus can generate a residual sample for the current block based on the received residual information (S640). The image decoding apparatus can generate a restored sample for the current block based on the prediction sample and the residual sample, and generate a restored picture based on this (S650). Thereafter, as described above, an in-loop filter procedure and the like can be further applied to the restored picture.

[0109] As described above, the inter prediction procedure can include an inter prediction mode determination step, a motion information derivation step according to the determined prediction mode, and a prediction execution (generation of a prediction sample) step based on the derived motion information. The inter prediction procedure can be performed by the image encoding apparatus and the image decoding apparatus as described above.

[0110] Overview of Sub - Picture

[0111] Hereinafter, the sub-picture according to the present disclosure will be described.

[0112] One picture can be divided into tile units, and each tile can be further divided into sub-picture units. Each sub-picture can include one or more slices and can form a rectangular area within the picture.

[0113] FIG. 8 is a diagram showing an example of a sub-picture.

[0114] Referring to FIG. 8, one picture can be divided into 18 tiles. Twelve tiles can be arranged on the left side of the picture, and each of the tiles can include one sub-picture / slice composed of 16 CTUs. Also, six tiles can be arranged on the right side of the picture, and each of the tiles can include two sub-pictures / slices composed of 4 CTUs. As a result, the picture is divided into 24 sub-pictures, and each of the sub-pictures can include one slice.

[0115] Information about the sub-picture (e.g., the number and size of the sub-pictures, etc.) can be encoded / signaled via a higher-level syntax, such as SPS, PPS, and / or a slice header.

[0116] FIG. 9 is a diagram showing an example of SPS including information about the sub-picture.

[0117] Referring to FIG. 9, the SPS can include a syntax element subpic_info_present_flag that indicates the presence or absence of subpicture information for a CLVS (coded layer video sequence). For example, a subpic_info_present_flag having a first value (e.g., 0) can indicate that there is no subpicture information for the CLVS and that only one subpicture exists within each picture of the CLVS. In contrast, a subpic_info_present_flag having a second value (e.g., 1) can indicate that there is subpicture information for the CLVS and that one or more subpictures exist within each picture of the CLVS. In one example, when the picture spatial resolution can be changed within the CLVS that refers to the SPS (e.g., res_change_in_clvs_allowed_flag == 1), the value of the subpic_info_present_flag can be restricted to the first value (e.g., 0). On the other hand, when the bitstream, as a result of a sub-bitstream extraction process, includes only a subset of subpictures of the input bitstream for the sub-bitstream extraction process, the value of the subpic_info_present_flag can be restricted to the second value (e.g., 1).

[0118] In addition, the SPS may include a syntax element sps_num_subpics_minus1 indicating the number of sub-pictures. For example, the value obtained by adding 1 to sps_num_subpics_inus1 can indicate the number of sub-pictures included in each picture within the CLVS. In one example, the value of sps_num_subpics_minus1 can be restricted to have a range of 0 or more and Ceil(pic_width_max_in_luma_samples / CtbSizeY)*Ceil(pic_height_max_in_luma_samples / CtbSizeY) or less. Here, Ceil(x) can be a ceiling function that outputs the smallest integer value that is the same as or greater than x. Also, pic_width_max_in_luma_samples means the maximum width in terms of luma samples of each picture, pic_height_max_in_luma_samples means the maximum height in terms of luma samples of each picture, and CtbSizeY can mean the array size of each luma component CTB (coding tree block) in both width and height. On the other hand, when sps_num_subpics_minus1 does not exist, the value of sps_num_subpics_minus1 can be inferred as a first value (e.g., 0).

[0119] In addition, the SPS may include a syntax element sps_independent_subpics_flag that indicates whether to treat sub-picture boundaries as picture boundaries. For example, sps_independent_subpics_flag having a second value (e.g., 1) can indicate that all sub-picture boundaries within the CLVS are treated as picture boundaries and that loop filtering across the sub-picture boundaries is not performed. In contrast, sps_independent_subpics_flag having a first value (e.g., 0) can indicate that the above-described constraints do not apply. On the other hand, if sps_independent_subpics_flag does not exist, the value of sps_independent_subpics_flag can be inferred as the first value (e.g., 0).

[0120] In addition, the SPS may include syntax elements subpic_ctu_top_left_x[i], subpic_ctu_top_left_y[i], subpic_width_minus1[i], and subpic_height_minus1[i] that indicate the position and size of the sub-picture.

[0121] subpic_ctu_top_left_x[i] can represent the horizontal position of the top left CTU of the i-th sub-picture in units of CtbSizeY. In one example, the length of subpic_ctu_top_left_x[i] can be Ceil(Log2((pic_width_max_in_luma_samples + CtbSizeY - 1) >> CtbLog2SizeY)) bits. On the other hand, if subpic_ctu_top_left_x[i] does not exist, the value of subpic_ctu_top_left_x[i] can be inferred as the first value (e.g., 0).

[0122] subpic_ctu_top_left_y[i] can represent the vertical position of the top - left CTU of the i - th sub - picture in units of CtbSizeY. In one example, the length of subpic_ctu_top_left_y[i] can be Ceil(Log2((pic_height_max_in_luma_samples + CtbSizeY - 1)>>CtbLog2SizeY)) bits. On the other hand, if subpic_ctu_top_left_y[i] does not exist, the value of subpic_ctu_top_left_y[i] can be inferred as the first value (e.g., 0).

[0123] The value obtained by adding 1 to subpic_width_minus1[i] can represent the width of the i - th sub - picture in units of CtbSizeY. In one example, the length of subpic_width_minus1[i] can be Ceil(Log2((pic_width_max_in_luma_samples + CtbSizeY - 1)>>CtbLog2SizeY)) bits. On the other hand, if subpic_width_minus1[i] does not exist, the value of subpic_width_minus1[i] can be inferred as ((pic_width_max_in_luma_samples + CtbSizeY - 1)>>CtbLog2SizeY)-subpic_ctu_top_left_x[i]-1.

[0124] The value obtained by adding 1 to subpic_height_minus1[i] can represent the height of the i-th subpicture in units of CtbSizeY. In one example, the length of subpic_height_minus1[i] can be Ceil(Log2((pic_height_max_in_luma_samples + CtbSizeY - 1) >> CtbLog2SizeY)) bits. On the other hand, if subpic_height_minus1[i] does not exist, the value of subpic_height_minus1[i] can be inferred as ((pic_height_max_in_luma_samples + CtbSizeY - 1) >> CtbLog2SizeY) - subpic_ctu_top_left_y[i] - 1.

[0125] Also, the SPS can include subpic_treated_as_pic_flag[i] indicating whether the subpicture is treated as one picture. For example, subpic_treated_as_pic_flag[i] having the first value (e.g., 0) can indicate that the i-th subpicture in each coded picture within the CLVS is not treated as one picture in the decoding process excluding the in-loop filtering operation. In contrast, subpic_treated_as_pic_flag[i] having the second value (e.g., 1) can indicate that the i-th subpicture in each coded picture within the CLVS is treated as one picture in the decoding process excluding the in-loop filtering operation. If subpic_treated_as_pic_flag[i] does not exist, the value of subpic_treated_as_pic_flag[i] can be inferred to be the same value as the aforementioned sps_independent_subpics_flag. In one example, subpic_treated_as_pic_flag[i] can be coded / signaled only when the aforementioned sps_independent_subpics_flag has the first value (e.g., 0) (i.e., when the subpicture boundary is not treated as a picture boundary).

[0126] On the other hand, when subpic_treated_as_pic_flag[i] has a second value (e.g., 1), for each output layer and its reference layer included in an output layer set (OLS) that includes the layer containing the i-th subpicture as an output layer, it may be a requirement for bitstream compliance that all of the following conditions are true.

[0127] -(Condition 1) All pictures in the output layer and its reference layer must have the same value of pic_width_in_luma_samples and the same value of pic_height_in_luma_samples.

[0128] -(Condition 2) All SPSs referenced by the output layer and its reference layer must have the same value of sps_num_subpics_minus1, and the same values of subpic_ctu_top_left_x[j], subpic_ctu_top_left_y[j], subpic_width_minus1[j], subpic_height_minus1[j], and loop_filter_across_subpic_enabled_flag[j], respectively. Here, j has a range from 0 or more to sps_num_subpics_minus1 or less.

[0129] In addition, the SPS may include a syntax element loop_filter_across_subpic_enabled_flag[i] indicating whether an in-loop filtering operation across sub-picture boundaries can be performed. For example, loop_filter_across_subpic_enabled_flag[i] having a first value (e.g., 0) can indicate that the in-loop filtering operation across the boundaries of the i-th sub-picture in each coded picture within the CLVS is not performed. In contrast, loop_filter_across_subpic_enabled_flag[i] having a second value (e.g., 1) can indicate that the in-loop filtering operation across the boundaries of the i-th sub-picture in each coded picture within the CLVS can be performed. If loop_filter_across_subpic_enabled_flag[i] does not exist, the value of loop_filter_across_subpic_enabled_flag[i] can be inferred to be the same as the value of 1-sps_independent_subpics_flag. In one example, loop_filter_across_subpic_enabled_flag[i] can be coded / signaled only when the aforementioned sps_independent_subpics_flag has a first value (e.g., 0) (i.e., when the sub-picture boundary is not treated as a picture boundary). On the other hand, as a requirement for bitstream compliance, the form of the sub-picture should be such that when each sub-picture is decoded, the entire left boundary and the entire upper boundary of each sub-picture are composed of the boundaries of the picture or the boundaries of the previously decoded sub-picture.

[0130] FIG. 10 is a diagram showing a method for an image encoding apparatus to encode an image using sub-pictures according to an embodiment of the present disclosure.

[0131] The image encoding device can encode the current picture based on the sub-picture structure. Alternatively, the image encoding device can encode at least one sub-picture that constitutes the current picture and output a (sub) bitstream including the (encoded) information regarding the (encoded) at least one sub-picture.

[0132] Referring to FIG. 10, the image encoding device can divide an input picture into a plurality of sub-pictures (S1010). Then, the image encoding device can generate information regarding the sub-pictures (S1020). Here, the information regarding the sub-pictures can include, for example, information regarding the area of the sub-pictures and / or information regarding the grid spacing for use in the sub-pictures. Also, the information regarding the sub-pictures can include information regarding whether each sub-picture can be treated as one picture and / or information (boundary) regarding whether in-loop filtering can be performed across the boundaries of each sub-picture.

[0133] The image encoding device can encode at least one sub-picture based on the information regarding the sub-pictures. For example, each sub-picture can be independently encoded based on the information regarding the sub-pictures. Then, the image encoding device can encode the image information including the information regarding the sub-pictures and output a bitstream (S1030). Here, the bitstream for the sub-pictures may be referred to as a sub-stream or a sub-bitstream.

[0134] FIG. 11 is a diagram showing a method by which an image decoding device according to an embodiment of the present disclosure decodes an image using sub-pictures.

[0135] The image decoding device can decode at least one sub-picture included in the current picture using the (encoded) information regarding the (encoded) at least one sub-picture obtained from the (sub) bitstream.

[0136] Referring to FIG. 11, the image decoding apparatus can obtain information regarding sub-pictures from a bit stream (S1110). Here, the bit stream can include a sub-stream or a sub-bit stream for sub-pictures. The information regarding sub-pictures can be configured in the upper-level syntax of the bit stream. Then, the image decoding apparatus can derive at least one sub-picture based on the information regarding sub-pictures (S1120).

[0137] The image decoding apparatus can decode at least one sub-picture based on the information regarding sub-pictures (S1130). For example, when the current sub-picture including the current block is treated as one picture, the current sub-picture can be decoded independently. Also, when in-loop filtering can be performed across the boundary of the current sub-picture, in-loop filtering (e.g., deblocking filtering) can be performed on the boundary of the current sub-picture and the boundaries of adjacent sub-pictures adjacent to the boundary. Also, when the boundary of the current sub-picture coincides with the picture boundary, in-loop filtering across the boundary of the current sub-picture cannot be performed. The image decoding apparatus can decode sub-pictures based on methods such as CABAC method, prediction method, residual processing method (transformation, quantization), in-loop filtering method, etc. Then, the image decoding apparatus can output at least one decoded sub-picture or output the current picture including at least one sub-picture. The decoded sub-pictures can be output in the form of an OPS (output sub-picture set). For example, in relation to a 360-degree image or an omnidirectional image, when only a part of the current picture is rendered, only some of the sub-pictures among all the sub-pictures in the current picture can be decoded, and all or part of the decoded sub-pictures can be rendered according to the user's viewport.

[0138] Overview of Wrap - Around

[0139] When inter prediction is applied to a current block, a predicted block of the current block can be derived based on a reference block specified by a motion vector of the current block. At this time, when at least one reference sample in the reference block goes outside a boundary of a reference picture, a sample value of the reference sample can be replaced with a sample value of an adjacent sample existing at the boundary or the outermost contour of the reference picture. This is called padding, and the boundary of the reference picture can be extended through the padding.

[0140] On the other hand, when a reference picture is obtained from a 360-degree image, continuity can exist between a left boundary and a right boundary of the reference picture. Accordingly, a sample adjacent to the left boundary (or the right boundary) of the reference picture can have the same / similar sample value and / or motion information as a sample adjacent to the right boundary (or the left boundary) of the picture. Based on such a characteristic, at least one reference sample going outside a boundary of a reference picture within a reference block can be replaced with an adjacent sample within the reference picture corresponding to the reference sample. This is called (horizontal) wrap-around motion compensation, and a motion vector of a current block can be adjusted to point inside the reference picture through the wrap-around motion compensation.

[0141] Wrap-around motion compensation refers to a coding tool designed to improve the visual quality of a restored image / video, such as a 360-degree image / video projected in an ERP format. According to the existing motion compensation process, when the motion vector of the current block points to a sample outside the boundary of the reference picture, the sample value of the sample outside the boundary can be induced by copying the sample value of the adjacent sample closest to the boundary through repetitive padding. However, since a 360-degree image / video is acquired spherically and essentially has no image boundary, reference samples outside the boundary of the reference picture on the projected domain (2D domain) can always be obtained from adjacent samples adjacent to the reference sample on the old domain (3D domain). Therefore, repetitive padding is not suitable for 360-degree images / videos and may induce visual artifacts called seam artifacts in the restored viewport image / video.

[0142] On the other hand, when a general projection method is applied, 2D-to-3D and 3D-to-2D coordinate conversions are performed together with sample interpolation for fractional sample positions, so it can be difficult to obtain adjacent samples for wrap-around motion compensation on the old domain. However, when the ERP projection method is applied, spherical adjacent samples outside the left boundary (or right boundary) of the reference picture can be obtained relatively easily from samples within the right boundary (or left boundary) of the reference picture. Therefore, considering the relative ease of implementation and its wide use of the ERP projection method, wrap-around motion compensation can be more effective for 360-degree images / videos encoded in the ERP format.

[0143] FIG. 12 is a diagram showing an example of a 360-degree image converted into a 2D picture.

[0144] Referring to FIG. 12, the 360-degree image 1210 can be converted into a two-dimensional picture 1230 through a projection process. The two-dimensional picture 1230 can have various projection formats such as an ERP (equi-rectangular projection) format or a PERP (Padded ERP) format according to the projection method applied to the 360-degree image 1210.

[0145] Due to the image characteristics obtained from all directions, the 360-degree image 1210 has no image boundary. However, the two-dimensional picture 1230 obtained from the 360-degree image 1210 has an image boundary due to the projection process. At this time, the left boundary LBd and the right boundary RBd of the two-dimensional picture 1230 may form a single line RL within the 360-degree image 1210 and be adjacent to each other. Therefore, the similarity between samples adjacent to the left boundary LBd and the right boundary RBd within the two-dimensional picture 1230 may be relatively high.

[0146] On the other hand, a predetermined region within the 360-degree image 1210 can correspond to an internal region or an external region of the two-dimensional picture 1230 according to a reference image boundary. For example, when the left boundary LBd of the two-dimensional picture 1230 is used as a reference, the A region within the 360-degree image 1210 can correspond to the A1 region existing outside the two-dimensional picture 1230. Conversely, when the right boundary RBd of the two-dimensional picture 1230 is used as a reference, the A region within the 360-degree image 1210 can correspond to the A2 region existing inside the two-dimensional picture 1230. The A1 region and the A2 region can have the same / similar sample attributes in that they correspond to the same A region based on the 360-degree image 1210.

[0147] Based on such characteristics, an external sample outside the left boundary LBd of the two-dimensional picture 1230 can be replaced by an internal sample of the two-dimensional picture 1230 located at a position separated by a predetermined distance in the first direction DIR1 through wrap-around motion compensation. For example, an external sample of the two-dimensional picture 1230 included in the A1 region can be replaced by an internal sample of the two-dimensional picture 1230 included in the A2 region. Similarly, an external sample outside the right boundary RBd of the two-dimensional picture 1230 can be replaced by an internal sample of the two-dimensional picture 1230 located at a position separated by a predetermined distance in the second direction DIR2 through wrap-around motion compensation.

[0148] FIG. 13 is a diagram showing an example of the wrap-around motion compensation process.

[0149] Referring to FIG. 13, when inter prediction is applied to the current block 1310, the predicted block of the current block 1310 can be induced based on the reference block 1330.

[0150] The reference block 1330 can be specified by the motion vector 1320 of the current block 1310. In one example, the motion vector 1320 can indicate the upper left position of the reference block 1330 based on the upper left position of the same position block 1315 existing at the same position as the current block 1310 within the reference picture.

[0151] As shown in FIG. 13, the reference block 1330 can include a first region 1335 outside the left boundary of the reference picture. Since the first region 1335 cannot be used for the inter prediction of the current block 1310, it can be replaced by a second region 1340 within the reference picture through wrap-around motion compensation. The second region 1340 can correspond to the same region as the first region 1335 on the old domain (3D domain), and the position of the second region 1340 can be specified by adding a wrap-around offset to a predetermined position (e.g., the upper left position) of the first region 1335.

[0152] The wrap - around offset can be set to the ERP width before padding of the current picture. Here, the ERP width can be meant as the width of the original picture in the ERP format (i.e., the ERP picture) obtained from the 360 - degree image. A horizontal padding process can be performed with respect to the left and right boundaries of the ERP picture. Thereby, the width of the current picture (PicWidth) can be determined as the value obtained by combining all of the ERP width, the padding amount for the left boundary of the ERP picture (left padding), and the padding amount for the right boundary of the ERP picture (right padding). On the other hand, the wrap - around offset can be encoded / signaled using a predetermined syntax element (e.g., pps_ref_wraparound_offset) within the upper - level syntax. The syntax element is not affected by the padding amounts for the left and right boundaries of the ERP picture, and as a result, asymmetric padding for the original picture can be supported. That is, the padding amount for the left boundary of the ERP picture (left padding) and the padding amount for the right boundary of the ERP picture (right padding) can be different from each other.

[0153] The information regarding the wrap - around motion compensation described above (e.g., presence or absence of activation, wrap - around offset, etc.) can be encoded / signaled via an upper - level syntax, such as SPS and / or PPS.

[0154] FIG. 14a is a diagram showing an example of an SPS including information regarding wrap - around motion compensation.

[0155] Referring to FIG. 14a, the SPS may include a syntax element sps_ref_wraparound_enabled_flag that indicates whether to apply wrap-around motion compensation at the sequence level. For example, the sps_ref_wraparound_enabled_flag having a first value (e.g., 0) may indicate that wrap-around motion compensation is not applied to the current video sequence including the current block. In contrast, the sps_ref_wraparound_enabled_flag having a second value (e.g., 1) may indicate that wrap-around motion compensation is applied to the current video sequence including the current block. In one example, wrap-around motion compensation for the current video sequence can be applied only when the picture width (e.g., pic_width_in_luma_samples) and the CTB width (CtbSizeY) satisfy the following conditions.

[0156] -(Condition) (CtbSizeY / MinCbSizeY + 1) ≥ (pic_width_in_luma_samples / MinCbSizeY - 1)

[0157] If the above conditions are not satisfied, for example, if the value of (CtbSizeY / MinCbSizeY + 1) is greater than the value of (pic_width_in_luma_samples / MinCbSizeY - 1), the sps_ref_wraparound_enabled_flag can be limited to the first value (e.g., 0). Here, CtbSizeY can mean the width or height of the luma component CTB, and MinCbSizeY can mean the minimum width or minimum height of the luma component CB (coding block). Also, pic_width_max_in_luma_samples can mean the maximum width in luma sample units of each picture.

[0158] FIG. 14b is a diagram showing an example of a PPS including information related to wrap-around motion compensation.

[0159] Referring to FIG. 14b, the PPS can include a syntax element pps_ref_wraparound_enabled_flag that indicates whether to apply wrap-around motion compensation at the picture level.

[0160] Referring to FIG. 14b, the PPS can include a syntax element pps_ref_wraparound_enabled_flag that indicates whether to apply wrap-around motion compensation at the sequence level. For example, a pps_ref_wraparound_enabled_flag having a first value (e.g., 0) can indicate that wrap-around motion compensation is not applied to the current picture including the current block. In contrast, a pps_ref_wraparound_enabled_flag having a second value (e.g., 1) can indicate that wrap-around motion compensation is applied to the current picture including the current block. In one example, wrap-around motion compensation for the current picture can be applied only when the picture width (e.g., pic_width_in_luma_samples) is greater than the CTB width (CtbSizeY). For example, when the value of (CtbSizeY / MinCbSizeY + 1) is greater than the value of (pic_width_in_luma_samples / MinCbSizeY - 1), the pps_ref_wraparound_enabled_flag can be limited to the first value (e.g., 0). In other examples, when the sps_ref_wraparoud_enabled_flag has the first value (e.g., 0), the value of the pps_ref_wraparound_enabled_flag can be limited to the first value (e.g., 0).

[0161] In addition, PPS can include a syntax element pps_ref_wraparound_offset that indicates the offset for wrap-around motion compensation. For example, the value obtained by adding ((CtbSizeY / MinCbSizeY)+2) to pps_ref_wraparound_offset can indicate the wrap-around offset for calculating the wrap-around position in luma sample units. The value of pps_ref_wraparound_offset can be a value greater than or equal to 0 and less than or equal to ((pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)-2). On the other hand, the variable PpsRefWraparoundOffset can be set to the same value as (pps_ref_wraparound_offset+(CtbSizeY / MinCbSizeY)+2). The variable PpsRefWraparoundOffset can be used in the process of clipping reference samples that are outside the boundary of the reference picture.

[0162] On the other hand, when the current picture is divided into a plurality of sub-pictures, the wrap-around motion compensation can be selectively performed based on the attributes of each sub-picture.

[0163] FIG. 15 is a flowchart showing a method in which an image decoding apparatus performs wrap-around motion compensation based on sub-picture attributes.

[0164] Referring to FIG. 15, the image decoding apparatus can determine whether the current sub-picture is independently coded (S1510).

[0165] When the current sub-picture is independently coded (\"YES\" in S1510), the image decoding apparatus can clip the position of the reference sample based on the sub-picture boundary for motion compensation for the current block (S1520). The operation can be performed using any one of a luma sample bilinear interpolation process, a luma sample interpolation filtering process, a luma integer sample fetching process, and a chroma sample interpolation process.

[0166] When the current sub-picture is not independently coded (\"NO\" in S1510), the image encoding apparatus can determine whether wraparound motion compensation is available for the current block (S1530).

[0167] Whether wraparound motion compensation is available for the current block can be determined based on a predetermined variable (e.g., refWraparoundEnabledFlag). For example, when refWraparoundEnabledFlag has a first value (e.g., 0), wraparound motion compensation may not be available for the current block. In contrast, when refWraparoundEnabledFlag has a second value (e.g., 1), wraparound motion compensation may be available for the current block. In one example, the value of refWraparoundEnabledFlag can be derived based on a predetermined flag (e.g., pps_ref_wraparound_enabled_flag) obtained from a higher-level syntax, such as a picture parameter set.

[0168] When wraparound motion compensation is available for the current block (the answer to S1530 is "YES"), the video decoder can correct the position of the reference sample using the wraparound offset (S1540). For example, the video decoder can correct the position of the reference sample by shifting the x coordinate of the reference sample in the positive or negative direction by the wraparound offset (e.g., PpsRefWraparoundOffset*MinCbSizeY). Then, for motion compensation for the current block, the video decoder can clip the position of the corrected reference sample based on the boundary of the reference picture (S1550).

[0169] In contrast, when wraparound motion compensation is not available for the current block (the answer to S1530 is "NO"), the video decoder can clip the position of the reference sample based on the boundary of the reference picture for motion compensation for the current block (S1560).

[0170] On the other hand, the above-described clipping operation can be performed using the luma sample bilinear interpolation procedure. A specific example is as shown in Table 1 below.

[0171]

Table 1

[0172] Referring to Table 1, the luma position (xInti, yInti) of the reference sample in integer sample units can be adjusted within the reference picture boundary or sub-picture boundary using a predetermined clipping function (Clip3, ClipH). Here, Clip3(x, y, z) means a function that outputs x when z is smaller than x, outputs y when z is larger than y, and outputs z in other cases. Also, ClipH(x, y, z) means a function that outputs z + x when z is smaller than 0, outputs z - x when z is larger than y - 1, and outputs z in other cases.

[0173] When the sub - picture is currently independently coded, Process 1 can be performed. Specifically, a clipping operation based on the sub - picture boundary can be performed on the x - coordinate and y - coordinate of the reference sample (A110, A120). In Process 1, SubpicLeftBoundaryPos can indicate the left boundary of the sub - picture, SubpicRightBoundaryPos can indicate the right boundary of the sub - picture, SubpicTopBoundaryPos can indicate the upper boundary of the sub - picture, and SubpicBotBoundaryPos can indicate the lower boundary of the sub - picture.

[0174] When the sub - picture is not currently independently coded, Process 2 can be performed. Specifically, for the x - coordinate of the reference sample, depending on whether wrap - around motion compensation is available for the current block (e.g., refWraparoundEnabledFlag == 1), a wrap - around offset (e.g., PpsRefWraparoundOffset) can be selectively applied, and a clipping operation based on the reference picture boundary can be performed on the selectively applied x - coordinate (A130). In contrast, for the y - coordinate of the reference sample, a clipping operation based on the reference picture boundary can be performed regardless of whether the wrap - around motion compensation is available (A140). In Process 2, picW can indicate the width of the reference picture, and picH can indicate the height of the reference picture.

[0175] Alternatively, the above - mentioned clipping operation can be performed using a luma sample interpolation filtering procedure. A specific example is as shown in Table 2 below.

[0176] [Table 2]

[0177] Referring to Table 2, the luma positions (xInti, yInti) of the reference samples in integer sample units can be adjusted within the reference picture boundary or sub-picture boundary using a predetermined clipping function (Clip3, ClipH). Duplicate explanations with Table 1 are omitted.

[0178] Depending on whether the current sub-picture is independently coded or not, Process 1 or Process 2 can be selectively performed. Also, the wrap-around motion compensation can be performed only on the x coordinate of the reference sample when the current sub-picture is not independently coded (Process 2) (A230).

[0179] Alternatively, the above-described clipping operation can be performed using the luma integer sample fetching procedure. A specific example is as shown in Table 3 below.

[0180]

Table 3

[0181] Referring to Table 3, the luma positions (xInt, yInt) of the reference samples in integer sample units can be adjusted within the reference picture boundary or sub-picture boundary using a predetermined clipping function (Clip3, ClipH). Duplicate explanations with Table 1 are omitted.

[0182] Depending on whether the current sub-picture is independently coded or not, Process 1 or Process 2 can be selectively performed. Also, the wrap-around motion compensation can be performed only on the x coordinate of the reference sample when the current sub-picture is not independently coded (Process 2) (A330).

[0183] Also, the above-described clipping operation can be performed using the chroma sample interpolation procedure. A specific example is as shown in Table 4 below.

[0184]

Table 4

[0185] Referring to Table 4, the luma positions (xInt, yInt) of the reference samples in integer sample units can be adjusted within the reference picture boundary or sub-picture boundary using a predetermined clipping function (Clip3, ClipH). At this time, the reference picture boundary or sub-picture boundary can be determined based on the chroma samples, unlike the method described above with reference to Tables 1 to 3. For example, SubWidthC and SubHeightC can indicate the width ratio and height ratio between the luma samples and the chroma samples, respectively. Then, based on the chroma samples, the sub-picture left boundary can be determined as (SubpicLeftBoundaryPos / SubWidthC), the sub-picture right boundary as (SubpicRightBoundaryPos / SubWidthC), the sub-picture upper boundary as (SubpicTopBoundaryPos / SubWidthC), and the sub-picture lower boundary as (SubpicBotBoundaryPos / SubWidthC).

[0186] Depending on whether the current sub-picture is independently coded or not, Process 1 or Process 2 can be selectively performed. Also, the wrap-around motion compensation can be performed only for the x coordinate of the reference samples when the current sub-picture is not independently coded (Process 2) (A430).

[0187] On the other hand, the method of FIG. 15 can also be performed by an image encoding apparatus, which is obvious to an ordinary technician.

[0188] According to the method of FIG. 15 above, wraparound motion compensation for the current block can only be performed when the current sub-picture is not independently coded. As a result, there is a problem that wraparound-related coding tools cannot be used together with various sub-picture-related coding tools that assume independent coding of sub-pictures. This can act as a factor reducing the encoding / decoding performance for pictures with inter-boundary continuity, such as ERP pictures or PERP pictures.

[0189] To solve such a problem, according to an embodiment of the present disclosure, wraparound motion compensation can be performed according to a predetermined condition even when the current sub-picture is independently coded. Hereinafter, embodiments of the present disclosure will be described in detail.

[0190] According to an embodiment of the present disclosure, when all independently coded sub-pictures in the current video sequence have the same width as the picture width, wraparound motion compensation may be available for all sub-pictures in the current video sequence.

[0191] FIG. 16 is a flowchart showing a method by which an image encoding apparatus according to an embodiment of the present disclosure determines whether to use wraparound motion compensation.

[0192] Referring to FIG. 16, the image encoding apparatus can determine whether there is one or more sub-pictures independently coded in the current video sequence (S1610).

[0193] If, as a result of the determination, there are no one or more sub-pictures independently coded within the current video sequence ("NO" in S1610), the image coding device can determine whether wrap-around motion compensation is available for the current video sequence based on predetermined wrap-around constraint conditions (S1640). In this case, the image coding device can code flag information (e.g., sps_wraparound_enabled_flag) indicating whether wrap-around motion compensation is available for the current video sequence into a first value (e.g., 0) or a second value (e.g., 1) based on the determination.

[0194] As an example of the wrap-around constraint conditions, when wrap-around motion compensation is restricted for one or more output layer sets (OLSs) specified by a video parameter set (VPS), wrap-around motion compensation for the current video sequence can be restricted from being available. As another example of the wrap-around constraint conditions, when all sub-pictures within the current video sequence have discontinuous sub-picture boundaries, wrap-around motion compensation for the current video sequence can be restricted from being available.

[0195] In contrast, if there are one or more sub-pictures independently coded within the current video sequence ("YES" in S1610), the image coding device can determine whether there is a sub-picture having a width different from the picture width among the independently coded sub-pictures (S1620).

[0196] In one embodiment, the picture width can be derived as in Equation 1 based on the maximum width that a picture can have within the current video sequence.

[0197]

Equation

[0198] Here, pic_width_max_in_luma_samples represents the maximum width of the picture in luma samples, CtbSizeY represents the width of the CTB (coding tree block) in the picture in luma samples, and CtbLog2SizeY can represent the logarithmic scale value of CtbSizeY.

[0199] If, as a result of the determination, at least one of the independently coded sub-pictures has a width different from the picture width (\"YES\" in S1620), the image coding device can determine that wraparound motion compensation is not available for the current video sequence (S1630). In this case, the image coding device can code sps_ref_wraparound_enabled_flag to a first value (e.g., 0) based on the determination.

[0200] In contrast, if all of the independently coded sub-pictures have the same width as the picture width (\"NO\" in S1620), the image coding device can determine whether wraparound motion compensation is available for the current video sequence based on the above-described wraparound constraint conditions (S1640). In this case, the image coding device can code sps_ref_wraparound_enabled_flag to a first value (e.g., 0) or a second value (e.g., 1) based on the determination.

[0201] As described above, in FIG. 16, it is illustrated that steps S1610 and S1620 are sequentially performed, but this is merely exemplary and does not limit the embodiments of the present disclosure. For example, step S1620 may be performed simultaneously with step S1610 or may be performed prior to step S1610.

[0202] On one hand, the sps_ref_wraparound_enabled_flag encoded by the image encoding device can be saved in the bitstream and signaled to the image decoding device. In this case, the image decoding device can determine whether wraparound motion compensation is available for the current video sequence based on the sps_ref_wraparound_enabled_flag obtained from the bitstream.

[0203] For example, when the sps_ref_wraparound_enabled_flag has a first value (e.g., 0), the image decoding device determines that wraparound motion compensation is not available for the current video sequence and cannot perform wraparound motion compensation on the current block. In this case, the reference sample position of the current block can be clipped based on the reference picture boundary or sub-picture boundary, and motion compensation can be performed using the reference samples at the clipped positions.

[0204] That is, the image decoding device can perform correct motion compensation according to the present disclosure without separately determining whether there is one or more sub-pictures independently coded within the current video sequence and having a width different from the picture width. However, the operation of the image decoding device is not limited thereto. For example, the image decoding device can also perform motion compensation based on the result of the determination after determining whether there is one or more sub-pictures independently coded within the current video sequence and having a width different from the picture width. More specifically, the image decoding device determines whether there is one or more sub-pictures independently decoded within the current video sequence and having a width different from the picture width. If such sub-pictures exist, the sps_ref_wraparound_enabled_flag can be regarded as the first value (e.g., 0), and wraparound motion compensation can be not performed.

[0205] In contrast, when the sps_ref_wraparound_enabled_flag has a second value (e.g., 1), the image decoding device can determine that wraparound motion compensation is available for the current video sequence. In this case, the image decoding device additionally obtains from the bitstream a wraparound flag (e.g., pps_ref_wraparound_enabled_flag) indicating whether wraparound motion compensation is available for the current picture, and can determine whether to perform wraparound motion compensation on the current block based on the obtained wraparound flag information.

[0206] For example, when the pps_ref_wraparound_enabled_flag has a first value (e.g., 0), the image decoding device cannot perform wraparound motion compensation on the current block. In this case, the reference sample position of the current block can be clipped based on the reference picture boundary or the sub-picture boundary, and motion compensation can be performed using the reference samples at the clipped positions. In contrast, when the pps_ref_wraparound_enabled_flag has a second value (e.g., 1), the image decoding device can perform wraparound motion compensation on the current block.

[0207] As described above, when all independently coded sub-pictures in the current video sequence have the same width as the picture width, wraparound motion compensation can be applied to all sub-pictures in the current video sequence. Thereby, since sub-picture related coding tools and wraparound motion compensation related coding tools can be used together, the coding / decoding efficiency can be further improved.

[0208] FIG. 17 is a flowchart showing a method by which an image decoding device according to an embodiment of the present disclosure performs wraparound motion compensation based on sub-picture attributes.

[0209] Referring to FIG. 17, the image decoding apparatus can determine whether the current sub-picture is independently coded (S1710). In one example, whether the current sub-picture is independently coded can be determined based on a higher-level syntax, for example, a predetermined flag (e.g., subpic_treated_as_pic_flag) in the SPS. For example, when subpic_treated_as_pic_flag has a first value (e.g., 0), the current sub-picture cannot be independently coded. In contrast, when subpic_treated_as_pic_flag has a second value (e.g., 1), the current sub-picture can be independently coded.

[0210] When the current sub-picture is independently coded (\"YES\" in S1710), the image decoding apparatus can determine whether wrap-around motion compensation is available for the current block (S1720).

[0211] Whether wrap-around motion compensation is available for the current block can be determined based on a predetermined variable (e.g., refWraparoundEnabledFlag). For example, when refWraparoundEnabledFlag has a first value (e.g., 0), wrap-around motion compensation may not be available for the current block. In contrast, when refWraparoundEnabledFlag has a second value (e.g., 1), wrap-around motion compensation may be available for the current block. In one example, refWraparoundEnabledFlag can be set as shown in Equation 2 below.

[0212]

Equation

[0213] Here, the pps_ref_wraparound_enabled_flag can indicate whether wraparound motion compensation is available for the current picture including the current block (1 means available, 0 means not available). The pps_ref_wraparound_enabled_flag can be obtained via a higher-level syntax, such as a Picture Parameter Set (PPS). Also, the variable refPicIsScaled can indicate whether reference picture scaling is performed (or whether reference picture scaling is necessary). For example, when reference picture scaling is performed, refPicIsScaled can have a first value (e.g., 0), and when reference picture scaling is not performed, refPicIsScale can have a second value (e.g., 1).

[0214] Referring to Equation 2, if wraparound motion compensation is not available for the current picture (e.g., pps_ref_wraparound_enabled_flag == 0), or when reference picture scaling is performed (e.g., refPicIsScaled == 1), refWraparoundEnabledFlag can have a first value (e.g., 0) indicating that wraparound motion compensation is not available for the current block. In contrast, if wraparound motion compensation is available for the current picture (e.g., pps_ref_wraparound_enabled_flag == 1) and reference picture scaling is not performed (e.g., refPicIsScaled == 0), refWraparoundEnabledFlag can have a second value (e.g., 1) indicating that wraparound motion compensation is available for the current block. Thus, wraparound motion compensation for the reference picture and reference picture scaling can be selectively performed.

[0215] In one embodiment, a variable SliceRefWraparoundEnabledFlag is defined in relation to whether wraparound motion compensation is available and can be derived as follows.

[0216] - SliceRefWraparoundEnabledFlag can be set to be the same as the aforementioned pps_ref_wrap_around_enabled_flag.

[0217] - If SliceRefWraparoundEnabledFlag has a second value (e.g., 1), and for the current subpicture (i.e., the subpicture to which the current slice belongs), subpic_treated_as_pic_flag[i] has a second value (e.g., 1) (i.e., when the current subpicture is coded independently), and the width of the current subpicture is different from the width of the picture, the value of SliceRefWraparoundEnabledFlag can be set to a first value (e.g., 0).

[0218] - If SliceRefWraparoundEnabledFlag has a second value (e.g., 1) and the reference picture is not scaled, SliceRefWraparoundEnabledFlag can be used to activate the aforementioned refWraparoundEnabledFlag (i.e., refWraparoundEnabledFlag = 1).

[0219] The derivation process of the above-mentioned SliceRefWraparoundEnabledFlag can be defined as semantics within the slice header. A specific example is as shown in Table 5.

[0220]

Table 5

[0221] Referring to Table 5, the SliceRefWraparoundEnabledFlag can be set to the same value as pps_ref_wraparound_enabled_flag. And when the SliceRefWraparoundEnabledFlag has the second value (e.g., 1), subpic_treated_as_pic_flag[CurrSubpicIdx] has the second value (e.g., 1), and subpic_width_minus1[CurrSubpicIdx] + 1 is smaller than Ceil(pic_width_in_luma_samples ÷ CtbSizeY), the value of the SliceRefWraparoundEnabledFlag can be reset to the first value (e.g., 0).

[0222] When wrap-around motion compensation is available for the current block (\"YES\" in S1720), the image decoding apparatus can correct the position of the reference sample using the wrap-around offset (S1730).

[0223] The wrap-around offset can be set to the ERP width before padding of the current picture. Here, the ERP width can mean the width of the original picture (i.e., the ERP picture) in the ERP format obtained from the 360-degree picture. The wrap-around offset can be determined based on a predetermined syntax element (e.g., pps_ref_wraparound_offset) obtained via a higher-level syntax, such as the picture parameter set (PPS). And the x coordinate of the reference sample can be shifted in the positive or negative direction using the wrap-around offset.

[0224] The image decoding apparatus can clip the position of the corrected reference sample based on the sub-picture boundary (S1740). Thereby, the position of the reference sample can be included within the sub-picture boundary.

[0225] In contrast, when wrap-around motion compensation is not available for the current block (``NO'' in S1720), the video decoder can clip the position of the reference sample based on the sub-picture boundary (S1750). As a result, the position of the reference sample can be included within the sub-picture boundary.

[0226] As described above, when the current sub-picture is independently coded, the clipping operation for the position of the reference sample can be performed based on the boundary of the sub-picture regardless of whether wrap-around motion compensation is available for the current block (S1740, S1750). This may be due to the fact that an independently coded sub-picture has continuity based on the boundary of the sub-picture rather than the boundary of the reference picture.

[0227] Returning to step S1710 again, when the current sub-picture is not independently coded (``NO'' in S1710), the video decoder can determine whether wrap-around motion compensation is available for the current block (S1760). As described above, whether wrap-around motion compensation is available for the current block can be determined based on a predetermined variable (e.g., refWraparoundEnabledFlag).

[0228] When wrap-around motion compensation is available for the current block (``YES'' in S1760), the video decoder can correct the position of the reference sample using the wrap-around offset (S1770). Then, the video decoder can clip the corrected position of the reference sample based on the reference picture boundary (S1780). As a result, the position of the reference sample can be included within the reference picture boundary.

[0229] In contrast, when wrap-around motion compensation is not available for the current block (\"NO\" in S1760), the image decoding apparatus can clip the position of the reference sample based on the reference picture boundary (S1790).

[0230] As described above, when the current sub-picture is not independently coded, the clipping operation for the position of the reference sample can be performed based on the reference picture boundary regardless of whether wrap-around motion compensation is available for the current block (S1780, S1790). This may be due to the fact that a sub-picture that is not independently coded has continuity based on the reference picture boundary rather than the sub-picture boundary.

[0231] On the other hand, the above-described clipping operation can be performed using any one of a luma sample bilinear interpolation procedure, a luma sample interpolation filtering procedure, a luma integer sample fetching procedure, and a chroma sample interpolation procedure. And according to an embodiment of the present disclosure, wrap-around motion compensation can also be performed when the current sub-picture is independently coded. Therefore, Process 1 in Tables 1 to 4 described above can be modified as shown in Tables 6 to 9 below.

[0232]

Table 6

[0233]

Table 7

[0234]

Table 8

[0235]

Table 9

[0236] Referring to Tables 6 to 9, unlike the existing clipping operation, when the current sub-picture is independently coded, if wraparound motion compensation is available, the position of the reference sample can be clipped based on the sub-picture boundary (A610 to A910).

[0237] Specifically, when the current sub-picture is independently coded, Process 1 can be performed. Depending on whether wraparound motion compensation for the current block is available for the x coordinate of the reference sample (e.g., refWraparoundEnabledFlag == 1), the wraparound offset (e.g., PpsRefWraparoundOffset) can be selectively applied, and a clipping operation based on the sub-picture boundary can be performed on the selectively applied x coordinate. However, for the y coordinate of the reference sample, as described above with reference to Tables 1 to 4, a general clipping operation (or padding operation) based on the sub-picture boundary can be performed. This can mean that wraparound motion compensation is applied only to the left and right boundaries of the reference picture (i.e., horizontal wraparound motion compensation).

[0238] In Table 9, the xOffset, which is the wraparound offset used for wraparound motion compensation using the chroma sample interpolation procedure (see Table 8), can be derived as shown in Equation 3 below.

[0239]

Equation

[0240] Referring to Equation 3, the wrap-around offset (xOffset) can be calculated by multiplying the offset information (PpsRefWraparoundOffset) obtained from a higher-level syntax, such as a Picture Parameter Set (PPS), by the minimum width (MinCbSizeY) of a coding block (CB), and then dividing the result by the width ratio of luma-chroma samples (SubWidthC).

[0241] In one embodiment, when the variable SliceRefWraparoundEnabledFlag is newly defined as described above, the refWraparoundEnabledFlag in Tables 6 to 9 can be derived as shown in Equation 4 below without using the pps_ref_wraparound_enabled_flag obtained from the PPS.

[0242] [Equation]

[0243] Referring to Equation 4, when SliceRefWraparoundEnabledFlag has a first value (e.g., 0), or refPicIsScaled has a second value (e.g., 1), refWraparoundEnabledFlags can be set to the first value (e.g., 0). In contrast, when SliceRefWraparoundEnabledFlag has a second value (e.g., 1) and refPicIsScaled has a first value (e.g., 0), refWraparoundEnabledFlags can be set to the second value (e.g., 1).

[0244] In one embodiment, variables LeftBoundaryPos, RightBoundaryPos, TopBoundaryPos, and / or BottomBoundaryPos indicating the boundaries of the subpicture can be newly defined. The new variables can indicate the positions of the boundaries of the subpicture when subpic_treated_as_flag has a first value (e.g., 0) or a second value (e.g., 1). A specific example of the method for deriving the new variables is as shown in Table 10.

[0245] [Table 10]

[0246] Referring to Table 10, the initial value of the variable LeftBoundaryPos can be set to 0. Also, the initial value of the variable RightBoundaryPos can be set to pic_width_in_luma_samples - 1. Also, the initial value of the variable TopBoundaryPos can be set to 0. Also, the initial value of BottomBoundaryPos can be set to pic_height_in_luma_samples - 1.

[0247] When the current sub-picture is independently coded (i.e., subpic_treated_as_pic_flag[CurrSubpicIdx] == 1), the values of LeftBoundaryPos, RightBoundaryPos, TopBoundaryPos, and BottomBoundaryPos can be updated. Specifically, the value of LeftBoundaryPos can be updated to subpic_ctu_top_left_x[CurrSubpicIdx] * CtbSizeY. Also, the value of RightBoundaryPos can be updated to Min(pic_width_max_in_luma_samples - 1, (subpic_ctu_top_left_x[CurrSubpicIdx] + subpic_width_minus1[CurrSubpicIdx] + 1) * CtbSizeY - 1). Here, Min(x, y) means a function that outputs the smaller value of x and y. Also, the value of TopBoundaryPos can be updated to subpic_ctu_top_left_y[CurrSubpicIdx] * CtbSizeY. Also, the value of BottomBoundaryPos can be updated to Min(pic_height_max_in_luma_samples - 1, (subpic_ctu_top_left_y[CurrSubpicIdx] + subpic_height_minus1[CurrSubpicIdx] + 1) * CtbSizeY - 1).

[0248] The procedures in Tables 6 to 9 described above can be modified as shown in Tables 11 to 14 below using new variables related to the sub-picture boundary.

[0249]

Table 11

[0250]

Table 12

[0251]

Table 13

[0252]

Table 14

[0253] Each procedure of Tables 11 to 14 is the same as each procedure of Tables 6 to 9 except for using new variables related to sub-picture boundaries, so duplicate explanations are omitted. On the other hand, procedures such as the TMVP (temporal motion vector prediction) derivation procedure, the sub-block-based temporal merge candidate derivation procedure, and the restored affine control point motion vector merge candidate derivation procedure can also be performed using the new variables related to sub-picture boundaries.

[0254] FIG. 18 is a diagram showing an example of an SPS according to an embodiment of the present invention. Duplicate explanations overlapping with the SPS described above with reference to FIGS. 9 and 14a are omitted.

[0255] Referring to FIG. 18, the SPS can include a flag sps_ref_wraparound_enabled_flag indicating whether wrap-around motion compensation is available at the sequence level. For example, the sps_ref_wraparound_enabled_flag with a first value (e.g., 0) can indicate that (horizontal) wrap-around motion compensation is applied to inter prediction. In contrast, the sps_ref_wraparound_enabled_flag with a second value (e.g., 1) can indicate that (horizontal) wrap-around motion compensation is not applied.

[0256] If the value obtained by adding 1 to the syntax element subpic_width_minus1[i] indicating the width of the i-th subpicture is different from (pic_width_max_in_luma_samples + CtbSizeY - 1) >> CtbLog2SizeY) for all subpictures within the picture, it may be a requirement for bitstream compliance that the value of sps_ref_wraparound_enabled_flag is forced to a first value (e.g., 0). Here, i is 0 or more and may be less than or equal to the value of the syntax element sps_num_subpics_minus1 indicating the number of subpictures within the picture.

[0257] In one embodiment, sps_ref_wraparound_enabled_flag can be signaled prior to syntax elements related to subpictures (e.g., subpic_info_present_flag, sps_num_subpics_minus1, etc.). And when sps_ref_wraparound_enabled_flag has a second value (e.g., 1), syntax elements related to the position and width of subpictures (e.g., subpic_ctu_top_left_x[i], subpic_ctu_top_left_y[i], and subpic_width_minus1[i]) can be not signaled.

[0258] In this case, the value of subpic_ctu_top_left_x[i], which is a syntax element representing the horizontal position of the top left CTU of the i-th subpicture in units of CtbSizeY, can be inferred as the first value (e.g., 0). Also, subpic_ctu_top_left_y[i], which is a syntax element representing the vertical position of the top left CTU of the i-th subpicture in units of CtbSizeY, can be inferred as subpic_ctu_top_left_y[i-1]+subpic_height_minus1[i-1]. Here, subpic_height_minus1[i-1] can indicate the value obtained by subtracting 1 from the height of the (i-1)-th subpicture. Also, the syntax element subpic_width_minus1[i] indicating the width of the subpicture can be inferred as ((pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)-1. Here, pic_width_max_in_luma_samples indicates the maximum width of the picture in luma sample units, CtbSizeY indicates the CTB (coding tree block) width, and CtbLog2SizeY can indicate the log scale value of the CTB width.

[0259] In other embodiments, the syntax elements regarding the position and width of the subpicture described above can be restricted to be inferred as the above-described values when the sps_ref_wraparound_enabled_flag has a second value (e.g., 1).

[0260] Continuing to refer to FIG. 18, subpic_ctu_top_left_x[i], subpic_ctu_top_left_y[i], and subpic_width_minus1[i] can be signaled only when the sps_ref_wraparound_enabled_flag has a first value (e.g., 0) (1810).

[0261] subpic_ctu_top_left_x[i] can represent the horizontal position of the top - left CTU of the i - th sub - picture in units of CtbSizeY. The length of subpic_ctu_top_left_x[i] can be Ceil(Log2((pic_width_max_in_luma_samples + CtbSizeY - 1)>>CtbLog2SizeY)) bits. If subpic_ctu_top_left_x[i] does not exist in the bit - stream (i.e., is not signaled), the value of subpic_ctu_top_left_x[i] can be inferred as 0.

[0262] Also, subpic_ctu_top_left_y[i] can represent the vertical position of the top - left CTU of the i - th sub - picture in units of CtbSizeY. The length of subpic_ctu_top_left_y[i] can be Ceil(Log2((pic_height_max_in_luma_samples + CtbSizeY - 1)>>CtbLog2SizeY)) bits. If subpic_ctu_top_left_y[i] does not exist in the bit - stream (i.e., is not signaled), the following applies.

[0263] - If the i value is greater than 0 and the sps_ref_wraparound_enabled_flag has a second value (e.g., 1), the value of subpic_ctu_top_left_y[i] can be inferred to be the same as (subpic_ctu_top_left_y[i - 1]+subpic_height_minus1[i - 1]+1).

[0264] - Otherwise, the value of subpic_ctu_top_left_y[i] can be inferred to be the same as (((pic_width_max_in_luma_samples + CtbSizeY - 1)>>CtbLog2SizeY)-subpic_ctu_top_left_x[i]-1).

[0265] Also, the value obtained by adding 1 to subpic_width_minus1[i] can represent the width of the i-th subpicture in units of CtbSizeY. The length of subpic_width_minus1[i] can be Ceil(Log2((pic_width_max_in_luma_samples + CtbSizeY - 1) >> CtbLog2SizeY)) bits. If subpic_width_minus1[i] does not exist in the bitstream (i.e., is not signaled), the following applies.

[0266] - If sps_ref_wraparound_enabled_flag has a second value (e.g., 1), the value of subpic_width_minus1[i] can be inferred to be the same as (((pic_width_max_in_luma_samples + CtbSizeY - 1) >> CtbLog2SizeY) - 1).

[0267] - Otherwise, the value of subpic_width_minus1[i] can be inferred to be the same as (((pic_width_max_in_luma_samples + CtbSizeY - 1) >> CtbLog2SizeY) - subpic_ctu_top_left_x[i] - 1).

[0268] On the other hand, according to another embodiment, the above-described inference conditions for subpic_ctu_top_left_x[i], subpic_ctu_top_left_y[i], and subpic_width_minus1[i] can be enforced. For example, when sps_ref_wraparound_enabled_flag has a second value (e.g., 1), the inference that the value of subpic_ctu_top_left_x[i] is a first value (e.g., 0) can be a constraint condition for bitstream compliance. Also, when sps_ref_wraparound_enabled_flag has a second value (e.g., 1), the inference that subpic_ctu_top_left_y[i] is subpic_ctu_top_left_y[i-1]+subpic_height_minus1[i-1] can be a constraint condition for bitstream compliance. Also, when sps_ref_wraparound_enabled_flag has a second value (e.g., 1), the inference that subpic_width_minus1[i] is ((pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)-1 can be a constraint condition for bitstream compliance.

[0269] As described above, according to the embodiments of the present disclosure, regardless of whether the current subpicture is independently coded or not, wrap-around motion compensation can be applied to the current subpicture. At this time, when the current subpicture is independently coded, the wrap-around motion compensation can be performed based on the subpicture boundary. In contrast, when the current subpicture is not independently coded, the wrap-around motion compensation can be performed based on the reference picture boundary. Thereby, since subpicture-related coding tools and wrap-around motion compensation-related coding tools can be used together, the coding / decoding efficiency can be further improved.

[0270] Hereinafter, with reference to FIGS. 19 and 20, an image encoding / decoding method according to an embodiment of the present disclosure will be described in detail.

[0271] FIG. 19 is a flowchart showing an image encoding method according to an embodiment of the present disclosure.

[0272] The image encoding method of FIG. 19 can be performed by the image encoding apparatus of FIG. 2. For example, steps S1910 and S1920 can be performed by the inter prediction unit 180, and step S1930 can be performed by the entropy encoding unit 190.

[0273] Referring to FIG. 19, the image encoding apparatus can determine whether to apply wrap-around motion compensation to the current block (S1910).

[0274] In one embodiment, the image encoding apparatus determines whether wrap-around motion compensation is available based on whether there is one or more sub-pictures that are independently coded within the current video sequence including the current block and have a width different from the picture width. For example, if there is one or more sub-pictures that are independently coded within the current video sequence and have a width different from the picture width, the image encoding apparatus can determine that wrap-around motion compensation is not available. In contrast, if all sub-pictures independently coded within the current video sequence have the same width as the picture width, the image encoding apparatus can determine that wrap-around motion compensation is available based on a predetermined wrap-around constraint condition. Here, an example of the wrap-around constraint condition is as described above with reference to FIGS. 16 to 18.

[0275] Then, based on the determination, the image encoding device can determine whether to apply wrap-around motion compensation to the current block. For example, when wrap-around motion compensation is not available for the current picture, the image encoding device can determine not to perform wrap-around motion compensation on the current block. In contrast, when wrap-around motion compensation is available for the current picture, the image encoding device can determine to perform wrap-around motion compensation on the current block.

[0276] In one embodiment, based on the fact that the current sub-picture is independently encoded and the width of the current sub-picture is different from the width of the current picture, wrap-around motion compensation for the current block can be skipped.

[0277] The image encoding device can generate a predicted block of the current block by performing inter prediction based on the determination result of step S1910 (S1920).

[0278] In one embodiment, wrap-around motion compensation for the current block can be performed based on either the boundary of the current sub-picture including the current block or the boundary of the reference picture of the current block, based on whether the current sub-picture including the current block is independently encoded. For example, when the current sub-picture is independently encoded (e.g., subpic_treated_as_pic_flag == 1), wrap-around motion compensation for the current block can be performed based on the boundary of the current sub-picture. In contrast, when the current sub-picture is not independently encoded (e.g., subpic_treated_as_pic_flag == 0), wrap-around motion compensation for the current block can be performed based on the boundary of the reference picture.

[0279] In one embodiment, the wrap-around motion compensation for the current block can be performed by changing the x coordinate of the reference block specified by the motion vector of the current block within the reference picture of the current block based on a predetermined wrap-around offset. In one example, the wrap-around offset can be set to the ERP width before padding of the current picture. Here, the ERP width can mean the width of the original picture in the ERP format (i.e., the ERP picture) obtained from the 360-degree image. Then, the wrap-around motion compensation can be performed by clipping the changed x coordinate of the reference block within the range of either the boundary position of the current sub-picture or the boundary position of the reference picture.

[0280] In one embodiment, the wrap-around motion compensation for the current block can be performed based on the left boundary position and the right boundary position, which are set based on whether the current sub-picture is independently coded. For example, when the current sub-picture is independently coded, the left boundary position and the right boundary position can be set to the left boundary position and the right boundary position of the current sub-picture. In contrast, when the current sub-picture is not independently coded, the left boundary position and the right boundary position can be set to the left boundary position and the right boundary position of the reference picture. Then, the wrap-around motion compensation for the current block can be performed based on the set left boundary position and right boundary position. In this way, by reflecting whether the current sub-picture is independently coded in the boundary position used to clip the position of the reference block, there is no need to separately determine whether the current sub-picture is independently coded. As a result, the motion compensation process for generating the predicted block of the current block can be simplified. The image encoding device can encode the inter-prediction information of the current block and the wrap-around information related to the wrap-around motion compensation to generate a bitstream (S1930).

[0281] In one embodiment, the wrap-around information may include a first flag (e.g., pps_ref_wraparound_enabled_flag) indicating whether wrap-around motion compensation is available for the current picture. The first flag may have a first value (e.g., 0) indicating that wrap-around motion compensation is not available for the current picture based on that wrap-around motion compensation is not available for the current video sequence including the current block (e.g., sps_ref_waraparound_enabled_flag == 0). Also, the first flag may have a first value (e.g., 0) indicating that wrap-around motion compensation for the current picture is not available based on a predetermined condition regarding the width of a CTB (coding tree block) in the current picture and the width of the current picture. For example, when the width of the CTB in the current picture (e.g., CtbSizeY) is larger than the width of the picture (e.g., pic_width_in_luma_samples), pps_ref_wraparound_enabled_flag may be limited to the first value (e.g., 0).

[0282] In one embodiment, the wrap-around information may further include a wrap-around offset (e.g., pps_ref_wraparound_offset) based on that the wrap-around motion compensation is available for the current picture. The image encoding device may perform wrap-around motion compensation based on the wrap-around offset.

[0283] FIG. 20 is a flowchart showing an image decoding method according to an embodiment of the present disclosure.

[0284] The image decoding method of FIG. 20 can be performed by the image decoding device of FIG. 3. For example, steps S2010 and S2020 can be performed by the inter prediction unit 260.

[0285] Referring to FIG. 20, the image decoding apparatus can obtain the inter prediction information and the wrap-around information of the current block from the bitstream (S2010). Here, the inter prediction information of the current block can include the motion information of the current block, for example, the reference picture index, the differential motion vector information, and the like. The wrap-around information can include a first flag (for example, pps_ref_wraparound_enabled_flag) indicating whether wrap-around motion compensation is available for the current picture including the current block. The first flag can have a first value (for example, 0) indicating that wrap-around motion compensation is not available for the current picture based on the fact that wrap-around motion compensation is not available for the current video sequence including the current block (for example, sps_ref_wraparound_enabled_flag == 0). Also, the first flag can have a first value (for example, 0) indicating that wrap-around motion compensation is not available for the current picture based on a predetermined condition regarding the width of the CTB (coding tree block) in the current picture and the width of the current picture. For example, when the width of the CTB in the current picture (for example, CtbSizeY) is larger than the width of the picture (for example, pic_width_in_luma_samples), pps_ref_wraparound_enabled_flag can be limited to the first value (for example, 0).

[0286] In one embodiment, based on the fact that wraparound motion compensation is available for the current video sequence including the current block (e.g., sps_wraparound_enabled_flag == 1), the upper left position and width of the current subpicture can be set (or inferred) to predetermined values. For example, the x coordinate of the upper left position of the current subpicture (e.g., subpic_ctu_top_left_x[i]) is set to 0, and the y coordinate of the upper left position of the current subpicture (e.g., subpic_ctu_top_left_y[i]) can be set to a value obtained by adding the height of the first subpicture decoded before the current subpicture (e.g., subpicture index = i - 1) (e.g., subpic_height_minus1[i - 1]) to the y coordinate of the first subpicture (e.g., subpic_ctu_top_left_y[i - 1]). Also, the width of the current subpicture (e.g., subpic_width_minus[i]) can be set to the same value as the width of the current picture (e.g., ((pic_width_max_in_luma_samples + CtbSizeY - 1) >> CtbLog2SizeY) - 1).

[0287] The image decoding apparatus can generate a predicted block of the current block based on the inter prediction information and wraparound information obtained from the bitstream (S2020).

[0288] In one embodiment, based on the fact that the above-described first flag has a predetermined value (e.g., 1) indicating that wraparound motion compensation is available for the current picture, the predicted block of the current block can be generated by wraparound motion compensation.

[0289] In one embodiment, the wrap-around motion compensation for the current block can be performed by modifying the x coordinate of the reference block specified by the motion vector of the current block within the reference picture of the current block based on the wrap-around offset. Then, the wrap-around motion compensation can be performed by clipping the x coordinate of the modified reference block within the range of either the boundary position of the current sub-picture or the boundary position of the reference picture. At this time, the boundary position serving as the clipping criterion can be determined based on whether the current sub-picture including the current block is independently coded (e.g., subpic_treatead_as_pic_flag). For example, when the current sub-picture is independently coded (e.g., subpic_treated_as_pic_flag == 1), the x coordinate of the reference block can be clipped based on the boundary of the current sub-picture. In contrast, when the current sub-picture is not independently coded (e.g., subpic_treated_as_pic_flag == 0), the x coordinate of the reference block can be clipped based on the boundary of the reference picture.

[0290] On the other hand, in one embodiment, based on the fact that the reference picture of the current block is scaled, the wrap-around motion compensation for the current block can be skipped. Also, based on the fact that the current sub-picture is independently coded and the width of the current sub-picture is different from the width of the current picture, the wrap-around motion compensation for the current block can be skipped.

[0291] In one embodiment, the wrap-around motion compensation for the current block can be performed based on the left boundary position and the right boundary position, which are set based on whether the current sub-picture is independently coded. For example, when the current sub-picture is independently coded, the left boundary position and the right boundary position can be set to the left boundary position and the right boundary position of the current sub-picture. In contrast, when the current sub-picture is not independently coded, the left boundary position and the right boundary position can be set to the left boundary position and the right boundary position of the reference picture. Then, the wrap-around motion compensation for the current block can be performed based on the set left boundary position and right boundary position. In this way, whether the current sub-picture is independently coded is reflected in the boundary position used to clip the reference block, so there is no need to separately determine whether the current sub-picture is independently coded. As a result, the motion compensation process for generating the predicted block of the current block can be more simplified.

[0292] As described above, according to the image encoding / decoding method according to an embodiment of the present disclosure, wrap-around motion compensation can be used according to predetermined conditions even when the current sub-picture is independently coded. As a result, sub-picture related coding tools and wrap-around motion compensation related coding tools can be used together, so that the encoding / decoding efficiency can be further improved.

[0293] The names of the syntax elements described in the present disclosure can include information regarding the position where the syntax element is signaled. For example, a syntax element starting with "sps_" can mean that the syntax element is signaled in a sequence parameter set (SPS). Further, syntax elements starting with "pps_", "ph_", "sh_", etc. can mean that the syntax element is signaled in a picture parameter set (PPS), a picture header, a slice header, etc., respectively.

[0294] Exemplary methods of the present disclosure are presented as a series of operations for clarity of explanation, but this is not intended to limit the order in which the steps are performed. If necessary, each step can also be performed simultaneously or in a different order. To implement the methods according to the present disclosure, other steps can be further included in the exemplary steps, or the remaining steps can be included excluding some steps, or additional other steps can be included excluding some steps.

[0295] In the present disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) can perform an operation (step) of checking the execution conditions and situations of the operation (step). For example, when it is described that a predetermined operation is performed if a predetermined condition is satisfied, the image encoding device or the image decoding device can perform the predetermined operation after performing an operation of checking whether the predetermined condition is satisfied.

[0296] The various embodiments of the present disclosure do not list all possible combinations, but are for explaining representative aspects of the present disclosure. The matters described in the various embodiments may be applied independently or in combinations of two or more.

[0297] Also, the various embodiments of the present disclosure can be implemented by hardware, firmware, software, or a combination thereof. In the case of implementation by hardware, it can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), general processors, controllers, microcontrollers, microprocessors, etc.

[0298] In addition, the image decoding device and the image encoding device to which the embodiments of the present disclosure are applied can be included in a multimedia broadcast transmission / reception device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an over-the-top video (OTT) device, an Internet streaming service providing device, a three-dimensional (3D) video device, an image phone video device, and a medical video device, etc., and can be used to process video signals or data signals. For example, as an over-the-top video (OTT) device, it can include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.

[0299] FIG. 21 is a diagram illustrating a content streaming system to which the embodiments of the present disclosure can be applied.

[0300] As shown in FIG. 21, the content streaming system to which the embodiments of the present disclosure are applied can generally include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0301] The encoding server compresses the content input from a multimedia input device such as a smartphone, a camera, or a camcorder into digital data to generate a bitstream, and plays a role of transmitting this to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, or a video camera directly generates a bitstream, the encoding server can be omitted.

[0302] The bitstream can be generated by an image encoding method and / or an image encoding apparatus to which the embodiments of the present disclosure are applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0303] The streaming server transmits multimedia data to a user device based on a user request via a Web server, and the Web server can serve as a medium for informing the user of available services. When the user requests a desired service from the Web server, the Web server transmits this to the streaming server, and the streaming server can transmit multimedia data to the user. At this time, the content streaming system can include a separate control server, and in this case, the control server can control commands / responses between each device in the content streaming system.

[0304] The streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0305] Examples of the user device may include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device, for example, a smartwatch, smart glass, an HMD (head mounted display), a digital TV, a desktop computer, a digital signage, and the like.

[0306] Each server in the content streaming system can be operated as a distributed server, and in this case, the data received from each server can be distributedly processed.

[0307] FIG. 22 is a diagram schematically showing an architecture for providing a three-dimensional image / video service that can utilize an embodiment of the present disclosure.

[0308] FIG. 22 can show a 360-degree or omnidirectional video / image processing system. Further, the system of FIG. 22 can be realized by, for example, an extended reality (XR) support device. That is, the system can provide a solution for providing virtual reality to a user.

[0309] Extended Reality refers to the general term for Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR). VR technology only provides objects and backgrounds in the real world as CG images. AR technology provides both real-world object images and virtual CG images created on top of them. MR technology is a computer graphics technology that mixes and combines virtual objects with the real world.

[0310] MR technology is similar to AR technology in that it shows real objects and virtual objects together. However, there is a difference in that in AR technology, virtual objects are used to complement real objects, while in MR technology, virtual objects and real objects are used with equal status.

[0311] XR technology can be applied to Head-Mount Displays (HMDs), Head-Up Displays (HUDs), mobile phones, tablets, laptop computers, desktops, TVs, digital signage, etc. Devices to which XR technology is applied can be called XR devices (XR Devices). An XR device can include a first digital device and / or a second digital device, which will be described later.

[0312] 360-degree content refers to all content for realizing and providing VR, and can include 360-degree videos and / or 360-degree audios. 360-degree videos can mean video or image content that is captured or played back simultaneously in all directions (360 degrees or less) and is necessary for providing VR. Hereinafter, 360-degree videos can refer to 360-degree videos. 360-degree audios are also audio content for providing VR, and can mean spatial audio content where the sound generation location can be recognized as being located in a specific three-dimensional space. 360-degree content can be generated, processed, and transmitted to users, and users can consume VR experiences using 360-degree content. 360-degree videos may be called omnidirectional videos, and 360-degree images may be called omnidirectional images. Also, hereinafter, an explanation will be given based on 360-degree videos, and the embodiments of this document are not limited to VR, and can include processing for video / image content such as AR and MR. 360-degree videos can mean videos or images displayed on various forms of 3D spaces according to 3D models. For example, 360-degree videos can be displayed on a spherical surface.

[0313] This method proposes a method for effectively providing 360-degree videos in particular. To provide 360-degree videos, first, 360-degree videos can be captured via one or more cameras. The captured 360-degree videos are transmitted through a series of processes, and on the receiving side, the received data can be processed back into the original 360-degree videos for rendering. Thereby, 360-degree videos can be provided to users.

[0314] Specifically, the overall process for providing 360-degree videos can include a capture process, a preparation process, a transmission process, a processing process, a rendering process, and / or a feedback process.

[0315] The capture process can mean the process of capturing images or videos for each of a plurality of viewpoints via one or more cameras. Image / video data such as 2210 of FIG. 22 shown by the capture process can be generated. Each plane of 2210 of the illustrated FIG. 22 can mean an image / video for each viewpoint. These captured multiple images / videos can also be called raw data. Metadata related to the capture can be generated during the capture process.

[0316] For this capture, special cameras for VR can be used. According to an embodiment, when attempting to provide a 360-degree video for a virtual space generated by a computer, capture via an actual camera may not be possible. In this case, the capture process can simply be replaced by a process in which relevant data is generated.

[0317] The preparation process can be a process of processing the captured images / videos and the metadata generated during the capture process. The captured images / videos can undergo a stitching process, a projection process, a region-wise packing process, and / or an encoding process, etc. during the preparation process.

[0318] First, each image / video can undergo a stitching process. The stitching process can be a process of connecting each captured image / video to create one panoramic image / video or a spherical image / video.

[0319] After that, the stitched image / video can go through the Projection process. In the projection process, the stitched image / video can be projected onto a 2D image. This 2D image may sometimes be called a 2D image frame depending on the context. Projecting onto a 2D image can also be expressed as mapping to a 2D image. The projected image / video data can be in the form of a 2D image such as 2220 in FIG. 22.

[0320] The video data projected onto the 2D image can go through a Region-wise Packing process to improve video coding efficiency, etc. Region-wise Packing can mean a process of dividing the video data projected onto the 2D image into regions and processing them separately. Here, a Region can mean the area into which the 2D image onto which 360-degree video data is projected is divided. These regions can be divided and segmented by equally dividing the 2D image or arbitrarily dividing it according to the embodiment. Also, according to the embodiment, these regions can be distinguished by the projection scheme. The Region-wise Packing process is an optional process and can be omitted in the preparation process.

[0321] According to the embodiment, this processing process can include a process of rotating each region or rearranging it on the 2D image to improve video coding efficiency. For example, by rotating these regions so that specific sides of the regions are close to each other, the coding efficiency can be increased.

[0322] According to an embodiment, this processing process can include a process of increasing or decreasing the resolution for a specific region in order to equalize the resolution for each region on the 360-degree video. For example, a region corresponding to a relatively more important region on the 360-degree video can have a higher resolution than other regions. The video data projected on the 2D image or the video data packed for each region can go through an encoding process via a video codec.

[0323] According to an embodiment, the preparation process can further include an editing process and the like. In this editing process, editing of the image / video data before and after projection can be further performed. Similarly, in the preparation process, metadata for stitching / projection / encoding / editing, etc. can be generated. Also, metadata regarding the initial viewpoint of the video data projected on the 2D image, or the ROI (Region of Interest), etc. can be generated.

[0324] The transmission process can be a process of processing and transmitting the image / video data and metadata that have gone through the preparation process. Processing by any transmission protocol can be performed for transmission. The data that has completed the processing for transmission can be transmitted via a broadcast network and / or broadband. These data can also be transmitted to the receiving side in an on-demand manner. On the receiving side, the corresponding data can be received via various paths.

[0325] The processing process can be meant to decrypt the received data and re-project the projected image / video data onto a 3D model. In this process, the image / video data projected on a 2D image can be re-projected onto a 3D space. This process may also be called mapping or projection according to the context. At this time, the 3D space to be mapped can have different forms depending on the 3D model. For example, the 3D model can be a sphere, a cube, a cylinder, or a pyramid, etc.

[0326] According to the embodiment, the processing process can further include an editing process, an upscaling process, etc. In this editing process, editing of the image / video data before and after re-projection can be further performed. When the image / video data is reduced, its size can be enlarged through sample upscaling in the upscaling process. If necessary, the operation of reducing the size by downscaling can also be performed.

[0327] The rendering process can be meant to render and display the image / video data re-projected onto a 3D space. Depending on the representation, it can also be expressed as rendering onto a 3D model by combining re-projection and rendering. The image / video re-projected (or rendered) onto the 3D model can have a form like 2230 in FIG. 22 shown. FIG. 22 shown is the case of being re-projected onto a spherical 3D model. The user can view a partial area of the image / video rendered through a VR display or the like. At this time, the area viewed by the user can be in a form like 2240 in FIG. 22 shown.

[0328] The feedback process can be meant as a process of transmitting various feedback information that can be obtained in the display process to the transmitting side. Through the feedback process, interactivity can be provided in 360-degree video consumption. According to an embodiment, in the feedback process, head orientation information, viewport information indicating the area that the user is currently viewing, etc. can be transmitted to the transmitting side. According to an embodiment, the user can also interact with what is realized on the VR environment, and in this case, information related to the interaction can also be transmitted to the transmitting side or the service provider side in the feedback process. According to an embodiment, the feedback process can also not be performed.

[0329] Head orientation information can mean information regarding the position, angle, movement, etc. of the user's head. Based on this information, information regarding the area that the user is currently viewing within the 360-degree video, that is, viewport information, can be calculated.

[0330] Viewport information can be information regarding the area that the user is currently viewing in the 360-degree video. By performing gaze analysis, it is also possible to confirm how the user consumes the 360-degree video, how long the user stares at which area of the 360-degree video, etc. Gaze analysis is performed on the receiving side and can also be transmitted to the transmitting side via the feedback channel. Devices such as VR displays can extract the viewport area based on the position / direction of the user's head, vertical or horizontal FOV (Field Of View) information supported by the device, etc.

[0331] On the one hand, 360-degree videos / images can be processed based on sub-pictures. A projected picture or a packed picture including a 2D image can be divided into sub-pictures, and processing can be performed in units of sub-pictures. For example, a high resolution can be given to a specific sub-picture by a user viewport or the like, or only a specific sub-picture can be encoded and signaled to a receiving device (on the decoder side). In this case, the decoder can receive a sub-picture bit stream, restore / decode the specific sub-picture, and render it according to the user viewport.

[0332] According to an embodiment, the above-described feedback information can not only be transmitted to the transmission side, but also be consumed on the reception side. That is, the above-described feedback information can be used to perform processes such as decoding, reprojection, and rendering on the reception side. For example, using the head orientation information and / or the viewport information, only the 360-degree video for the area currently viewed by the user can be preferentially decoded and rendered.

[0333] Here, a viewport or a viewport area can mean an area that a user is viewing in a 360-degree video. A viewpoint can be a point that a user is viewing in a 360-degree video and can mean the center point of the viewport area. That is, the viewport is an area centered on the viewpoint, but the size and form occupied by the area can be determined by the FOV (Field Of View).

[0334] In the overall architecture for providing the above-described 360-degree video, image / video data that goes through a series of processes of capture / projection / encoding / transmission / decoding / reprojection / rendering can be called 360-degree video data. Also, the term 360-degree video data may be used as a concept including metadata or signaling information related to such image / video data.

[0335] To store and transmit the above-described media data such as audio or video, a standardized media file format can be defined. According to an embodiment, the media file can have a file format based on ISO BMFF (ISO base media file format).

[0336] The scope of the present disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) that enable the operations of various example methods to be executed on a device or computer, and non-transitory computer-readable media on which such software or commands are stored and can be executed on a device or computer.

Industrial Applicability

[0337] Examples according to the present disclosure can be used to encode / decrypt images.

Claims

1. An image decoding method performed by an image decoding device, comprising: obtaining inter prediction information and wraparound information of a current block from a bitstream; generating a prediction block of the current block based on the inter prediction information and the wraparound information; the wraparound information includes a first flag indicating whether wraparound motion compensation is available for a current picture that includes the current block; the first flag has a first value indicating that the wraparound motion compensation is not available for the current picture based on at least one sub-picture included in the current picture being treated as a picture and having a width different from a value derived based on information about a maximum width of a picture in a current video sequence that includes the current picture; based on the first flag having a second value indicating that the wraparound motion compensation is available for the current picture, the wraparound information further comprises a wraparound offset used to perform the wraparound motion compensation; 11. The image decoding method of claim 10, wherein, depending on whether a current sub-picture is treated as a picture, the wrap-around motion compensation is performed based on a boundary of the current sub-picture containing the current block or a boundary of a reference picture of the current block.

2. the wraparound motion compensation for the current block is performed based on the wraparound offset; The image decoding method according to claim 1 , wherein the wraparound offset is obtained from the bitstream based on the first flag having the second value.

3. the wraparound motion compensation for the current block is performed by modifying an x-coordinate of a reference block in the reference picture based on the wraparound offset; The image decoding method according to claim 2 , wherein the reference block is identified by a motion vector of the current block.

4. The image decoding method of claim 3 , wherein the wraparound motion compensation for the current block is performed by clipping the modified x-coordinate of the reference block within a boundary of the current sub-picture or the reference picture.

5. The image decoding method of claim 1 , wherein the wraparound motion compensation for the current block is skipped based on the reference picture being scaled.

6. The image decoding method of claim 1 , wherein the wraparound motion compensation for the current block is skipped based on the current sub-picture being independently coded and having a width different from a width of the current picture.

7. the wraparound motion compensation for the current block is performed based on a left boundary position and a right boundary position; The image decoding method of claim 1 , wherein the left and right boundary positions are set based on whether the current sub-picture is independently coded.

8. The image decoding method of claim 1 , wherein a top-left position and a width of the current sub-picture are set to predetermined values ​​based on the wrap-around motion compensation being available for a current video sequence that includes the current block.

9. An image coding method performed by an image coding device, comprising: determining whether wraparound motion compensation is applied to the current block; generating a predicted block of the current block by performing inter prediction based on the determination; encoding inter prediction information of the current block and wrap-around information related to the wrap-around motion compensation; the wraparound information includes a first flag indicating whether wraparound motion compensation is available for a current picture that includes the current block; the first flag has a first value indicating that the wraparound motion compensation is not available for the current picture based on at least one sub-picture included in the current picture being treated as a picture and having a width different from a value derived based on a maximum width of a picture in a current video sequence that includes the current picture; based on the first flag having a second value indicating that the wraparound motion compensation is available for the current picture, the wraparound information further comprises a wraparound offset used to perform the wraparound motion compensation; 11. A method for coding an image, wherein, depending on whether a current sub-picture is treated as a picture, the wrap-around motion compensation is performed based on a boundary of the current sub-picture containing the current block or a boundary of a reference picture of the current block.

10. the wraparound motion compensation for the current block is performed by modifying an x-coordinate of a reference block in the reference picture based on the wraparound offset; The image coding method according to claim 9 , wherein the reference block is identified by a motion vector of the current block.

11. The image coding method of claim 10 , wherein the wraparound motion compensation for the current block is performed by clipping the modified x-coordinate of the reference block within a boundary of the current sub-picture or the reference picture.

12. 10. The image coding method of claim 9, wherein the wraparound motion compensation for the current block is skipped based on the current sub-picture being independently coded and having a different width than a width of the current picture.

13. the wraparound motion compensation for the current block is performed based on a left boundary position and a right boundary position; The image encoding method of claim 9 , wherein the left and right boundary positions are set based on whether the current sub-picture is independently coded.

14. 1. A method for transmitting a bitstream, the method comprising: determining whether wraparound motion compensation is applied to the current block; generating a predicted block of the current block by performing inter prediction based on the determination; encoding inter prediction information of the current block and wrap-around information related to the wrap-around motion compensation in the bitstream; transmitting the bitstream to a receiving party; the wraparound information includes a first flag indicating whether wraparound motion compensation is available for a current picture that includes the current block; the first flag has a first value indicating that the wraparound motion compensation is not available for the current picture based on at least one sub-picture included in the current picture being treated as a picture and having a width different from a value derived based on a maximum width of a picture in a current video sequence that includes the current picture; based on the first flag having a second value indicating that the wraparound motion compensation is available for the current picture, the wraparound information further comprises a wraparound offset used to perform the wraparound motion compensation; A method according to claim 1, wherein, depending on whether a current subpicture is treated as a picture, the wraparound motion compensation is performed based on a boundary of the current subpicture containing the current block or a boundary of a reference picture of the current block.

Citation Information

Patent Citations

  • Image encoding / decoding method and device based on wraparound motion compensation, and recording medium storing bitstream

    JP7662876B2