Image encoding / decoding method and device based on wraparound motion compensation, and recording medium storing bitstream

The image encoding/decoding method and apparatus address the high cost of transmitting and storing high-resolution images by employing wraparound motion compensation, enhancing efficiency through improved encoding/decoding techniques for independently coded sub-pictures.

JP7801397B2Active Publication Date: 2026-01-16LG ELECTRONICS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024097509
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-04-14
Filing Date
2024-06-17
Publication Date
2026-01-16
Estimated Expiration
2041-03-26

AI Technical Summary

Technical Problem

The increasing demand for high-resolution, high-quality images leads to a significant increase in transmission and storage costs due to the higher amount of information, necessitating highly efficient image compression techniques.

Method used

An image encoding/decoding method and apparatus based on wraparound motion compensation, including inter-prediction information and wrap-around information, with flags indicating the availability of wrap-around motion compensation for current video sequences, especially for independently coded sub-pictures with different widths.

Benefits of technology

Improves encoding/decoding efficiency and enables efficient transmission and storage of high-resolution, high-quality images by utilizing wraparound motion compensation for independently coded sub-pictures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007801397000002
    Figure 0007801397000002
  • Figure 0007801397000003
    Figure 0007801397000003
  • Figure 0007801397000004
    Figure 0007801397000004
Patent Text Reader

Abstract

To provide an image encoding / decoding method and apparatus.SOLUTION: An image encoding / decoding method and apparatus are provided. An image decoding method includes obtaining inter prediction information of a current block and wraparound information from a bitstream, and generating a prediction block of the current block based on the inter prediction information and the wraparound information. The wraparound information may include a first flag specifying whether wraparound motion compensation is available for a current video sequence including the current block. The first flag may have a first value specifying that the wraparound motion compensation is not available, based on that one or more subpictures, which are coded independently and have a width different from a width of a current picture including the current block, are present in the current video sequence.SELECTED DRAWING: Figure 20
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly to an image encoding / decoding method and apparatus based on wraparound motion compensation, and a recording medium storing a bitstream generated by the image encoding method / apparatus of the present disclosure. [Background technology]

[0002] Recently, demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, has been increasing in various fields. As image data becomes higher in resolution and quality, the amount of information or bits to be transmitted increases relatively compared to conventional image data. The increase in the amount of information or bits to be transmitted results in an increase in transmission costs and storage costs.

[0003] This requires highly efficient image compression techniques for effectively transmitting, storing, and reproducing high-resolution, high-quality image information. Summary of the Invention [Problem to be solved by the invention]

[0004] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0005] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus based on wraparound motion compensation.

[0006] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus based on wraparound motion compensation for independently coded sub-pictures.

[0007] Another object of the present disclosure is to provide a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0008] Another object of the present disclosure is to provide a recording medium storing a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0009] Another object of the present disclosure is to provide a recording medium storing a bitstream that is received by an image decoding device according to the present disclosure, decoded, and used to restore an image.

[0010] The technical problems to be solved by the present disclosure are not limited to the above-mentioned technical problems, and other technical problems not mentioned above will be clearly understood by a person having ordinary skill in the technical field to which the present disclosure pertains from the following description. [Means for solving the problem]

[0011] An image decoding method according to one aspect of the present disclosure includes the steps of obtaining inter-prediction information and wrap-around information for a current block from a bitstream, and generating a predicted block for the current block based on the inter-prediction information and the wrap-around information, wherein the wrap-around information includes a first flag indicating whether wrap-around motion compensation is available for a current video sequence including the current block, and the first flag may have a first value indicating that the wrap-around motion compensation is not available based on the presence of one or more sub-pictures in the current video sequence that are independently coded and have a width different from the width of the current picture including the current block.

[0012] An image decoding device according to another aspect of the present disclosure includes a memory and at least one processor, wherein the at least one processor obtains inter-prediction information and wrap-around information for a current block from a bitstream, and generates a prediction block for the current block based on the inter-prediction information and the wrap-around information, and the wrap-around information includes a first flag indicating whether wrap-around motion compensation is available for a current video sequence including the current block, and the first flag may have a first value indicating that the wrap-around motion compensation is not available based on the presence of one or more sub-pictures in the current video sequence that are independently coded and have a width different from the width of the current picture including the current block.

[0013] An image encoding method according to another aspect of the present disclosure includes the steps of determining whether to apply wrap-around motion compensation to a current block, generating a predicted block of the current block by performing inter prediction based on the decision, and encoding inter prediction information of the current block and wrap-around information related to the wrap-around motion compensation, wherein the wrap-around information includes a first flag indicating whether the wrap-around motion compensation is available for a current video sequence including the current block, and the first flag may have a first value indicating that the wrap-around motion compensation is not available based on the presence of one or more sub-pictures in the current video sequence that are coded independently and have a width different from the width of the current picture including the current block.

[0014] A computer-readable recording medium according to another aspect of the present disclosure can store a bitstream generated by the image encoding method or image encoding device of the present disclosure.

[0015] A transmission method according to another aspect of the present disclosure can transmit a bitstream generated by the image encoding device or image encoding method of the present disclosure.

[0016] The features described above in this brief summary of the present disclosure are merely exemplary embodiments of the detailed description of the present disclosure that follows and are not intended to limit the scope of the present disclosure. [Effects of the Invention]

[0017] According to the present disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.

[0018] Furthermore, according to the present disclosure, an image encoding / decoding method and apparatus based on wraparound motion compensation can be provided.

[0019] Furthermore, according to the present disclosure, an image encoding / decoding method and apparatus based on wraparound motion compensation for independently coded sub-pictures can be provided.

[0020] The present disclosure also provides a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.

[0021] Furthermore, according to the present disclosure, a recording medium storing a bitstream generated by the image encoding method or apparatus according to the present disclosure can be provided.

[0022] The present disclosure may also provide a recording medium storing a bitstream that is received by an image decoding device according to the present disclosure, decoded, and used to restore an image.

[0023] The effects obtained by the present disclosure are not limited to the effects described above, and other effects not described above will be clearly understood by those having ordinary skill in the art to which the present disclosure pertains from the following description. [Brief explanation of the drawings]

[0024] [Figure 1] 1 is a diagram illustrating a video coding system to which embodiments of the present disclosure can be applied; [Figure 2] 1 is a diagram schematically illustrating an image encoding device to which an embodiment of the present disclosure can be applied. [Figure 3] FIG. 1 is a diagram schematically illustrating an image decoding device to which an embodiment of the present disclosure can be applied. [Figure 4] 1 is a flowchart illustrating an outline of an image decoding procedure to which an embodiment of the present disclosure can be applied. [Figure 5] 1 is a flowchart illustrating an outline of an image encoding procedure to which an embodiment of the present disclosure can be applied. [Figure 6] 1 is a flowchart illustrating an inter-prediction based video / image decoding method. [Figure 7] 10 is a diagram illustrating an example of the configuration of an inter prediction unit 260 according to the present disclosure. [Figure 8] FIG. 10 is a diagram showing an example of a sub-picture. [Figure 9] FIG. 10 is a diagram showing an example of an SPS including information about subpictures. [Figure 10] FIG. 1 is a diagram illustrating a method for encoding an image using sub-pictures by an image encoding device according to an embodiment of the present disclosure. [Figure 11] FIG. 10 is a diagram illustrating a method in which an image decoding device according to an embodiment of the present disclosure decodes an image using sub-pictures. [Figure 12] FIG. 10 is a diagram showing an example of a 360-degree image converted into a two-dimensional picture. [Figure 13] FIG. 1 illustrates an example of a horizontal wraparound motion compensation process. [Figure 14a] FIG. 10 is a diagram showing an example of an SPS including information about wraparound motion compensation. [Figure 14b] FIG. 10 is a diagram showing an example of a PPS including information about wrap-around motion compensation. [Figure 15] 10 is a flowchart illustrating a method for an image coding apparatus to perform wraparound motion compensation. [Figure 16] 10 is a flowchart illustrating a method for determining whether or not wraparound motion compensation is available by an image encoding device according to an embodiment of the present disclosure. [Figure 17] 10 is a flowchart illustrating a method for determining whether or not wraparound motion compensation is available by an image encoding device according to an embodiment of the present disclosure. [Figure 18] 10 is a flowchart illustrating a method for performing wraparound motion compensation by an image encoding device according to an embodiment of the present disclosure. [Figure 19] 1 is a flowchart illustrating an image encoding method according to an embodiment of the present disclosure. [Figure 20] 1 is a flowchart illustrating an image decoding method according to an embodiment of the present disclosure. [Figure 21] 1 is a diagram illustrating an exemplary content streaming system to which an embodiment of the present disclosure can be applied; [Figure 22] 1 is a diagram illustrating a schematic architecture for providing 3D image / video services that can be utilized by embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0025] The present disclosure will be described in detail below with reference to the accompanying drawings, so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein.

[0026] In describing the embodiments of the present disclosure, if it is determined that a detailed description of a known configuration or function may obscure the gist of the present disclosure, the detailed description thereof will be omitted. In addition, in the drawings, parts that are not related to the description of the present disclosure will be omitted, and similar parts will be designated by similar reference numerals.

[0027] In this disclosure, when a component is referred to as being "coupled," "coupled," or "connected" to another component, this includes not only a direct connection, but also an indirect connection where another component exists between them. Furthermore, when a component is referred to as "including" or "having" another component, this does not mean that the other component is excluded, but that the component can further include the other component, unless otherwise specified.

[0028] In this disclosure, terms such as "first" and "second" are used only to distinguish one component from another component, and do not limit the order or importance of the components unless otherwise specified. Therefore, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.

[0029] In this disclosure, components that are distinguished from one another are used to clearly describe the characteristics of each component and do not necessarily mean that the components are separate. In other words, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed into multiple hardware or software units. Therefore, even if not otherwise specified, such integrated or distributed embodiments are also included within the scope of this disclosure.

[0030] In this disclosure, the components described in various embodiments are not necessarily essential components, and some may be optional components. Therefore, an embodiment consisting of a subset of the components described in one embodiment is also within the scope of this disclosure. Furthermore, an embodiment including other components in addition to the components described in various embodiments is also within the scope of this disclosure.

[0031] The present disclosure relates to image encoding and decoding, and terms used in this disclosure may have their ordinary meaning in the technical field to which the present disclosure belongs unless they are newly defined in this disclosure.

[0032] In this disclosure, a "picture" generally refers to a unit representing any one image in a specific time period, and a slice / tile is a coding unit constituting a part of a picture, and one picture may be composed of one or more slices / tiles. Furthermore, a slice / tile may include one or more coding tree units (CTUs).

[0033] In this disclosure, "pixel" or "pel" may refer to the smallest unit constituting one picture (or image). Also, "sample" may be used as a term corresponding to pixel. A sample may generally indicate a pixel or a pixel value, may indicate only a pixel / pixel value of a luma component, or may indicate only a pixel / pixel value of a chroma component.

[0034] In this disclosure, the term "unit" may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to that region. The term "unit" may be used interchangeably with terms such as "sample array," "block," or "area," depending on the situation. In general, an M×N block may include a set (or array) of samples or transform coefficients consisting of M columns and N rows.

[0035] In the present disclosure, a "current block" may refer to any one of a "current coding block," a "current coding unit," a "block to be coded," a "block to be decoded," or a "block to be processed." When prediction is performed, a "current block" may refer to a "current predicted block" or a "block to be predicted." When transformation (inverse transformation) / quantization (inverse quantization) is performed, a "current block" may refer to a "current transformed block" or a "block to be transformed." When filtering is performed, a "current block" may refer to a "block to be filtered."

[0036] Furthermore, in this disclosure, unless explicitly stated as a chroma block, the term "current block" may refer to a block including both a luma component block and a chroma component block, or to the "luma block of the current block." The luma component block of the current block may be expressed by explicitly including the term "luma block" or "current luma block." The chroma component block of the current block may be expressed by explicitly including the term "chroma block" or "current chroma block."

[0037] In the present disclosure, " / " and "," can be interpreted as "and / or." For example, "A / B" and "A, B" can be interpreted as "A and / or B." Also, "A / B / C" and "A, B, C" can mean "at least one of A, B, and / or C."

[0038] In this disclosure, "or" can be interpreted as "and / or." For example, "A or B" can mean 1) only "A," 2) only "B," or 3) "A and B." Alternatively, in this disclosure, "or" can mean "additionally or alternatively."

[0039] Video Coding System Overview

[0040] FIG. 1 is a diagram illustrating a video coding system to which embodiments of the present disclosure can be applied.

[0041] A video coding system according to one embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may transmit encoded video and / or image information or data to the decoding device 20 in a file or streaming format via a digital storage medium or a network.

[0042] An encoding device 10 according to an embodiment may include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. A decoding device 20 according to an embodiment may include a reception unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmission unit 13 may be included in the encoding unit 12. The reception unit 21 may be included in the decoding unit 22. The rendering unit 23 may include a display unit, which may be configured as a separate device or an external component.

[0043] The video source generation unit 11 can acquire video / images through a video / image capture, synthesis, or generation process. The video source generation unit 11 can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, virtual video / images can be generated via a computer, etc., in which case the video / image capture process can be replaced with a process in which related data is generated.

[0044] The encoder 12 may encode the input video / image. The encoder 12 may perform a series of steps such as prediction, transformation, and quantization for compression and coding efficiency. The encoder 12 may output the encoded data (encoded video / image information) in a bitstream format.

[0045] The transmitter 13 may transmit the encoded video / image information or data output in a bitstream format to the receiver 21 of the decoding device 20 in a file or streaming format via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter 13 may include elements for generating a media file in a predetermined file format and elements for transmitting via a broadcasting / communication network. The receiver 21 may extract / receive the bitstream from the storage medium or network and transmit it to the decoder 22.

[0046] The decoding unit 22 can decode the video / image by performing a series of steps such as inverse quantization, inverse transformation, and prediction corresponding to the operations of the encoding unit 12.

[0047] The rendering unit 23 can render the decoded video / images, and the rendered video / images can be displayed via the display unit.

[0048] Overview of the image encoding device

[0049] FIG. 2 is a diagram schematically illustrating an image encoding device to which an embodiment of the present disclosure can be applied.

[0050] 2, the image encoding device 100 may include an image division unit 110, a subtraction unit 115, a transform unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse transform unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter prediction unit 180, an intra prediction unit 185, and an entropy encoding unit 190. The inter prediction unit 180 and the intra prediction unit 185 may be collectively referred to as a "prediction unit." The transform unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse transform unit 150 may be included in a residual processing unit. The residual processing unit may further include a subtraction unit 115.

[0051] Depending on the embodiment, all or at least some of the components constituting the image encoding device 100 may be realized by a single hardware component (e.g., an encoder or a processor). Also, the memory 170 may include a decoded picture buffer (DPB) and may be realized by a digital storage medium.

[0052] The image division unit 110 may divide an input image (or picture, frame) input to the image encoding device 100 into one or more processing units. As an example, the processing units may be called coding units (CUs). The coding units may be obtained by recursively dividing a coding tree unit (CTU) or a largest coding unit (LCU) using a QT / BT / TT (quad-tree / binary-tree / ternary-tree) structure. For example, one coding unit may be divided into multiple coding units at deeper depths based on a quad-tree structure, a binary-tree structure, and / or a ternary-tree structure. To divide the coding units, the quad-tree structure may be applied first, and then the binary-tree structure and / or the ternary-tree structure may be applied later. The coding procedure according to the present disclosure may be performed based on the final coding unit that is not further divided. The maximum coding unit may be used as the final coding unit, or a lower-depth coding unit obtained by dividing the maximum coding unit may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and / or reconstruction, which will be described later. As another example, a processing unit of the coding procedure may be a prediction unit (PU) or a transform unit (TU). The prediction unit and the transform unit may be divided or partitioned from the final coding unit, respectively. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.

[0053] The prediction unit (inter prediction unit 180 or intra prediction unit 185) may perform prediction on a current block (current block) to generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block or CU. The prediction unit may generate various information related to prediction of the current block and transmit it to the entropy coding unit 190. The prediction information may be coded by the entropy coding unit 190 and output in a bitstream format.

[0054] The intra prediction unit 185 may predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located far away from the current block according to the intra prediction mode and / or intra prediction technique. The intra prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the granularity of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 185 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.

[0055] The inter prediction unit 180 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation between the motion information of neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a collocated CU (colCU), etc. The reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter predictor 180 may construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter predictor 180 may use motion information of neighboring blocks as motion information for the current block. In the case of skip mode, unlike in merge mode, a residual signal may not be transmitted.In the case of a motion vector prediction (MVP) mode, the motion vector of a neighboring block is used as a motion vector predictor, and the motion vector of the current block can be signaled by encoding a motion vector difference and an indicator for the motion vector predictor. The motion vector difference may mean the difference between the motion vector of the current block and the motion vector predictor.

[0056] The predictor may generate a prediction signal based on various prediction methods and / or prediction techniques, which will be described later. For example, the predictor may apply intra prediction or inter prediction to predict the current block, or may simultaneously apply intra prediction and inter prediction. A prediction method that simultaneously applies intra prediction and inter prediction to predict the current block may be referred to as combined inter and intra prediction (CIIP). The predictor may also perform intra block copy (IBC) to predict the current block. Intra block copy can be used for content image / video coding, such as screen content coding (SCC), for games. IBC is a method of predicting a current block using an already reconstructed reference block in a current picture that is located a predetermined distance away from the current block. When IBC is applied, the position of the reference block in the current picture may be coded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that a reference block is derived within the current picture. That is, the IBC may use at least one of the inter prediction techniques described in this disclosure.

[0057] The prediction signal generated by the prediction unit may be used to generate a restored signal or a residual signal. The subtraction unit 115 may subtract the prediction signal (predicted block, predicted sample array) output from the prediction unit from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array). The generated residual signal may be transmitted to the conversion unit 120.

[0058] The transform unit 120 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loeve transform (KLT), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, the GBT refers to a transform obtained from a graph representing inter-pixel relationship information. The CNT refers to a transform obtained based on a predicted signal generated using all previously reconstructed pixels. The transform process may be applied to pixel blocks having the same square size or to non-square blocks of variable size.

[0059] The quantization unit 130 may quantize the transform coefficients and transmit the quantized transform coefficients to the entropy coding unit 190. The entropy coding unit 190 may encode the quantized signal (information about the quantized transform coefficients) and output the encoded signal in a bitstream format. The information about the quantized transform coefficients may be referred to as residual information. The quantization unit 130 may rearrange the quantized transform coefficients in a block format into a one-dimensional vector format based on a coefficient scan order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector format.

[0060] The entropy coding unit 190 may perform various coding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy coding unit 190 may also code information necessary for video / image restoration (e.g., values ​​of syntax elements) together with or separately from the quantized transform coefficients. The coded information (e.g., coded video / image information) may be transmitted or stored in a bitstream format in network abstraction layer (NAL) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The signaling information, transmitted information and / or syntax elements mentioned in this disclosure may be encoded through the above-described encoding procedure and included in the bitstream.

[0061] The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitting unit (not shown) that transmits and / or a storing unit (not shown) that stores the signal output from the entropy encoding unit 190 may be provided as an internal / external element of the image encoding device 100, or the transmitting unit may be provided as a component of the entropy encoding unit 190.

[0062] The quantized transform coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via the inverse quantization unit 140 and the inverse transform unit 150.

[0063] The adder 155 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit 180 or the intra prediction unit 185. When there is no residual for the current block to be processed, such as when a skip mode is applied, the predicted block may be used as the reconstructed block. The adder 155 may be referred to as a reconstruction unit or a reconstructed block generation unit. The generated reconstructed signal may be used for intra prediction of the next current block to be processed in the current picture, and may also be used for inter prediction of the next picture after filtering, as will be described later.

[0064] The filtering unit 160 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 160 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 170, specifically, in the DPB of the memory 170. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filtering unit 160 may generate various information related to filtering and transmit it to the entropy coding unit 190, as will be described later in connection with each filtering method. The filtering information may be coded by the entropy coding unit 190 and output in a bitstream format.

[0065] The modified reconstructed picture transmitted to the memory 170 can be used as a reference picture in the inter prediction unit 180. When inter prediction is applied through this, the image encoding device 100 can avoid a prediction mismatch between the image encoding device 100 and the image decoding device, and can also improve encoding efficiency.

[0066] The DPB in the memory 170 may store modified reconstructed pictures for use as reference pictures in the inter predictor 180. The memory 170 may store motion information of blocks from which motion information in the current picture is derived (or coded) and / or motion information of already reconstructed intra-picture blocks. The stored motion information may be transmitted to the inter predictor 180 to be used as motion information of spatially surrounding blocks or temporally surrounding blocks. The memory 170 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 185.

[0067] Overview of the image decoding device

[0068] FIG. 3 is a diagram schematically illustrating an image decoding device to which an embodiment of the present disclosure can be applied.

[0069] 3, the image decoding apparatus 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265. The inter prediction unit 260 and the intra prediction unit 265 may be collectively referred to as a "prediction unit." The inverse quantization unit 220 and the inverse transform unit 230 may be included in a residual processing unit.

[0070] Depending on the embodiment, all or at least some of the components constituting the image decoding device 200 may be realized by a single hardware component (e.g., a decoder or a processor). Also, the memory 170 may include a DPB and may be realized by a digital storage medium.

[0071] The image decoding device 200, which receives a bitstream including video / image information, can reconstruct an image by performing a process corresponding to the process performed by the image encoding device 100 of FIG. 2. For example, the image decoding device 200 can perform decoding using a processing unit applied in the image encoding device. Therefore, the decoding processing unit can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. The reconstructed image signal decoded and output by the image decoding device 200 can be reproduced by a reproduction device (not shown).

[0072] The image decoding apparatus 200 may receive a signal output from the image encoding apparatus of FIG. 2 in a bitstream format. The received signal may be decoded via an entropy decoding unit 210. For example, the entropy decoding unit 210 may parse the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / image information may also include general constraint information. The image decoding apparatus may further use the information on the parameter sets and / or the general constraint information to decode an image. The signaling information, received information, and / or syntax elements referred to in the present disclosure may be obtained from the bitstream by being decoded via the decoding procedure. For example, the entropy decoding unit 210 may decode information in a bitstream based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values ​​of syntax elements required for image restoration and quantized values ​​of transform coefficients related to residuals. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element from the bitstream, determines a context model using information on the syntax element to be decoded and decoded information on neighboring blocks and the block to be decoded, or information on symbols / bins decoded in a previous step, predicts the occurrence probability of the bins based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values ​​of each syntax element. After determining the context model, the CABAC entropy decoding method may update the context model using information on the decoded symbol / bin for the context model of the next symbol / bin.Among the information decoded by the entropy decoding unit 210, information related to prediction is provided to the prediction units (inter prediction unit 260 and intra prediction unit 265), and residual values ​​entropy decoded by the entropy decoding unit 210, i.e., quantized transform coefficients and related parameter information, may be input to the inverse quantization unit 220. Also, among the information decoded by the entropy decoding unit 210, information related to filtering may be provided to the filtering unit 240. Meanwhile, a receiving unit (not shown) for receiving a signal output from the image encoding device may be further provided as an internal / external element of the image decoding device 200, or the receiving unit may be provided as a component of the entropy decoding unit 210.

[0073] Meanwhile, the image decoding apparatus according to the present disclosure may be referred to as a video / image / picture decoding apparatus. The image decoding apparatus may include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoding unit 210, and the sample decoder may include at least one of an inverse quantization unit 220, an inverse transform unit 230, an adder 235, a filtering unit 240, a memory 250, an inter prediction unit 260, and an intra prediction unit 265.

[0074] The inverse quantization unit 220 may inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit 220 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the image encoding device. The inverse quantization unit 220 may perform inverse quantization on the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.

[0075] The inverse transform unit 230 can inversely transform the transform coefficients to obtain a residual signal (residual block, residual sample array).

[0076] The prediction unit may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The prediction unit may determine whether intra prediction or inter prediction is applied to the current block based on information about the prediction output from the entropy decoding unit 210, and may determine a specific intra / inter prediction mode (prediction technique).

[0077] The prediction unit can generate a prediction signal based on various prediction methods (techniques) described below, as described in the description of the prediction unit of the image encoding device 100.

[0078] The intra predictor 265 may predict the current block by referring to samples in the current picture. The description of the intra predictor 185 may also be applied to the intra predictor 265.

[0079] The inter prediction unit 260 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on correlations between motion information of neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter prediction unit 260 may construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction may be performed based on various prediction modes (techniques), and the prediction information may include information indicating the inter prediction mode (technique) for the current block.

[0080] The adder 235 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to a prediction signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 260 and / or intra prediction unit 265). When there is no residual for the current block, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The description of the adder 155 also applies to the adder 235. The adder 235 may also be referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next current block in the current picture, and may also be used for inter prediction of the next picture via filtering, as described below.

[0081] The filtering unit 240 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 240 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may store the modified reconstructed picture in the memory 250, specifically, in a DPB of the memory 250. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.

[0082] The (modified) reconstructed picture stored in the DPB of the memory 250 can be used as a reference picture in the inter predictor 260. The memory 250 can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information can be transmitted to the inter predictor 260 to be used as motion information of a spatially surrounding block or a temporally surrounding block. The memory 250 can store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 265.

[0083] In this specification, the embodiments described for the filtering unit 160, inter prediction unit 180 and intra prediction unit 185 of the image encoding device 100 can also be applied in a similar or corresponding manner to the filtering unit 240, inter prediction unit 260 and intra prediction unit 265 of the image decoding device 200, respectively.

[0084] Inter Prediction Overview

[0085] Inter prediction according to the present disclosure will be described below.

[0086] The prediction unit of the image encoding / decoding apparatus according to the present disclosure may perform inter prediction on a block-by-block basis to derive predicted samples. Inter prediction may refer to prediction derived in a manner dependent on data elements (e.g., sample values ​​or motion information) of pictures other than the current picture. When inter prediction is applied to the current block, a predicted block (prediction block or prediction sample array) for the current block may be derived based on a reference block (reference sample array) identified by a motion vector in a reference picture indicated by a reference picture index. In this case, to reduce the amount of motion information transmitted in the inter prediction mode, motion information of the current block may be predicted on a block, sub-block, or sample-by-block basis based on correlation between motion information of neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction type (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). When inter prediction is applied, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporal peripheral block may be the same or different. The temporal peripheral block may be called a collocated reference block, a collocated CU (colCU), a col block, etc., and the reference picture including the temporal peripheral block may be called a collocated picture (colPic), a col picture, etc. For example, a motion information candidate list may be constructed based on the peripheral blocks of the current block, and flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block may be signaled.

[0087] Inter prediction may be performed based on various prediction modes. For example, in skip mode and merge mode, the motion information of the current block may be the same as the motion information of a selected neighboring block. Unlike merge mode, in skip mode, a residual signal may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of a selected neighboring block may be used as a motion vector predictor, and a motion vector difference may be signaled. In this case, the motion vector of the current block may be derived using the sum of the motion vector predictor and the motion vector difference. In this disclosure, MVP mode may be used interchangeably with AMVP (Advanced Motion Vector Prediction).

[0088] The motion information may include L0 motion information and / or L1 motion information based on an inter-prediction type (such as L0 prediction, L1 prediction, or Bi prediction). A motion vector in the L0 direction may be referred to as an L0 motion vector or MVL0, and a motion vector in the L1 direction may be referred to as an L1 motion vector or MVL1. Prediction based on an L0 motion vector may be referred to as L0 prediction, prediction based on an L1 motion vector may be referred to as L1 prediction, and prediction based on both the L0 motion vector and the L1 motion vector may be referred to as bi-prediction (Bi) prediction. Here, an L0 motion vector may indicate a motion vector associated with a reference picture list L0 (L0), and an L1 motion vector may indicate a motion vector associated with a reference picture list L1 (L1). The reference picture list L0 may include, as reference pictures, pictures that are earlier in output order than the current picture, and the reference picture list L1 may include pictures that are later in output order than the current picture. The previous picture may be referred to as a forward (reference) picture, and the subsequent picture may be referred to as a backward (reference picture). The reference picture list L0 may further include, as reference pictures, pictures that are subsequent to the current picture in output order. In this case, the previous picture may be indexed first in the reference picture list L0, and the subsequent picture may be indexed next. The reference picture list L1 may further include, as reference pictures, pictures that are prior to the current picture in output order. In this case, the subsequent picture may be indexed first in the reference picture list L1, and the previous picture may be indexed next. Here, the output order may correspond to a picture order count (POC) order.

[0089] FIG. 4 is a flowchart illustrating an inter-prediction based video / image coding method.

[0090] FIG. 5 is a diagram illustrating an exemplary configuration of the inter prediction unit 180 according to the present disclosure.

[0091] The encoding method of FIG. 4 may be performed by the image encoding apparatus of FIG. 2. Specifically, step S410 may be performed by the inter prediction unit 180, and step S420 may be performed by the residual processing unit. Specifically, step S420 may be performed by the subtraction unit 115. Step S430 may be performed by the entropy encoding unit 190. The prediction information of step S430 may be derived by the inter prediction unit 180, and the residual information of step S430 may be derived by the residual processing unit. The residual information is information about the residual sample. The residual information may include information about quantized transform coefficients for the residual sample. As described above, the residual sample may be derived as transform coefficients via the transform unit 120 of the image encoding apparatus, and the transform coefficients may be derived as quantized transform coefficients via the quantization unit 130. Information about the quantized transform coefficients may be coded by the entropy encoding unit 190 through a residual coding procedure.

[0092] 4 and 5, an image encoding apparatus may perform inter prediction on a current block (S410). The image encoding apparatus may derive an inter prediction mode and motion information of the current block and generate a prediction sample for the current block. Here, the inter prediction mode determination, motion information derivation, and prediction sample generation procedures may be performed simultaneously, or one procedure may be performed before the other procedures. For example, as shown in FIG. 5, an inter prediction unit 180 of the image encoding apparatus may include a prediction mode determination unit 181, a motion information derivation unit 182, and a prediction sample derivation unit 183. The prediction mode determination unit 181 may determine a prediction mode for the current block, the motion information derivation unit 182 may derive motion information of the current block, and the prediction sample derivation unit 183 may derive a prediction sample for the current block. For example, the inter prediction unit 180 of the image encoding device may search for blocks similar to the current block within a certain region (search region) of a reference picture through motion estimation and derive a reference block whose difference from the current block is minimum or equal to or less than a certain criterion. Based on this, a reference picture index indicating a reference picture in which the reference block is located may be derived, and a motion vector may be derived based on the position difference between the reference block and the current block. The image encoding device may determine a mode to be applied to the current block from various prediction modes. The image encoding device may compare rate-distortion (RD) costs for the various prediction modes and determine an optimal prediction mode for the current block. However, the method by which the image encoding device determines a prediction mode for the current block is not limited to the above example, and various methods may be used.

[0093] For example, when a skip mode or a merge mode is applied to a current block, the image encoding apparatus may derive merge candidates from neighboring blocks of the current block and construct a merge candidate list using the derived merge candidates. Furthermore, the image encoding apparatus may derive a reference block whose difference from the current block is minimum or equal to or less than a certain criterion among reference blocks indicated by merge candidates included in the merge candidate list. In this case, a merge candidate associated with the derived reference block may be selected, and merge index information indicating the selected merge candidate may be generated and signaled to the image decoding apparatus. Motion information of the current block may be derived using motion information of the selected merge candidate.

[0094] As another example, when the MVP mode is applied to the current block, the image encoding apparatus may derive motion vector predictor (MVP) candidates from neighboring blocks of the current block and construct an MVP candidate list using the induced MVP candidates. The image encoding apparatus may also use the motion vector of an MVP candidate selected from the MVP candidates included in the MVP candidate list as the MVP of the current block. In this case, for example, a motion vector pointing to a reference block derived by the motion estimation described above may be used as the motion vector of the current block, and the MVP candidate having the smallest difference from the motion vector of the current block may be the selected MVP candidate. A motion vector difference (MVD), which is the difference obtained by subtracting the MVP from the motion vector of the current block, may be derived. In this case, index information indicating the selected MVP candidate and information regarding the MVD may be signaled to the image decoding apparatus. Also, when the MVP mode is applied, the value of the reference picture index may be configured as reference picture index information and separately signaled to the image decoding apparatus.

[0095] The image encoding apparatus may derive residual samples based on the predicted samples (S420). The image encoding apparatus may derive the residual samples by comparing the original samples of the current block with the predicted samples. For example, the residual samples may be derived by subtracting corresponding predicted samples from the original samples.

[0096] The image encoding apparatus may encode image information including prediction information and residual information (S430). The image encoding apparatus may output the encoded image information in a bitstream format. The prediction information may be information related to the prediction procedure and may include prediction mode information (e.g., a skip flag, a merge flag, or a mode index) and information about motion information. Among the prediction mode information, the skip flag is information indicating whether a skip mode is applied to a current block, and the merge flag is information indicating whether a merge mode is applied to the current block. Alternatively, the prediction mode information may be information indicating one of a plurality of prediction modes, such as a mode index. If the skip flag and the merge flag are both 0, it may be determined that the MVP mode is applied to the current block. The information about the motion information may include candidate selection information (e.g., a merge index, an MVP flag, or an MVP index) that is information for deriving a motion vector. Among the candidate selection information, the merge index may be signaled when a merge mode is applied to the current block, and may be information for selecting one of merge candidates included in a merge candidate list. Among the candidate selection information, the MVP flag or MVP index may be signaled when the MVP mode is applied to the current block, and may be information for selecting one of the MVP candidates included in the MVP candidate list. Furthermore, the information about the motion information may include the above-mentioned information about MVD and / or reference picture index information. Furthermore, the information about the motion information may include information indicating whether L0 prediction, L1 prediction, or Bi-prediction is applied. The residual information is information about the residual sample. The residual information may include information about quantized transform coefficients for the residual sample.

[0097] The output bitstream can be stored in a (digital) storage medium and transmitted to the image decoding device, or can be transmitted to the image decoding device via a network.

[0098] Meanwhile, as described above, the image coding apparatus can generate a reconstructed picture (a picture including reconstructed samples and reconstructed blocks) based on the reference samples and the residual samples. This is because the image coding apparatus derives the same prediction result as that performed in the image decoding apparatus, thereby improving coding efficiency. Therefore, the image coding apparatus can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in a memory and use it as a picture for inter prediction. As described above, an in-loop filtering procedure can be further applied to the reconstructed picture.

[0099] FIG. 6 is a flowchart illustrating an inter-prediction-based video / image decoding method, and FIG. 7 is a diagram illustrating an exemplary configuration of the inter prediction unit 260 according to the present disclosure.

[0100] The image decoding apparatus may perform operations corresponding to those performed by the image encoding apparatus, such as performing prediction on a current block based on received prediction information and deriving predicted samples.

[0101] The decoding method of FIG. 6 may be performed by the image decoding apparatus of FIG. 3. Steps S610 to S630 may be performed by the inter prediction unit 260, and the prediction information of step S610 and the residual information of step S640 may be obtained from a bitstream by the entropy decoding unit 210. The residual processing unit of the image decoding apparatus may derive residual samples for the current block based on the residual information (S640). Specifically, the inverse quantization unit 220 of the residual processing unit may derive transform coefficients by performing inverse quantization on the quantized transform coefficients derived based on the residual information, and the inverse transform unit 230 of the residual processing unit may derive residual samples for the current block by performing inverse transform on the transform coefficients. Step S650 may be performed by the adder 235 or a reconstruction unit.

[0102] 6 and 7, the image decoding apparatus may determine a prediction mode for the current block based on received prediction information (S610). The image decoding apparatus may determine which inter prediction mode is applied to the current block based on prediction mode information in the prediction information.

[0103] For example, it may determine whether the skip mode is applied to the current block based on the skip flag. Also, it may determine whether the merge mode or the MVP mode is applied to the current block based on the merge flag. Or, it may select one of various inter prediction mode candidates based on the mode index. The inter prediction mode candidates may include skip mode, merge mode, and / or MVP mode, or various inter prediction modes described below.

[0104] The image decoding apparatus may derive motion information of the current block based on the determined inter prediction mode (S620). For example, when a skip mode or a merge mode is applied to the current block, the image decoding apparatus may construct a merge candidate list (described below) and select one of the merge candidates included in the merge candidate list. The selection may be made based on the candidate selection information (merge index) described above. The image decoding apparatus may derive motion information of the current block using motion information of the selected merge candidate. For example, the motion information of the selected merge candidate may be used as motion information of the current block.

[0105] As another example, when the MVP mode is applied to the current block, the image decoding apparatus may construct an MVP candidate list and use a motion vector of an MVP candidate selected from the MVP candidates included in the MVP candidate list as the MVP of the current block. The selection may be made based on the candidate selection information (MVP flag or MVP index). In this case, the MVD of the current block may be derived based on information related to the MVD, and the motion vector of the current block may be derived based on the MVP of the current block and the MVD. Also, the image decoding apparatus may derive a reference picture index of the current block based on the reference picture index information. A picture pointed to by the reference picture index in the associated reference picture list for the current block may be derived as a reference picture referenced for inter-prediction of the current block.

[0106] The image decoding apparatus may generate prediction samples for the current block based on the motion information of the current block (S630). In this case, the reference picture may be derived based on a reference picture index of the current block, and the prediction samples of the current block may be derived using samples of a reference block pointed to in the reference picture by the motion vector of the current block. Depending on the circumstances, a prediction sample filtering procedure may further be performed on all or some of the prediction samples of the current block.

[0107] 7, the inter prediction unit 260 of the image decoding apparatus may include a prediction mode determination unit 261, a motion information derivation unit 262, and a prediction sample derivation unit 263. The inter prediction unit 260 of the image decoding apparatus may determine a prediction mode for the current block based on prediction mode information received from the prediction mode determination unit 261, derive motion information (such as a motion vector and / or a reference picture index) of the current block based on information related to the motion information received from the motion information derivation unit 262, and derive a prediction sample of the current block via the prediction sample derivation unit 263.

[0108] The image decoding apparatus may generate residual samples for the current block based on the received residual information (S640). The image decoding apparatus may generate reconstructed samples for the current block based on the predicted samples and the residual samples, and generate a reconstructed picture based on the reconstructed samples (S650). Thereafter, an in-loop filtering procedure may be further applied to the reconstructed picture, as described above.

[0109] As described above, the inter prediction procedure may include an inter prediction mode determination step, a motion information deriving step according to the determined prediction mode, and a prediction execution step (generation of prediction samples) based on the derived motion information. The inter prediction procedure may be performed in an image encoding device and an image decoding device, as described above.

[0110] Subpicture Overview

[0111] Subpictures according to the present disclosure will now be described.

[0112] A picture can be divided into tiles, and each tile can be further divided into sub-pictures, each of which can contain one or more slices and form a rectangular area within the picture.

[0113] FIG. 8 is a diagram showing an example of a sub-picture.

[0114] Referring to Figure 8, one picture may be divided into 18 tiles. 12 tiles may be arranged on the left side of the picture, and each of the tiles may include one sub-picture / slice consisting of 16 CTUs. 6 tiles may be arranged on the right side of the picture, and each of the tiles may include two sub-pictures / slices consisting of four CTUs. As a result, the picture is divided into 24 sub-pictures, and each of the sub-pictures may include one slice.

[0115] Information about the sub-pictures (eg, the number and size of the sub-pictures) can be coded / signaled via higher level syntax, such as the SPS, PPS and / or slice header.

[0116] FIG. 9 is a diagram showing an example of an SPS including information about sub-pictures.

[0117] Referring to FIG. 9, an SPS may include a syntax element subpic_info_present_flag indicating whether subpicture information exists for a coded layer video sequence (CLVS). For example, a subpic_info_present_flag having a first value (e.g., 0) may indicate that subpicture information does not exist for a CLVS and that only one subpicture exists in each picture of the CLVS. Conversely, a subpic_info_present_flag having a second value (e.g., 1) may indicate that subpicture information exists for a CLVS and that one or more subpictures exist in each picture of the CLVS. In one example, if picture spatial resolution can be changed in a CLVS that references an SPS (e.g., res_change_in_clvs_allowed_flag==1), the value of subpic_info_present_flag may be limited to a first value (e.g., 0). On the other hand, if the bitstream contains only a subset of the sub-pictures of the input bitstream to the sub-bitstream extraction process as a result of the sub-bitstream extraction process, the value of subpic_info_present_flag can be limited to a second value (e.g., 1).

[0118] The SPS may also include a syntax element sps_num_subpics_minus1 that indicates the number of subpictures. For example, a value of sps_num_subpics_inus1 plus 1 may indicate the number of subpictures included in each picture in the CLVS. In one example, the value of sps_num_subpics_minus1 may be limited to a range of 0 to Ceil(pic_width_max_in_luma_samples x CtbSizeY) x Ceil(pic_height_max_in_luma_samples / CtbSizeY), where Ceil(x) may be a ceiling function that outputs the smallest integer value equal to or greater than x. Furthermore, pic_width_max_in_luma_samples may mean the maximum width in luma sample units of each picture, pic_height_max_in_luma_samples may mean the maximum height in luma sample units of each picture, and CtbSizeY may mean the array size of each luma component CTB (coding tree block) in both width and height. Meanwhile, if sps_num_subpics_minus1 is not present, the value of sps_num_subpics_minus1 may be inferred to be a first value (e.g., 0).

[0119] The SPS may also include a syntax element sps_independent_subpics_flag that indicates whether subpicture boundaries are treated as picture boundaries. For example, sps_independent_subpics_flag having a second value (e.g., 1) may indicate that all subpicture boundaries in the CLVS are treated as picture boundaries and that loop filtering across the subpicture boundaries is not performed. In contrast, sps_independent_subpics_flag having a first value (e.g., 0) may indicate that the above-mentioned constraints do not apply. On the other hand, if sps_independent_subpics_flag is not present, the value of sps_independent_subpics_flag may be inferred to be the first value (e.g., 0).

[0120] The SPS may also include syntax elements subpic_ctu_top_left_x[i], subpic_ctu_top_left_y[i], subpic_width_minus1[i], and subpic_height_minus1[i] that indicate the position and size of the subpicture.

[0121] subpic_ctu_top_left_x[i] may represent the horizontal position of the top left CTU of the ith subpicture in CtbSizeY units. In one example, the length of subpic_ctu_top_left_x[i] may be Ceil(Log2((pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)) bits. On the other hand, if subpic_ctu_top_left_x[i] does not exist, the value of subpic_ctu_top_left_x[i] may be inferred to be a first value (e.g., 0).

[0122] subpic_ctu_top_left_y[i] may represent the vertical position of the top-left CTU of the ith subpicture in CtbSizeY units. In one example, the length of subpic_ctu_top_left_y[i] may be Ceil(Log2((pic_height_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)) bits. On the other hand, if subpic_ctu_top_left_y[i] does not exist, the value of subpic_ctu_top_left_y[i] may be inferred to be a first value (e.g., 0).

[0123] The value of subpic_width_minus1[i] plus 1 may represent the width of the i-th subpicture in CtbSizeY units. In one example, the length of subpic_width_minus1[i] may be Ceil(Log2((pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)) bits. On the other hand, if subpic_width_minus1[i] is not present, the value of subpic_width_minus1[i] may be inferred as ((pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)-subpic_ctu_top_left_x[i]-1.

[0124] The value of subpic_height_minus1[i] plus 1 may represent the height of the i-th subpicture in CtbSizeY units. In one example, the length of subpic_height_minus1[i] may be Ceil(Log2((pic_height_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)) bits. On the other hand, if subpic_height_minus1[i] is not present, the value of subpic_height_minus1[i] may be inferred as ((pic_height_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)-subpic_ctu_top_left_y[i]-1.

[0125] The SPS may also include subpic_treated_as_pic_flag[i], which indicates whether a subpicture is treated as one picture. For example, subpic_treated_as_pic_flag[i] having a first value (e.g., 0) may indicate that the i-th subpicture in each coded picture in the CLVS is not treated as one picture in the decoding process excluding in-loop filtering. In contrast, subpic_treated_as_pic_flag[i] having a second value (e.g., 1) may indicate that the i-th subpicture in each coded picture in the CLVS is treated as one picture in the decoding process excluding in-loop filtering. If subpic_treated_as_pic_flag[i] is not present, the value of subpic_treated_as_pic_flag[i] may be inferred to be the same as the above-mentioned sps_independent_subpics_flag. In one example, subpic_treated_as_pic_flag[i] can be coded / signaled only if the above-mentioned sps_independent_subpics_flag has the first value (e.g., 0) (i.e., the subpicture boundary is not treated as a picture boundary).

[0126] On the other hand, if subpic_treated_as_pic_flag[i] has a second value (e.g., 1), it may be a requirement for bitstream conformance that all of the following conditions be true for each output layer and its reference layer in an output layer set (OLS) that includes the layer containing the i-th subpicture as an output layer:

[0127] - (Condition 1) All pictures in the output layer and its reference layers must have the same value of pic_width_in_luma_samples and the same value of pic_height_in_luma_samples.

[0128] (Condition 2) All SPSs referenced by the output layer and its reference layers must have the same value of sps_num_subpics_minus1, and the same values ​​of subpic_ctu_top_left_x[j], subpic_ctu_top_left_y[j], subpic_width_minus1[j], subpic_height_minus1[j], and loop_filter_across_subpic_enabled_flag[j], where j is in the range from 0 to sps_num_subpics_minus1.

[0129] The SPS may also include a syntax element loop_filter_across_subpic_enabled_flag[i] indicating whether an in-loop filtering operation across a subpicture boundary is permitted. For example, a loop_filter_across_subpic_enabled_flag[i] having a first value (e.g., 0) may indicate that an in-loop filtering operation across the boundary of the i-th subpicture in each coded picture in the CLVS is not permitted. In contrast, a loop_filter_across_subpic_enabled_flag[i] having a second value (e.g., 1) may indicate that an in-loop filtering operation across the boundary of the i-th subpicture in each coded picture in the CLVS is permitted. If loop_filter_across_subpic_enabled_flag[i] is not present, the value of loop_filter_across_subpic_enabled_flag[i] may be inferred to be the same as 1-sps_independent_subpics_flag. In one example, loop_filter_across_subpic_enabled_flag[i] can be coded / signaled only when the above-mentioned sps_independent_subpics_flag has the first value (e.g., 0) (i.e., subpicture boundaries are not treated as picture boundaries). Meanwhile, as a requirement for bitstream consistency, the subpicture format must be such that when each subpicture is decoded, the entire left boundary and the entire top boundary of each subpicture are configured with the picture boundary or the boundary of a previously decoded subpicture.

[0130] FIG. 10 is a diagram showing a method for encoding an image using sub-pictures by an image encoding device according to an embodiment of the present disclosure.

[0131] The image coding device can code the current picture based on the sub-picture structure, or can code at least one sub-picture that constitutes the current picture and output a (sub)bitstream containing (coded) information about the (coded) at least one sub-picture.

[0132] 10, the image coding apparatus may divide an input picture into a plurality of sub-pictures (S1010). Then, the image coding apparatus may generate information about the sub-pictures (S1020). Here, the information about the sub-pictures may include, for example, information about the area of ​​the sub-pictures and / or information about the grid spacing to be used for the sub-pictures. The information about the sub-pictures may also include information about whether each sub-picture can be treated as one picture and / or information about whether in-loop filtering can be performed across the boundaries of each sub-picture (boundary).

[0133] The image coding apparatus may encode at least one subpicture based on information about the subpicture. For example, each subpicture may be independently encoded based on information about the subpicture. The image coding apparatus may then encode image information including information about the subpicture and output a bitstream (S1030). Here, a bitstream for a subpicture may be referred to as a substream or sub-bitstream.

[0134] FIG. 11 is a diagram showing a method in which an image decoding device according to an embodiment of the present disclosure decodes an image using sub-pictures.

[0135] The image decoding device can decode at least one sub-picture included in the current picture using (encoded) information about the at least one sub-picture obtained from the (sub)bitstream.

[0136] 11, the image decoding apparatus may obtain information about a sub-picture from a bitstream (S1110). Here, the bitstream may include a sub-stream or sub-bitstream for the sub-picture. The information about the sub-picture may be configured in a higher-level syntax of the bitstream. Then, the image decoding apparatus may derive at least one sub-picture based on the information about the sub-picture (S1120).

[0137] The image decoding apparatus may decode at least one sub-picture based on information about the sub-picture (S1130). For example, if a current sub-picture including a current block is treated as a single picture, the current sub-picture may be decoded independently. If in-loop filtering can be performed across the boundary of the current sub-picture, in-loop filtering (e.g., deblocking filtering) may be performed on the boundary of the current sub-picture and the boundary of an adjacent sub-picture adjacent to the boundary. If the boundary of the current sub-picture coincides with a picture boundary, in-loop filtering across the boundary of the current sub-picture may not be performed. The image decoding apparatus may decode the sub-picture based on a CABAC method, a prediction method, a residual processing method (transform, quantization), an in-loop filtering scheme, etc. Then, the image decoding apparatus may output the decoded at least one sub-picture or the current picture including the at least one sub-picture. The decoded sub-picture may be output in the form of an output sub-picture set (OPS). For example, in the context of a 360-degree or omnidirectional image, when only a portion of the current picture is rendered, only some of all sub-pictures in the current picture may be decoded, and all or some of the decoded sub-pictures may be rendered according to the user's viewport.

[0138] Wraparound Overview

[0139] When inter prediction is applied to a current block, a prediction block of the current block may be derived based on a reference block identified by a motion vector of the current block. In this case, if at least one reference sample in the reference block is outside the boundary of the reference picture, the sample value of the reference sample may be replaced with the sample value of an adjacent sample existing on the boundary or outermost edge of the reference picture. This is called padding, and the boundary of the reference picture may be extended through the padding.

[0140] On the other hand, when a reference picture is obtained from a 360-degree image, there may be continuity between the left and right boundaries of the reference picture. Accordingly, samples adjacent to the left boundary (or right boundary) of the reference picture may have the same / similar sample values ​​and / or motion information as samples adjacent to the right boundary (or left boundary) of the picture. Based on this characteristic, at least one reference sample in a reference block that falls outside the boundary of a reference picture may be replaced with an adjacent sample in the reference picture that corresponds to the reference sample. This is called (horizontal) wrap-around motion compensation, and the motion vector of the current block may be adjusted to point to the interior of the reference picture through the wrap-around motion compensation.

[0141] Wraparound motion compensation refers to a coding tool designed to improve the visual quality of restored images / videos, such as 360-degree images / videos projected in ERP format. According to existing motion compensation processes, when a motion vector of a current block points to a sample outside the boundary of a reference picture, the sample value of the outside sample can be derived by duplicating the sample value of an adjacent sample closest to the boundary through repeated padding. However, because 360-degree images / videos are acquired spherically and inherently have no image boundaries, a reference sample outside the boundary of a reference picture in the projected domain (two-dimensional domain) can always be obtained from an adjacent sample adjacent to the reference sample in the old domain (three-dimensional domain). Therefore, repeated padding is incompatible with 360-degree images / videos and may induce visual artifacts called seam artifacts in the restored viewport images / videos.

[0142] On the other hand, when a general projection method is applied, it may be difficult to obtain neighboring samples for wrap-around motion compensation in the old domain because 2D-to-3D and 3D-to-2D coordinate transformations are performed together with sample interpolation for fractional sample positions. However, when an ERP projection method is applied, spherical neighboring samples that are off the left boundary (or right boundary) of a reference picture can be relatively easily obtained from samples within the right boundary (or left boundary) of the reference picture. Therefore, considering the relative ease of implementation and widespread use of the ERP projection method, wrap-around motion compensation may be more effective for 360-degree images / videos coded in the ERP format.

[0143] FIG. 12 is a diagram showing an example of a 360-degree image converted into a two-dimensional picture.

[0144] 12, a 360-degree image 1210 can be converted into a 2D picture 1230 through a projection process. The 2D picture 1230 can have various projection formats, such as an equi-rectangular projection (ERP) format or a padded ERP (PERP) format, depending on the projection method applied to the 360-degree image 1210.

[0145] The 360-degree image 1210 does not have an image boundary due to the image characteristics acquired from all directions. However, the 2D picture 1230 acquired from the 360-degree image 1210 has an image boundary due to the projection process. In this case, the left boundary LBd and the right boundary RBd of the 2D picture 1230 may abut each other to form a line RL in the 360-degree image 1210. Therefore, the similarity between samples adjacent to the left boundary LBd and the right boundary RBd in the 2D picture 1230 may be relatively high.

[0146] Meanwhile, a predetermined region in the 360-degree image 1210 may correspond to an internal region or an external region of the 2D picture 1230 depending on the reference image boundary. For example, when the left boundary LBd of the 2D picture 1230 is used as a reference, region A in the 360-degree image 1210 may correspond to region A1 existing outside the 2D picture 1230. Conversely, when the right boundary RBd of the 2D picture 1230 is used as a reference, region A in the 360-degree image 1210 may correspond to region A2 existing inside the 2D picture 1230. Region A1 and region A2 may have the same / similar sample attributes to each other in that they correspond to the same region A using the 360-degree image 1210 as a reference.

[0147] Based on this characteristic, an external sample outside the left boundary LBd of the 2D picture 1230 can be replaced with an internal sample of the 2D picture 1230 located at a predetermined distance in the first direction DIR1 through wrap-around motion compensation. For example, an external sample of the 2D picture 1230 included in the A1 region can be replaced with an internal sample of the 2D picture 1230 included in the A2 region. Similarly, an external sample outside the right boundary RBd of the 2D picture 1230 can be replaced with an internal sample of the 2D picture 1230 located at a predetermined distance in the second direction DIR2 through wrap-around motion compensation.

[0148] FIG. 13 is a diagram illustrating an example of a wrap-around motion compensation process.

[0149] Referring to FIG. 13, when inter prediction is applied to a current block 1310 , a prediction block of the current block 1310 can be derived based on a reference block 1330 .

[0150] The reference block 1330 may be identified by a motion vector 1320 of the current block 1310. In one example, the motion vector 1320 may point to the upper left position of the reference block 1330 relative to the upper left position of a co-located block 1315 that is co-located with the current block 1310 in the reference picture.

[0151] The reference block 1330 may include a first region 1335 that is off the left boundary of the reference picture, as shown in Figure 13. Because the first region 1335 is unavailable for inter-prediction of the current block 1310, it may be replaced by a second region 1340 in the reference picture through wraparound motion compensation. The second region 1340 may correspond to the same region as the first region 1335 in the old domain (the three-dimensional domain), and the position of the second region 1340 may be identified by adding a wraparound offset to a predetermined position (e.g., the upper left position) of the first region 1335.

[0152] The wraparound offset can be set to the ERP width of the current picture before padding. Here, the ERP width may refer to the width of the original picture (i.e., the ERP picture) in the ERP format obtained from a 360-degree image. A horizontal padding process can be performed on the left and right boundaries of the ERP picture. Thus, the width of the current picture (PicWidth) can be determined as the sum of the ERP width, the amount of padding on the left boundary of the ERP picture (left padding), and the amount of padding on the right boundary of the ERP picture (right padding). Meanwhile, the wraparound offset can be coded / signaled using a predetermined syntax element (e.g., pps_ref_wraparound_offset) in the higher-level syntax. This syntax element is not affected by the amount of padding on the left and right boundaries of the ERP picture, thereby supporting asymmetric padding for the original picture. That is, the amount of padding on the left boundary of an ERP picture (left padding) and the amount of padding on the right boundary (right padding) may be different from each other.

[0153] Information regarding the wrap-around motion compensation (eg, whether or not activation is present, wrap-around offset, etc.) can be coded / signaled via higher level syntax, such as SPS and / or PPS.

[0154] FIG. 14a shows an example of an SPS including information about wrap-around motion compensation.

[0155] 14a, the SPS may include a syntax element sps_ref_wraparound_enabled_flag indicating whether wraparound motion compensation is applied at the sequence level. For example, sps_ref_wraparound_enabled_flag having a first value (e.g., 0) may indicate that wraparound motion compensation is not applied to the current video sequence including the current block. In contrast, sps_ref_wraparound_enabled_flag having a second value (e.g., 1) may indicate that wraparound motion compensation is applied to the current video sequence including the current block. In one example, wraparound motion compensation for the current video sequence may be applied only if the picture width (e.g., pic_width_in_luma_samples) and CTB width (CtbSizeY) satisfy the following conditions:

[0156] -(Condition):(CtbSizeY / MinCbSizeY+1)≧(pic_width_in_luma_samples / MinCbSizeY-1)

[0157] If the above condition is not satisfied, for example, if the value of (CtbSizeY / MinCbSizeY+1) is greater than the value of (pic_width_in_luma_samples / MinCbSizeY-1), sps_ref_wraparound_enabled_flag may be limited to a first value (e.g., 0). Here, CtbSizeY may represent the width or height of the luma component CTB, and MinCbSizeY may represent the minimum width or height of the luma component CB (coding block). Also, pic_width_max_in_luma_samples may represent the maximum width in luma sample units of each picture.

[0158] FIG. 14b shows an example of a PPS including information about wrap-around motion compensation.

[0159] Referring to FIG. 14b, the PPS may include a syntax element pps_ref_wraparound_enabled_flag that indicates whether wraparound motion compensation is applied at the picture level.

[0160] Referring to Figure 14b, the PPS may include a syntax element pps_ref_wraparound_enabled_flag indicating whether wraparound motion compensation is applied at the sequence level. For example, a pps_ref_wraparound_enabled_flag having a first value (e.g., 0) may indicate that wraparound motion compensation is not applied to the current picture including the current block. In contrast, a pps_ref_wraparound_enabled_flag having a second value (e.g., 1) may indicate that wraparound motion compensation is applied to the current picture including the current block. In one example, wraparound motion compensation for the current picture may be applied only when the picture width (e.g., pic_width_in_luma_samples) is greater than the CTB width (CtbSizeY). For example, if the value of (CtbSizeY / MinCbSizeY+1) is greater than the value of (pic_width_in_luma_samples / MinCbSizeY-1), pps_ref_wraparound_enabled_flag may be limited to a first value (e.g., 0). In another example, if sps_ref_wraparoud_enabled_flag has a first value (e.g., 0), the value of pps_ref_wraparound_enabled_flag may be limited to a first value (e.g., 0).

[0161] Furthermore, the PPS may include a syntax element pps_ref_wraparound_offset that indicates an offset for wraparound motion compensation. For example, a value obtained by adding ((CtbSizeY / MinCbSizeY)+2) to pps_ref_wraparound_offset may indicate a wraparound offset for calculating the wraparound position in units of luma samples. The value of pps_ref_wraparound_offset may be greater than or equal to 0 and less than or equal to ((pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)-2). Meanwhile, the variable PpsRefWraparoundOffset may be set to the same value as (pps_ref_wraparound_offset+(CtbSizeY / MinCbSizeY)+2). The variable PpsRefWraparoundOffset may be used in a process of clipping reference samples that deviate from the boundary of the reference picture.

[0162] On the other hand, if the current picture is divided into multiple sub-pictures, wrap-around motion compensation can be selectively performed based on the attributes of each sub-picture.

[0163] FIG. 15 is a flowchart showing a method for the image coding apparatus to perform wraparound motion compensation.

[0164] Referring to FIG. 15, the image coding apparatus may determine whether the current sub-picture is to be independently coded (S1510).

[0165] If the current sub-picture is independently coded ("YES" in S1510), the image coding apparatus may determine not to perform wraparound motion compensation on the current block (S1530). In this case, the image coding apparatus may clip the reference sample positions of the current block based on the sub-picture boundary and perform motion compensation using the reference samples at the clipped positions. This operation may be performed using, for example, a luma sample bilinear interpolation process, a luma sample interpolation filtering process, a luma integer sample fetching process, or a chroma sample interpolation process.

[0166] If the current subpicture is not independently coded ("NO" in S1510), the image coding apparatus may determine whether wraparound motion compensation is available (S1520). The image coding apparatus may determine whether wraparound motion compensation is available at the sequence level. For example, if all subpictures in the current video sequence have discontinuous subpicture boundaries, wraparound motion compensation may be restricted from being available for the current video sequence. If wraparound motion compensation is not available at the sequence level, wraparound motion compensation may be restricted from being available at the picture level. Alternatively, if wraparound motion compensation is available at the sequence level, the image coding apparatus may determine whether wraparound motion compensation is available at the picture level.

[0167] If wraparound motion compensation is available ("YES" in S1520), the image coding apparatus may perform wraparound motion compensation on the current block (S1540). In this case, the image coding apparatus may perform wraparound motion compensation by shifting the reference sample position of the current block by the wraparound offset and then clipping the shifted position based on the boundary of the reference picture.

[0168] Alternatively, if wraparound motion compensation is not available ("NO" in S1520), the image coding apparatus may not perform wraparound motion compensation on the current block (S1550). In this case, the image coding apparatus may clip the reference sample position of the current block based on the boundary of the reference picture, and perform motion compensation using the reference sample at the clipped position.

[0169] Meanwhile, the image decoding apparatus may determine whether the current subpicture is independently coded based on subpicture-related information (e.g., subpic_treated_as_pic_flag) acquired from the bitstream. For example, if subpic_treated_as_pic_flag has a first value (e.g., 0), the current subpicture may not be independently coded. In contrast, if subpic_treated_as_pic_flag has a second value (e.g., 1), the current subpicture may be independently coded.

[0170] The image decoding apparatus may then determine whether wraparound motion compensation is available based on wraparound-related information (e.g., pps_ref_wraparound_enabled_flag) obtained from the bitstream, and perform wraparound motion compensation on the current block based on the determination. For example, if pps_ref_wraparound_enabled_flag has a first value (e.g., 0), the image decoding apparatus may determine that wraparound motion compensation is not available for the current picture and may not perform wraparound motion compensation on the current block. In contrast, if pps_ref_wraparound_enabled_flag has a second value (e.g., 1), the image decoding apparatus may determine that wraparound motion compensation is available for the current picture and may perform wraparound motion compensation on the current block.

[0171] 15, wraparound motion compensation can be performed only when the current sub-picture is not treated as a single picture. As a result, a problem occurs in that wraparound-related coding tools cannot be used together with various sub-picture-related coding tools that assume independent coding of sub-pictures. This can cause a decrease in encoding / decoding performance for pictures with inter-boundary continuity, such as ERP pictures or PERP pictures.

[0172] To solve this problem, according to an embodiment of the present disclosure, even when the current sub-picture is independently coded, wraparound motion compensation can be performed according to a predetermined condition. Hereinafter, an embodiment of the present disclosure will be described in detail.

[0173] Example 1

[0174] According to the first embodiment of the present disclosure, if a current video sequence contains one or more independently coded sub-pictures with widths different from the picture width, wraparound motion compensation may be restricted from being enabled for the current video sequence. In this case, flag information (e.g., sps_ref_wraparound_enabled_flag) indicating whether wraparound motion compensation is enabled for the current video sequence may be restricted to have a first value (e.g., 0).

[0175] FIG. 16 is a flowchart showing a method for determining whether or not wraparound motion compensation is available in an image coding apparatus according to an embodiment of the present disclosure.

[0176] Referring to FIG. 16, the image coding apparatus may determine whether there are one or more independently coded sub-pictures in the current video sequence (S1610).

[0177] If the determination result indicates that there are no independently coded sub-pictures in the current video sequence ("NO" in S1610), the image coding apparatus may determine whether wraparound motion compensation is available for the current video sequence based on a predetermined wraparound constraint (S1640). In this case, the image coding apparatus may encode flag information (e.g., sps_wraparound_enabled_flag) indicating whether wraparound motion compensation is available for the current video sequence to a first value (e.g., 0) or a second value (e.g., 1) based on the determination.

[0178] As an example of the wraparound constraint, if wraparound motion compensation is restricted for one or more output layer sets (OLSs) specified by a video parameter set (VPS), wraparound motion compensation may be restricted to be unavailable for the current video sequence. As another example of the wraparound constraint, if all sub-pictures in the current video sequence have discontinuous sub-picture boundaries, wraparound motion compensation may be restricted to be unavailable for the current video sequence.

[0179] Alternatively, if there are one or more independently coded sub-pictures in the current video sequence ("YES" in S1610), the image coding device may determine whether at least one of the independently coded sub-pictures has a width different from the picture width (S1620).

[0180] In one embodiment, the picture width can be derived as shown in Equation 1 based on the maximum width that a picture can have in the current video sequence.

[0181]

number

[0182] Here, pic_width_max_in_luma_samples represents the maximum width of the picture in luma sample units, CtbSizeY represents the CTB (coding tree block) width in the picture in luma sample units, and CtbLog2SizeY can represent the log scale value of CtbSizeY.

[0183] If the determination result indicates that at least one of the independently coded sub-pictures has a width different from the picture width ("YES" in S1620), the image coding apparatus may determine that wraparound motion compensation is not available for the current video sequence (S1630). In this case, the image coding apparatus may code sps_ref_wraparound_enabled_flag to a first value (e.g., 0) based on the determination.

[0184] Alternatively, if all independently coded sub-pictures have the same width as the picture width ("NO" in S1620), the image coding apparatus may determine whether wraparound motion compensation is available for the current video sequence based on the wraparound constraints described above (S1640). In this case, the image coding apparatus may code sps_ref_wraparound_enabled_flag to a first value (e.g., 0) or a second value (e.g., 1) based on the determination.

[0185] 16 shows that steps S1610 and S1620 are performed sequentially, this is merely an example and does not limit the embodiments of the present disclosure. For example, step S1620 may be performed simultaneously with step S1610 or may be performed before step S1610.

[0186] Meanwhile, the sps_ref_wraparound_enabled_flag coded by the image coding device can be stored in the bitstream and signaled to the image decoding device, in which case the image decoding device can determine whether wraparound motion compensation is available for the current video sequence based on the sps_ref_wraparound_enabled_flag obtained from the bitstream.

[0187] For example, if sps_ref_wraparound_enabled_flag has a first value (e.g., 0), the image decoding apparatus may determine that wraparound motion compensation is not available for the current video sequence and may not perform wraparound motion compensation on the current block. In this case, the reference sample position of the current block may be clipped with respect to a reference picture boundary or a sub-picture boundary, and motion compensation may be performed using the reference sample at the clipped position.

[0188] That is, the image decoding apparatus can perform correct motion compensation according to the present disclosure without separately determining whether or not there are one or more sub-pictures in the current video sequence that are independently coded and have a width different from the picture width. However, the operation of the image decoding apparatus is not limited thereto. For example, the image decoding apparatus can determine whether or not there are one or more sub-pictures in the current video sequence that are independently coded and have a width different from the picture width, and then perform motion compensation based on the determination result. More specifically, the image decoding apparatus determines whether or not there are one or more sub-pictures in the current video sequence that are independently decoded and have a width different from the picture width, and if such sub-pictures exist, it may regard sps_ref_wraparound_enabled_flag as a first value (e.g., 0) and not perform wraparound motion compensation.

[0189] Alternatively, if sps_ref_wraparound_enabled_flag has a second value (e.g., 1), the image decoding apparatus may determine that wraparound motion compensation is available for the current video sequence. In this case, the image decoding apparatus may additionally acquire flag information (e.g., pps_ref_wraparound_enabled_flag) indicating whether wraparound motion compensation is available for the current picture from the bitstream, and may determine whether to perform wraparound motion compensation on the current block based on the additionally acquired flag information.

[0190] For example, if pps_ref_wraparound_enabled_flag has a first value (e.g., 0), the image decoding apparatus may not perform wraparound motion compensation on the current block. In this case, the reference sample position of the current block may be clipped with reference to a reference picture boundary or a sub-picture boundary, and motion compensation may be performed using the reference sample at the clipped position. In contrast, if pps_ref_wraparound_enabled_flag has a second value (e.g., 1), the image decoding apparatus may perform wraparound motion compensation on the current block.

[0191] As described above, according to the first embodiment of the present disclosure, if a current video sequence contains one or more independently coded sub-pictures with widths different from the picture width, wraparound motion compensation may be restricted from being used for the current video sequence. This may imply that if all independently coded sub-pictures in the current video sequence have the same width as the picture width, wraparound motion compensation may be applied to each sub-picture in the current video sequence regardless of whether it is independently coded. This allows sub-picture-related coding tools and wraparound motion compensation-related coding tools to be used together, thereby further improving encoding / decoding efficiency.

[0192] Example 2

[0193] According to a second embodiment of the present disclosure, if one or more sub-pictures having a width different from the picture width exist in the current video sequence, wraparound motion compensation may be restricted from being available for the current video sequence. In this case, flag information (e.g., sps_ref_wraparound_enabled_flag) indicating whether wraparound motion compensation is available for the current video sequence may be restricted to have a first value (e.g., 0).

[0194] FIG. 17 is a flowchart showing a method for determining whether or not wraparound motion compensation is available in an image coding apparatus according to an embodiment of the present disclosure.

[0195] Referring to FIG. 17, the image coding apparatus may determine whether one or more sub-pictures having a width different from the picture width are present in the current video sequence (S1710).

[0196] If the determination result indicates that one or more sub-pictures having a width different from the picture width exist in the current video sequence ("YES" in S1710), the image coding device may determine that wraparound motion compensation is not available for the current video sequence (S1720). In this case, the image coding device may code sps_ref_wraparound_enabled_flag to a first value (e.g., 0).

[0197] Alternatively, if there are no sub-pictures in the current video sequence that have a width different from the picture width (i.e., the widths of all sub-pictures are equal to the picture width) (“NO” in S1710), the image coding apparatus may determine whether wraparound motion compensation is available for the current video sequence based on a predetermined wraparound constraint (S1730). An example of the wraparound constraint is as described above with reference to FIG. 16. In this case, the image coding apparatus may code sps_ref_wraparound_enabled_flag to a first value (e.g., 0) or a second value (e.g., 1) based on the determination.

[0198] Meanwhile, the sps_ref_wraparound_enabled_flag coded by the image coding device can be stored in the bitstream and signaled to the image decoding device, in which case the image decoding device can determine whether wraparound motion compensation is available for the current video sequence based on the sps_ref_wraparound_enabled_flag obtained from the bitstream.

[0199] For example, if sps_ref_wraparound_enabled_flag has a first value (e.g., 0), the image decoding apparatus may determine that wraparound motion compensation is not available for the current video sequence and may not perform wraparound motion compensation on the current block. In contrast, if sps_ref_wraparound_enabled_flag has a second value (e.g., 1), the image decoding apparatus may determine that wraparound motion compensation is available for the current video sequence. In this case, the image decoding apparatus may additionally acquire flag information (e.g., pps_ref_wraparound_enabled_flag) indicating whether wraparound motion compensation is available for the current picture from the bitstream, and may determine whether to perform wraparound motion compensation on the current block based on the additionally acquired flag information.

[0200] As described above, according to the second embodiment of the present disclosure, if there are one or more sub-pictures in the current video sequence that have a width different from the picture width, wraparound motion compensation may be restricted from being used for the current video sequence. This may imply that if all sub-pictures in the current video sequence have the same width as the picture width, wraparound motion compensation may be applied to each sub-picture in the current video sequence regardless of whether it is independently coded. This allows sub-picture-related coding tools and wraparound motion compensation-related coding tools to be used together, thereby further improving encoding / decoding efficiency.

[0201] Example 3

[0202] According to Example 3 of the present disclosure, if the current sub-picture has the same width as the current picture even if it is coded independently, or if the current sub-picture is not coded independently, wraparound motion compensation can be performed on the current block.

[0203] FIG. 18 is a flowchart showing a method for performing wraparound motion compensation by an image encoding device according to an embodiment of the present disclosure.

[0204] Referring to FIG. 18, the image coding apparatus may determine whether the current sub-picture is to be independently coded (S1810).

[0205] If the current sub-picture is coded independently ("YES" in S1810), the image coding apparatus may determine whether the width of the current sub-picture is equal to the width of the current picture (S1820).

[0206] If the width of the current sub-picture is equal to the width of the current picture (“YES” in S1820), the image coding apparatus can determine whether wraparound motion compensation is available (S1830).

[0207] For example, if at least one of the independently coded sub-pictures in the current video sequence has a width different from the picture width, the image coding apparatus may determine that wraparound motion compensation is not available. In this case, the image coding apparatus may code flag information (e.g., sps_ref_wraparound_enabled_flag) indicating whether wraparound motion compensation is available for the current video sequence to a first value (e.g., 0) based on the determination.

[0208] Alternatively, if all independently coded sub-pictures in the current video sequence have the same width as the picture width, the image coding device can determine whether wraparound motion compensation is available based on predetermined wraparound constraints.

[0209] As an example of the wraparound constraint, if the CTB width (e.g., CtbSizeY) is greater than the picture width (e.g., pic_width_in_luma_samples), wraparound motion compensation may be restricted to not be available for the current picture. As another example of the wraparound constraint, if wraparound motion compensation is restricted for one or more output layer sets (OLSs) specified by a video parameter set (VPS), wraparound motion compensation may be restricted to not be available for the current picture. As another example of the wraparound constraint, if all sub-pictures in the current picture have discontinuous sub-picture boundaries, wraparound motion compensation may be restricted to not be available for the current picture.

[0210] If the determination result indicates that wraparound motion compensation is available ("YES" in S1830), the image coding apparatus may perform wraparound motion compensation on the current block based on the boundaries of the current subpicture (S1840). For example, if the reference sample position of the current block deviates from the left boundary of the reference picture (e.g., xInti<0), the image coding apparatus may perform wraparound motion compensation by shifting the x coordinate of the reference sample in the positive direction by the wraparound offset (e.g., PpsRefWraparoundOffset×MinCbSizeY) and then clipping the reference sample based on the left and right boundaries of the current subpicture. Alternatively, if the reference sample position of the current block deviates from the right boundary of the reference picture (e.g., xInti>picW-1), the image coding apparatus may perform wraparound motion compensation by shifting the x coordinate of the reference sample in the negative direction by the wraparound offset and then clipping the reference sample based on the left and right boundaries of the current subpicture.

[0211] Alternatively, if wraparound motion compensation is not available ("NO" in S1830), the image coding apparatus may not perform wraparound motion compensation on the current block (S1850). In this case, the image coding apparatus may clip the reference sample position of the current block based on the boundary of the current sub-picture, and perform motion compensation using the reference sample at the clipped position.

[0212] Returning to step S1810 again, if the current subpicture is not independently coded ("NO" in S1810), the image coding apparatus may determine whether wraparound motion compensation is available (S1860). The specific method for this determination is as described above in step S1830.

[0213] If the determination result indicates that wraparound motion compensation is available ("YES" in S1860), the image coding apparatus may perform wraparound motion compensation on the current block based on the reference picture boundary (S1870). In this case, the image coding apparatus may perform wraparound motion compensation by shifting the reference sample position of the current block by the wraparound offset and then clipping the shifted position based on the reference picture boundary.

[0214] Alternatively, if wraparound motion compensation is not available ("NO" in S1860), the image coding apparatus may not perform wraparound motion compensation on the current block (S1880). In this case, the image coding apparatus may clip the reference sample position of the current block based on the boundary of the reference picture, and perform motion compensation using the reference sample at the clipped position.

[0215] Meanwhile, the image decoding apparatus may determine whether the current subpicture is independently coded based on subpicture-related information (e.g., subpic_treated_as_pic_flag) acquired from the bitstream. For example, if subpic_treated_as_pic_flag has a first value (e.g., 0), the current subpicture may not be independently coded. In contrast, if subpic_treated_as_pic_flag has a second value (e.g., 1), the current subpicture may be independently coded.

[0216] The image decoding apparatus may then determine whether wraparound motion compensation is available based on wraparound-related information (e.g., pps_ref_wraparound_enabled_flag) obtained from the bitstream, and perform wraparound motion compensation on the current block based on the determination. For example, if pps_ref_wraparound_enabled_flag has a first value (e.g., 0), the image decoding apparatus may determine that wraparound motion compensation is not available for the current picture and may not perform wraparound motion compensation on the current block. In contrast, if pps_ref_wraparound_enabled_flag has a second value (e.g., 1), the image decoding apparatus may determine that wraparound motion compensation is available for the current picture and may perform wraparound motion compensation on the current block.

[0217] The image decoding apparatus can perform correct motion compensation according to the present disclosure without separately determining whether the width of the current sub-picture is equal to the width of the current picture. However, the operation of the image decoding apparatus is not limited thereto. For example, as shown in FIG. 18, the image decoding apparatus can determine whether the width of the independently coded current sub-picture is equal to the width of the current picture, and then determine whether wraparound motion compensation is available based on the result of the determination.

[0218] As described above, according to the third embodiment of the present disclosure, if the current sub-picture has the same width as the current picture even if it is coded independently, or if the current sub-picture is not coded independently, wrap-around motion compensation can be performed on the current block. This allows sub-picture-related coding tools and wrap-around motion compensation-related coding tools to be used together, thereby further improving encoding / decoding efficiency.

[0219] Hereinafter, an image encoding / decoding method according to an embodiment of the present disclosure will be described in detail with reference to FIGS.

[0220] FIG. 19 is a flowchart illustrating an image encoding method according to an embodiment of the present disclosure.

[0221] The image encoding method of Fig. 19 may be performed by the image encoding apparatus of Fig. 2. For example, steps S1910 and S1920 may be performed by the inter prediction unit 180, and step S1930 may be performed by the entropy encoding unit 190.

[0222] Referring to FIG. 19, the image coding apparatus may determine whether to apply wrap-around motion compensation to the current block (S1910).

[0223] In one embodiment, the image coding apparatus may determine whether wraparound motion compensation is available based on whether there are one or more independently coded sub-pictures in a current video sequence containing the current block that have a width different from the picture width. For example, if there are one or more independently coded sub-pictures in the current video sequence that have a width different from the picture width, the image coding apparatus may determine that wraparound motion compensation is unavailable. Alternatively, if all independently coded sub-pictures in the current video sequence have the same width as the picture width, the image coding apparatus may determine that wraparound motion compensation is available based on a predetermined wraparound constraint. Here, examples of the wraparound constraint are as described above with reference to Figures 16 to 18.

[0224] Then, the image coding apparatus can determine whether to apply wraparound motion compensation to the current block based on the determination. For example, if wraparound motion compensation is not available for the current picture, the image coding apparatus can determine not to perform wraparound motion compensation on the current block. Conversely, if wraparound motion compensation is available for the current picture, the image coding apparatus can determine to perform wraparound motion compensation on the current block.

[0225] The image coding apparatus may generate a predicted block of the current block by performing inter prediction based on the determination result of step S1910 (S1920). For example, when wrap-around motion compensation is applied to the current block, the image coding apparatus may shift the reference sample position of the current block by a wrap-around offset and clip the shifted position based on a reference picture boundary or a sub-picture boundary. The image coding apparatus may then generate a predicted block of the current block by performing motion compensation using the reference sample at the clipped position. On the other hand, when wrap-around motion compensation is not applied to the current block, the image coding apparatus may clip the reference sample position of the current block based on a reference picture boundary or a sub-picture boundary, and then perform motion compensation using the reference sample at the clipped position to generate a predicted block of the current block.

[0226] The image encoding apparatus may generate a bitstream by encoding inter prediction information of the current block and wrap-around information related to wrap-around motion compensation (S1930).

[0227] In one embodiment, the wraparound information may include a first flag (e.g., sps_ref_wraparound_enabled_flag) indicating whether wraparound motion compensation is available for a video sequence including the current block. The first flag may have a first value (e.g., 0) indicating that wraparound motion compensation is not available for the current video sequence based on the presence of one or more independently coded sub-pictures having a width different from the width of the current picture including the current block in the current video sequence. In this case, the width of the current picture may be derived based on information regarding the maximum width a picture can have in the current video sequence (e.g., pic_width_max_in_luma_samples), as described above with reference to Equation 1.

[0228] In one embodiment, the wraparound information may further include a second flag (e.g., pps_ref_wraparound_enabled_flag) indicating whether wraparound motion compensation is available for the current picture. The second flag may have a first value (e.g., 0) indicating that wraparound motion compensation is not available for the current picture based on the first flag having a first value (e.g., 0). The second flag may also have a first value (e.g., 0) indicating that wraparound motion compensation is not available for the current picture based on a predetermined condition related to the width of a coding tree block (CTB) in the current picture and the width of the current picture. For example, if the width of a CTB in the current picture (e.g., CtbSizeY) is greater than the picture width (e.g., pic_width_in_luma_samples), sps_ref_wraparound_enabled_flag may be limited to a first value (e.g., 0).

[0229] In one embodiment, the wraparound information may further include a wraparound offset (e.g., pps_ref_wraparound_offset) based on whether the wraparound motion compensation is available for the current picture. The image coding apparatus may perform wraparound motion compensation based on the wraparound offset.

[0230] FIG. 20 is a flowchart illustrating an image decoding method according to one embodiment of the present disclosure.

[0231] The image decoding method of Figure 20 may be performed by the image decoding apparatus of Figure 3. For example, steps S2010 and S2020 may be performed by the inter prediction unit 260.

[0232] Referring to FIG. 20, the image decoding apparatus may obtain inter prediction information and wrap-around information of a current block from a bitstream (S2010). Here, the inter prediction information of the current block may include motion information of the current block, such as a reference picture index and differential motion vector information. The wrap-around information may include information about wrap-around motion compensation and a first flag (e.g., sps_ref_wraparound_enabled_flag) indicating whether wrap-around motion compensation is available for a current video sequence including the current block. The first flag may have a first value (e.g., 0) indicating that wrap-around motion compensation is not available based on the presence in the current video sequence of one or more independently coded sub-pictures having a width different from that of the current picture including the current block. In this case, the width of the current picture may be derived based on information about the maximum width a picture can have in the current video sequence (e.g., pic_width_max_in_luma_samples), as described above with reference to Equation 1.

[0233] In one embodiment, the wraparound information may further include a second flag (e.g., pps_ref_wraparound_enabled_flag) indicating whether wraparound motion compensation is available for the current picture. The second flag may have a first value (e.g., 0) indicating that wraparound motion compensation is not available for the current picture based on the first flag having a first value (e.g., 0). The second flag may also have a first value (e.g., 0) indicating that wraparound motion compensation is not available for the current picture based on a predetermined condition regarding the width of a coding tree block (CTB) in the current picture and the width of the current picture. For example, if the width of the CTB in the current picture (e.g., CtbSizeY) is greater than the picture width (e.g., pic_width_in_luma_samples), sps_ref_wraparound_enabled_flag may be limited to a first value (e.g., 0).

[0234] The image decoding apparatus may generate a prediction block for the current block based on the inter prediction information and wraparound information obtained from the bitstream (S2020).

[0235] For example, if wraparound motion compensation is available for the current picture (e.g., pps_ref_wraparound_enabled_flag==1), wraparound motion compensation may be performed on the current block. In this case, the image decoding apparatus may shift the reference sample position of the current block by the wraparound offset and clip the shifted position based on the reference picture boundary or sub-picture boundary. Then, the image decoding apparatus may generate a prediction block for the current block by performing motion compensation using the reference sample at the clipped position.

[0236] Alternatively, if wraparound motion compensation is not available for the current picture (e.g., pps_ref_wraparound_enabled_flag==0), the image coding device may generate a prediction block for the current block by clipping the reference sample position of the current block based on the reference picture boundary or sub-picture boundary, and then performing motion compensation using the reference sample at the clipped position.

[0237] As described above, according to the image encoding / decoding method according to an embodiment of the present disclosure, if all independently coded sub-pictures in a current video sequence have the same width as the picture width, wrap-around motion compensation can be used for all sub-pictures in the current video sequence. This allows sub-picture-related coding tools and wrap-around motion compensation-related coding tools to be used together, thereby further improving encoding / decoding efficiency.

[0238] The names of syntax elements described in this disclosure may include information about the location where the syntax element is signaled. For example, a syntax element beginning with "sps_" may mean that the syntax element is signaled in a sequence parameter set (SPS). Furthermore, a syntax element beginning with "pps_", "ph_", "sh_", etc. may mean that the syntax element is signaled in a picture parameter set (PPS), picture header, slice header, etc., respectively.

[0239] Although the exemplary method of the present disclosure is expressed as a series of operations for clarity of explanation, this is not intended to limit the order in which the steps are performed, and the steps may be performed simultaneously or in a different order if necessary. To achieve the method according to the present disclosure, the steps illustrated may include other steps, or some steps may be omitted and the remaining steps may be included, or some steps may be omitted and additional other steps may be included.

[0240] In the present disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) can perform the operation (step) to check the execution conditions and circumstances of the operation (step). For example, if it is described that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or the image decoding device can perform the predetermined operation after performing an operation to check whether the predetermined condition is satisfied.

[0241] The various embodiments of the present disclosure are not intended to enumerate all possible combinations, but are intended to describe representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combination of two or more.

[0242] Additionally, various embodiments of the present disclosure may be implemented using hardware, firmware, software, or a combination thereof, etc. In the case of a hardware implementation, the implementation may be using one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), general processors, controllers, microcontrollers, microprocessors, etc.

[0243] In addition, an image decoding apparatus and an image encoding apparatus to which an embodiment of the present disclosure is applied may be included in a multimedia broadcast transmitting / receiving apparatus, a mobile communication terminal, a home cinema video apparatus, a digital cinema video apparatus, a surveillance camera, a video conversation apparatus, a real-time communication apparatus such as video communication, a mobile streaming apparatus, a storage medium, a camcorder, a video on demand (VoD) service providing apparatus, an over-the-top (OTT) video apparatus, an internet streaming service providing apparatus, a three-dimensional (3D) video apparatus, an image telephone video apparatus, a medical video apparatus, etc., and may be used to process a video signal or a data signal. For example, an over-the-top (OTT) video apparatus may include a game console, a Blu-ray player, an internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.

[0244] FIG. 21 is a diagram illustrating a content streaming system to which the embodiments of the present disclosure can be applied.

[0245] As shown in FIG. 21, a content streaming system to which an embodiment of the present disclosure is applied can broadly include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0246] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, or video camera directly generates a bitstream, the encoding server can be omitted.

[0247] The bitstream can be generated by an image encoding method and / or image encoding device to which an embodiment of the present disclosure is applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.

[0248] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server serves as an intermediary for informing the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which may control commands and responses between devices in the content streaming system.

[0249] The streaming server may receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content may be received in real time. In this case, the streaming server may store the bitstream for a certain period of time to provide a smooth streaming service.

[0250] Examples of the user device include a mobile phone, a smartphone, a laptop computer, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation system, a slate PC, a tablet PC, an ultrabook, a wearable device such as a smartwatch, smart glass, a head mounted display (HMD), a digital TV, a desktop computer, and digital signage.

[0251] Each server in the content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.

[0252] FIG. 22 is a diagram illustrating a schematic architecture for providing 3D image / video services that can be utilized by embodiments of the present disclosure.

[0253] 22 may show a 360-degree or omnidirectional video / image processing system. The system of FIG. 22 may be implemented, for example, in an extended reality (XR) assistive device. That is, the system may provide a method for providing a virtual reality experience to a user.

[0254] Extended reality is a general term for virtual reality (VR), augmented reality (AR), and mixed reality (MR). VR technology provides real-world objects and backgrounds only as CG images, AR technology provides virtual CG images on top of images of real objects, and MR technology is a computer graphics technology that combines and presents virtual objects in the real world.

[0255] MR technology is similar to AR technology in that it shows real and virtual objects together, but it differs in that in AR technology, virtual objects are used to complement real objects, while in MR technology, virtual objects and real objects are used equally.

[0256] XR technology can be applied to HMDs (Head-Mount Displays), HUDs (Head-Up Displays), mobile phones, tablet PCs, laptops, desktops, TVs, digital signage, etc. A device to which XR technology is applied can be called an XR device. The XR device can include a first digital device and / or a second digital device, which will be described later.

[0257] 360-degree content generally refers to content for realizing and providing VR and can include 360-degree video and / or 360-degree audio. 360-degree video may refer to video or image content simultaneously captured or played in all directions (360 degrees or less) necessary to provide VR. Hereinafter, 360-degree video may refer to 360-degree video. 360-degree audio may also refer to spatial audio content for providing VR, in which the sound source can be perceived as being located in a specific three-dimensional space. 360-degree content can be generated, processed, and transmitted to a user, who can then consume a VR experience using the 360-degree content. 360-degree video may also be referred to as omnidirectional video, and 360-degree images may also be referred to as omnidirectional images. While the following description will be based on 360-degree video, the embodiments of this document are not limited to VR and may include processing of video / image content such as AR and MR. 360-degree video can refer to video or images displayed in various forms of 3D space according to a 3D model; for example, 360-degree video can be displayed on a spherical surface.

[0258] This method particularly proposes a method for effectively providing 360-degree video. To provide 360-degree video, first, 360-degree video is captured using one or more cameras. The captured 360-degree video is transmitted through a series of processes, and the receiving side processes and renders the received data back into the original 360-degree video. In this way, the 360-degree video can be provided to the user.

[0259] Specifically, the entire process for providing a 360-degree video may include a capture process, a preparation process, a transmission process, a processing process, a rendering process, and / or a feedback process.

[0260] The capture process may refer to a process of capturing images or videos for each of a plurality of viewpoints via one or more cameras. Image / video data such as 2210 in FIG. 22 may be generated by the capture process. Each plane in 2210 in FIG. 22 may represent an image / video for each viewpoint. The captured images / videos may also be referred to as raw data. Metadata related to the capture may be generated during the capture process.

[0261] For this capture, a special camera for VR can be used. In some embodiments, when providing a 360-degree video of a computer-generated virtual space, capture via a real camera may not be performed. In this case, the capture process may be replaced by a process in which related data is simply generated.

[0262] The preparation process may be a process of processing the captured images / videos and metadata generated during the capture process. During the preparation process, the captured images / videos may undergo a stitching process, a projection process, a region-wise packing process, and / or an encoding process.

[0263] First, each image / video may undergo a stitching process, which may be a process of connecting each captured image / video to create one panoramic or spherical image / video.

[0264] The stitched image / video may then undergo a projection process. In the projection process, the stitched image / video may be projected onto a 2D image. This 2D image may also be called a 2D image frame depending on the context. Projecting as a 2D image may also be expressed as mapping onto a 2D image. The projected image / video data may be in the form of a 2D image such as 2220 in FIG. 22.

[0265] The video data projected onto the 2D image may undergo a region-wise packing process to improve video coding efficiency. The region-wise packing may refer to a process of dividing the video data projected onto the 2D image into regions and processing them. Here, a region may refer to an area into which a 2D image onto which 360-degree video data is projected is divided. Depending on the embodiment, these regions may be divided by equally dividing the 2D image or by arbitrarily dividing it. Depending on the embodiment, these regions may also be distinguished by a projection scheme. The region-wise packing process is an optional process and may be omitted in the preparation process.

[0266] Depending on the embodiment, this process may include rotating or rearranging each region in the 2D image to improve video coding efficiency, for example, by rotating the regions so that certain edges of the regions are closer to each other, thereby improving coding efficiency.

[0267] Depending on the embodiment, this processing process may include a process of increasing or decreasing the resolution for a specific region to differentially equalize the resolution for each region on the 360-degree video. For example, a region corresponding to a relatively more important region on the 360-degree video may have a higher resolution than other regions. The video data projected onto the 2D image or the video data packed for each region may undergo an encoding process using a video codec.

[0268] Depending on the embodiment, the preparation process may further include an editing process, etc. In this editing process, editing of image / video data before and after projection may be further performed. Similarly, in the preparation process, metadata regarding stitching / projection / encoding / editing, etc. may be generated. In addition, metadata regarding an initial viewpoint or ROI (Region of Interest) of the video data projected onto the 2D image may be generated.

[0269] The transmission process may be a process of processing and transmitting image / video data and metadata that have undergone a preparation process. Processing according to any transmission protocol may be performed for transmission. Data that has been processed for transmission may be transmitted via a broadcast network and / or broadband. This data may also be transmitted to a receiving side on an on-demand basis. The receiving side may receive the data via various routes.

[0270] The processing process may refer to a process of decoding received data and re-projecting the projected image / video data onto a 3D model. In this process, the image / video data projected onto a 2D image may be re-projected onto 3D space. This process may also be called mapping or projection depending on the context. In this case, the mapped 3D space may have different shapes depending on the 3D model. For example, the 3D model may be a sphere, cube, cylinder, pyramid, etc.

[0271] Depending on the embodiment, the processing process may further include an editing process, an upscaling process, etc. In this editing process, editing of image / video data before and after reprojection may be further performed. If the image / video data is reduced in size, the size may be increased through sample upscaling in the upscaling process. If necessary, size reduction may also be performed through downscaling.

[0272] The rendering process may refer to the process of rendering and displaying image / video data reprojected onto a 3D space. In some cases, a combination of reprojection and rendering may be expressed as rendering onto a 3D model. The image / video reprojected onto (or rendered onto) a 3D model may have a form such as 2230 in FIG. 22. 2230 in FIG. 22 illustrates a case where the image / video is reprojected onto a spherical 3D model. A user can view a portion of the rendered image / video through a VR display or the like. In this case, the area viewed by the user may have a form such as 2240 in FIG. 22.

[0273] The feedback process may refer to a process of transmitting various feedback information obtained in the display process to a transmitting side. Interactivity can be provided in 360-degree video consumption through the feedback process. Depending on the embodiment, head orientation information, viewport information indicating the area the user is currently viewing, etc. may be transmitted to the transmitting side during the feedback process. Depending on the embodiment, the user may interact with what is realized in the VR environment, and in this case, information related to the interaction may be transmitted to the transmitting side or the service provider during the feedback process. Depending on the embodiment, the feedback process may not be performed.

[0274] Head orientation information may refer to information about the position, angle, movement, etc. of the user's head. Based on this information, information about the area the user is currently viewing in the 360-degree video, i.e., viewport information, may be calculated.

[0275] Viewport information may be information about the area the user is currently looking at in a 360-degree video. This allows for gaze analysis to determine how the user consumes the 360-degree video and how much they gaze at which area of ​​the 360-degree video. Gaze analysis is performed on the receiving side and can be transmitted to the transmitting side via a feedback channel. A device such as a VR display can extract the viewport area based on the user's head position / direction, vertical or horizontal field of view (FOV) information supported by the device, etc.

[0276] Meanwhile, 360-degree videos / images can be processed based on subpictures. Projected pictures or packed pictures including 2D images can be divided into subpictures, and processing can be performed on a subpicture-by-subpicture basis. For example, a specific subpicture can be given a high resolution depending on the user viewport, or only a specific subpicture can be encoded and signaled to a receiving device (decoding device side). In this case, the decoding device can receive the subpicture bitstream, restore / decode the specific subpicture, and render it according to the user viewport.

[0277] According to an embodiment, the feedback information may be not only transmitted to the transmitting side but also consumed by the receiving side. That is, the receiving side may perform decoding, reprojection, rendering, etc. using the feedback information. For example, only the 360-degree video for the area currently viewed by the user may be preferentially decoded and rendered using head orientation information and / or viewport information.

[0278] Here, the term "viewport" or "viewport area" refers to the area a user views in a 360-degree video. The term "viewpoint" refers to the point a user views in a 360-degree video, which is the exact center of the viewport area. That is, the viewport is an area centered on the viewpoint, and the size and shape of the area can be determined by the FOV (Field Of View).

[0279] In the overall architecture for providing 360-degree video described above, image / video data that undergoes a series of processes of capture / projection / encoding / transmission / decoding / reprojection / rendering can be called 360-degree video data. The term 360-degree video data may also be used as a concept that includes metadata or signaling information related to such image / video data.

[0280] In order to store and transmit media data such as the above-mentioned audio or video, a standardized media file format may be defined. According to an embodiment, the media file may have a file format based on ISO BMFF (ISO base media file format).

[0281] The scope of the present disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of various embodiments to be performed on a device or computer, and non-transitory computer-readable medium on which such software or commands can be stored and executed on a device or computer. [Industrial Applicability]

[0282] The embodiments of the present disclosure can be used to encode / decode images.

Claims

1. An image decoding method performed by an image decoding device, comprising: obtaining inter prediction information and wraparound information for a current block in a current picture from a bitstream; generating a prediction block of the current block based on the inter prediction information and the wraparound information; the wraparound information includes a first flag and a second flag; the first flag indicates whether wraparound motion compensation is available for a current video sequence that includes the current picture; the second flag indicates whether the wraparound motion compensation is available for the current picture; the first flag obtained from the bitstream is constrained to have a first value indicating that the wraparound motion compensation is not available for the current video sequence based on the current video sequence including at least one sub-picture that is treated as a picture and has a width that differs from a value derived based on information about a maximum width of pictures in the current video sequence; the second flag has a first value indicating that the wraparound motion compensation is not available for the current picture based on the first flag having the first value.

2. An image coding method performed by an image coding device, comprising: determining whether wraparound motion compensation is applied to a current block in a current picture; generating a predicted block of the current block by performing inter prediction based on the determination; encoding inter prediction information of the current block and wrap-around information of the wrap-around motion compensation; the wraparound information includes a first flag and a second flag; the first flag indicates whether wraparound motion compensation is available for a current video sequence that includes the current picture; the second flag indicates whether the wraparound motion compensation is available for the current picture; the coded first flag has a first value indicating that the wraparound motion compensation is not available for the current video sequence based on the current video sequence including at least one sub-picture that is treated as a picture and has a width that differs from a value derived based on information about a maximum width of pictures in the current video sequence; the coded second flag has a first value indicating that wraparound motion compensation is not available for the current picture based on the first flag having the first value.

3. 1. A method for transmitting a bitstream generated by an image coding method, comprising: The image encoding method includes: determining whether wraparound motion compensation is applied to a current block in a current picture; generating a predicted block of the current block by performing inter prediction based on the determination; encoding inter prediction information of the current block and wrap-around information of the wrap-around motion compensation; the wraparound information includes a first flag and a second flag; the first flag indicates whether wraparound motion compensation is available for a current video sequence that includes the current picture; the second flag indicates whether the wraparound motion compensation is available for the current picture; the coded first flag has a first value indicating that the wraparound motion compensation is not available for the current video sequence based on the current video sequence including at least one sub-picture that is treated as a picture and has a width that differs from a value derived based on information about a maximum width of pictures in the current video sequence; the coded second flag has a first value indicating that the wraparound motion compensation is not available for the current picture based on the first flag having the first value.

Citation Information

Patent Citations

  • Wraparound offsets for reference picture resampling in video coding

    US20210203988A1

  • Wraparound offsets for reference picture resampling in video coding

    WO2021133979A1