Image encoding / decoding method and apparatus based on wrap-around motion compensation, and recording medium storing a bitstream.
Patent Information
- Application Number
- JP2026000137
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-04-14
- Filing Date
- 2026-01-05
- Publication Date
- 2026-10-01
- Estimated Expiration
- 2041-03-26
Smart Images

Figure 0007928021000002 
Figure 0007928021000003 
Figure 0007928021000004
Abstract
Description
[[TECHNICAL FIELD]]
[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and more particularly, to an image encoding / decoding method and apparatus based on wrap-around motion compensation, and to a recording medium storing a bitstream generated by the image encoding method / apparatus of the present disclosure. [[BACKGROUND ART]]
[0002] Recently, demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, has been increasing in various fields. As image data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to conventional image data. An increase in the amount of information or bits to be transmitted brings about an increase in transmission costs and storage costs.
[0003] Accordingly, high-efficiency image compression technology is required for effectively transmitting or storing and reproducing information of high-resolution, high-quality images. [[SUMMARY OF THE INVENTION]] [[PROBLEM TO BE SOLVED BY THE INVENTION]]
[0004] An object of the present disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0005] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus based on wrap-around motion compensation.
[0006] Another object of the present disclosure is to provide an image encoding / decoding method and apparatus based on wrap-around motion compensation for independently coded subpictures.
[0007] Another object of the present disclosure is to provide a method for transmitting a bitstream generated by the image encoding method or apparatus according to the present disclosure.
[0008] Furthermore, this disclosure aims to provide a recording medium that stores a bitstream generated by the image encoding method or apparatus according to this disclosure.
[0009] Furthermore, this disclosure aims to provide a recording medium that stores a bitstream received by the image decoding device provided herein, decoded, and used for image restoration.
[0010] The technical problems that this disclosure seeks to solve are not limited to those described above, and other technical problems not mentioned above will be clearly understood by a person with ordinary skill in the art to which this disclosure pertains from the following description. [Means for solving the problem]
[0011] A method for decoding an image according to one aspect of the present disclosure includes the steps of: obtaining inter-prediction information and wrap-around information for a current block from a bitstream; and generating a prediction block for the current block based on the inter-prediction information and the wrap-around information, wherein the wrap-around information includes a first flag indicating whether wrap-around motion compensation is available for the current video sequence containing the current block, is independently coded, and the first flag may have a first value indicating that wrap-around motion compensation is not available, based on the presence of one or more subpictures in the current video sequence having a width different from the width of the current picture containing the current block.
[0012] An image decoding apparatus according to another aspect of the present disclosure includes a memory and at least one processor, the at least one processor which obtains inter-prediction information and wrap-around information for the current block from a bitstream, generates a prediction block for the current block based on the inter-prediction information and the wrap-around information, the wrap-around information includes a first flag that is independently coded to indicate whether wrap-around motion compensation is available for the current video sequence containing the current block, and the first flag may have a first value indicating that wrap-around motion compensation is not available based on the presence of one or more subpictures in the current video sequence having a width different from the width of the current picture containing the current block.
[0013] A picture encoding method according to another aspect of the present disclosure includes the steps of: determining whether or not to apply wrap-around motion compensation to a current block; generating a predicted block of the current block by performing interpretation based on the determination; and encoding interpretation information of the current block and wrap-around information relating to the wrap-around motion compensation, wherein the wrap-around information includes a first flag indicating whether or not the wrap-around motion compensation is available for the current video sequence including the current block, is coded independently, and the first flag may have a first value indicating that the wrap-around motion compensation is not available, based on the presence of one or more subpictures in the current video sequence having a width different from the width of the current picture including the current block.
[0014] A computer-readable recording medium according to another aspect of the present disclosure can store a bitstream generated by an image encoding method or image encoding apparatus of the present disclosure.
[0015] Another aspect of the transmission method of the present disclosure can transmit a bitstream generated by an image encoding device or image encoding method of the present disclosure.
[0016] The features described above, which are a brief summary of this disclosure, are merely illustrative examples of the detailed description of this disclosure described below and do not limit the scope of this disclosure. [Effects of the Invention]
[0017] According to this disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.
[0018] Furthermore, according to this disclosure, an image coding / decoding method and apparatus based on wrap-around motion compensation can be provided.
[0019] Furthermore, this disclosure provides an image encoding / decoding method and apparatus based on wrap-around motion compensation for independently coded subpictures.
[0020] Furthermore, this disclosure provides a method for transmitting a bitstream generated by an image encoding method or apparatus according to this disclosure.
[0021] Furthermore, according to this disclosure, a recording medium storing a bitstream generated by the image encoding method or apparatus according to this disclosure can be provided.
[0022] Furthermore, this disclosure may provide a recording medium that stores a bitstream received by the image decoding device provided for this disclosure, decoded, and used for image reconstruction.
[0023] The effects obtained from this disclosure are not limited to those described above, and other effects not mentioned above will be clearly understood by a person with ordinary skill in the art to which this disclosure pertains from the following description. [Brief explanation of the drawing]
[0024] [Figure 1] FIG. 1 is a diagram schematically showing a video coding system to which an embodiment according to the present disclosure is applicable. [Figure 2] FIG. 2 is a diagram schematically showing an image encoding apparatus to which an embodiment according to the present disclosure is applicable. [Figure 3] FIG. 3 is a diagram schematically showing an image decoding apparatus to which an embodiment according to the present disclosure is applicable. [Figure 4] FIG. 4 is a flowchart schematically showing an image decoding procedure to which an embodiment according to the present disclosure is applicable. [Figure 5] FIG. 5 is a flowchart schematically showing an image encoding procedure to which an embodiment according to the present disclosure is applicable. [Figure 6] FIG. 6 is a flowchart illustrating an inter-prediction-based video / image decoding method. [Figure 7] FIG. 7 is a diagram exemplarily showing a configuration of an inter prediction unit 260 according to the present disclosure. [Figure 8] FIG. 8 is a diagram showing an example of a subpicture. [Figure 9] FIG. 9 is a diagram showing an example of an SPS including information on a subpicture. [Figure 10] FIG. 10 is a diagram showing a method of encoding an image using subpictures by an image encoding apparatus according to an embodiment of the present disclosure. [Figure 11] FIG. 11 is a diagram showing a method of decoding an image using subpictures by an image decoding apparatus according to an embodiment of the present disclosure. [Figure 12] FIG. 12 is a diagram showing an example of a 360-degree image converted into a two-dimensional picture. [Figure 13] FIG. 13 is a diagram showing an example of a horizontal wrap-around motion compensation process. [Figure 14a] FIG. 14 is a diagram showing an example of an SPS including information on wrap-around motion compensation. [Figure 14b] FIG. 15 is a diagram showing an example of a PPS including information on wrap-around motion compensation. [Figure 15] FIG. 16 is a flowchart showing a method for performing wrap-around motion compensation by an image encoding apparatus. [Figure 16] This flowchart shows how an image encoding device according to one embodiment of the present disclosure determines whether wrap-around motion compensation can be used. [Figure 17] This flowchart shows how an image encoding device according to one embodiment of the present disclosure determines whether wrap-around motion compensation can be used. [Figure 18] This flowchart shows a method by which an image encoding device according to one embodiment of the present disclosure performs wrap-around motion compensation. [Figure 19] This flowchart shows an image encoding method according to one embodiment of the present disclosure. [Figure 20] This is a flowchart showing an image decoding method according to one embodiment of the present disclosure. [Figure 21] This figure illustrates a content streaming system to which the embodiments described herein can be applied. [Figure 22] This figure schematically illustrates an architecture for providing a 3D image / video service that can be utilized in the embodiments of this disclosure. [Modes for carrying out the invention]
[0025] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings, so that they can be easily implemented by a person with ordinary skill in the art to which the present disclosure pertains. However, the present disclosure can be implemented in a variety of different forms and is not limited to the embodiments described herein.
[0026] In describing embodiments of this disclosure, if it is determined that a specific description of a known configuration or function would obscure the gist of this disclosure, such detailed description will be omitted. In the drawings, parts unrelated to the description of this disclosure will be omitted, and similar parts will be denoted by the same reference numerals.
[0027] In this disclosure, when one component is described as being “connected,” “joined,” or “linked” to another component, this can include not only direct connections but also indirect connections where another component exists between them. Furthermore, when one component is described as “containing” or “having” another component, this means, unless otherwise stated to the contrary, that it may include another component rather than excluding it.
[0028] In this disclosure, terms such as "first," "second," etc., are used solely for the purpose of distinguishing one component from another, and do not limit the order or importance of the components unless otherwise specified. Therefore, within the scope of this disclosure, a first component in one embodiment may be called a second component in another embodiment, and similarly, a second component in one embodiment may be called a first component in another embodiment.
[0029] In this disclosure, components that are distinguished from each other are used to clearly describe their respective characteristics and do not necessarily mean that the components are separate. In other words, multiple components may be integrated to constitute a single hardware or software unit, or a single component may be distributed to constitute multiple hardware or software units. Therefore, such integrated or distributed embodiments are also included in the scope of this disclosure, without needing to be specifically mentioned.
[0030] In this disclosure, the components described in various embodiments are not necessarily essential components, and some may be optional components. Therefore, embodiments consisting of a subset of the components described in one embodiment are also included in the scope of this disclosure. Furthermore, embodiments that include additional components in addition to the components described in various embodiments are also included in the scope of this disclosure.
[0031] This disclosure relates to the encoding and decoding of images, and the terms used in this disclosure may have their ordinary meanings in the art to which this disclosure pertains, unless they are newly defined in this disclosure.
[0032] In this disclosure, "picture" generally means a unit representing any one image within a specific time period, and "slice / tile" is an encoding unit that constitutes part of a picture, and a single picture can consist of one or more slices / tiles. Furthermore, a slice / tile may contain one or more CTUs (coding tree units).
[0033] In this disclosure, “pixel” or “pel” may mean the smallest unit that constitutes a picture (or image). The term “sample” may also be used as a counterpart to pixel. A sample may generally represent a pixel or a pixel value, or it may represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component.
[0034] In this disclosure, “unit” can refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information associated with that region. A unit may be used interchangeably with terms such as “sample array,” “block,” or “area,” as it may be used. Generally, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.
[0035] In this disclosure, “current block” can mean any one of the following: “current coding block,” “current coding unit,” “block to encode,” “block to decode,” or “block to process.” If prediction is performed, “current block” can mean “current prediction block” or “block to predict.” If transformation (inverse transformation) / quantization (inverse quantization) is performed, “current block” can mean “current transformation block” or “block to transform.” If filtering is performed, “current block” can mean “block to filter.”
[0036] Furthermore, in this disclosure, "current block" may mean the block containing all luma component blocks and chroma component blocks, or the "luma block of the current block," unless there is an explicit mention of a chroma block. The luma component block of the current block may be expressed with an explicit mention of a luma component block, such as "luma block" or "current luma block." Similarly, the chroma component block of the current block may be expressed with an explicit mention of a chroma component block, such as "chroma block" or "current chroma block."
[0037] In this disclosure, " / " and "," may be interpreted as "and / or." For example, "A / B" and "A, B" may be interpreted as "A and / or B." Also, "A / B / C" and "A, B, C" may mean "at least one of A, B and / or C."
[0038] In this disclosure, “or” may be interpreted as “and / or.” For example, “A or B” may mean 1) “A” only, 2) “B” only, or 3) “A and B.” Alternatively, in this disclosure, “or” may mean “additionally or alternatively.”
[0039] Overview of the video coding system
[0040] Figure 1 is a schematic diagram showing a video coding system to which the embodiments of this disclosure can be applied.
[0041] A video coding system according to one embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 can transmit encoded video and / or image information or data to the decoding device 20 via a digital storage medium or network in file or streaming format.
[0042] An encoding device 10 according to one embodiment may include a video source generation unit 11, an encoding unit 12, and a transmission unit 13. A decoding device 20 according to one embodiment may include a receiving unit 21, a decoding unit 22, and a rendering unit 23. The encoding unit 12 may be called a video / image encoding unit, and the decoding unit 22 may be called a video / image decoding unit. The transmission unit 13 may be included in the encoding unit 12. The receiving unit 21 may be included in the decoding unit 22. The rendering unit 23 may also include a display unit, which may be configured as a separate device or external component.
[0043] The video source generation unit 11 can acquire video / images through processes such as video / image capture, synthesis, or generation. The video source generation unit 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, or a video / image archive containing previously captured video / images. The video / image generation device may include, for example, a computer, tablet, and smartphone, and may generate video / images (electronically). For example, virtual video / images may be generated via a computer, in which case the video / image capture process may be replaced by a process in which the relevant data is generated.
[0044] The encoding unit 12 can encode the input video / image. The encoding unit 12 can perform a series of steps such as prediction, transformation, and quantization for compression and encoding efficiency. The encoding unit 12 can output the encoded data (encoded video / image information) in bitstream format.
[0045] The transmission unit 13 can transmit encoded video / image information or data, output in bitstream format, to the receiving unit 21 of the decoding device 20 via a digital storage medium or network in file or streaming format. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray®, HDD, and SSD. The transmission unit 13 may include elements for generating media files via a predetermined file format and elements for transmission via a broadcast / communication network. The receiving unit 21 can extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit 22.
[0046] The decoding unit 22 can decode the video / image by performing a series of steps such as inverse quantization, inverse transform, and prediction, corresponding to the operation of the encoding unit 12.
[0047] The rendering unit 23 can render the decoded video / image. The rendered video / image can be displayed via the display unit.
[0048] Overview of Image Encoding Devices
[0049] Figure 2 is a schematic diagram showing an image encoding device to which the embodiments of this disclosure can be applied.
[0050] As shown in Figure 2, the image coding device 100 may include an image splitting unit 110, a subtraction unit 115, a transformation unit 120, a quantization unit 130, an inverse quantization unit 140, an inverse transformation unit 150, an addition unit 155, a filtering unit 160, a memory 170, an inter-prediction unit 180, an intra-prediction unit 185, and an entropy coding unit 190. The inter-prediction unit 180 and the intra-prediction unit 185 can together be called the "prediction unit". The transformation unit 120, the quantization unit 130, the inverse quantization unit 140, and the inverse transformation unit 150 may be included in a residual processing unit. The residual processing unit may further include a subtraction unit 115.
[0051] All or at least some of the multiple components constituting the image encoding device 100 can be implemented by a single hardware component (e.g., an encoder or processor) depending on the embodiment. Furthermore, the memory 170 may include a DPB (decoded picture buffer) and can be implemented by a digital storage medium.
[0052] The image splitting unit 110 can split an input image (or picture, frame) input to the image encoding device 100 into one or more processing units. For example, the processing units may be called coding units (CUs). Coding units can be obtained by recursively splitting a coding tree unit (CTU) or the largest coding unit (LCU) using a QT / BT / TT (Quad-tree / binary-tree / ternary-tree) structure. For example, a single coding unit can be split into multiple coding units of deeper depth based on a quad-tree structure, a binary-tree structure and / or a ternary-tree structure. For the splitting of coding units, a quad-tree structure may be applied first, followed by a binary-tree structure and / or a ternary-tree structure. Based on the final coding unit that cannot be further split, the coding procedure according to this disclosure can be performed. The largest coding unit can be used as the final coding unit, or a lower-depth coding unit obtained by dividing the largest coding unit can be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and / or restoration, as described later. As another example, the processing units of the coding procedure may be prediction units (PU) or transformation units (TU). The prediction unit and the transformation unit may be divided or partitioned from the final coding unit, respectively. The prediction unit may be a unit of sample prediction, and the transformation unit may be a unit that derives transformation coefficients and / or a unit that derives a residual signal from transformation coefficients.
[0053] The prediction unit (inter-prediction unit 180 or intra-prediction unit 185) can make predictions for the block to be processed (current block) and generate a predicted block that includes prediction samples for the current block. The prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block or on a CU basis. The prediction unit can generate various information regarding the prediction of the current block and transmit it to the entropy coding unit 190. The prediction information can be encoded by the entropy coding unit 190 and output in bitstream format.
[0054] The intra-prediction unit 185 can predict the current block by referring to a sample in the current picture. The referenced sample may be located in the vicinity (neighbor) or at a distance from the current block, according to the intra-prediction mode and / or intra-prediction technique. The intra-prediction mode may include multiple non-directional modes and multiple directional modes. The non-directional modes may include, for example, a DC mode and a Planar mode. The directional modes may include, for example, 33 or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. The intra-prediction unit 185 may also determine the prediction mode to be applied to the current block using the prediction modes applied to the surrounding blocks.
[0055] The interprediction unit 180 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between the surrounding blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, the surrounding blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different from each other. The temporal neighboring block may be called a collocated reference block, collocated CU (colCU), etc. The reference picture containing the temporal neighboring block may be called a collocated picture (colPic). For example, the interpretation unit 180 can construct a motion information candidate list based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Interpretation can be performed based on various prediction modes; for example, in skip mode and merge mode, the interpretation unit 180 can use the motion information of surrounding blocks as the motion information of the current block. In skip mode, unlike merge mode, the residual signal may not be transmitted.In motion vector prediction (MVP) mode, the motion vector of the surrounding block is used as the motion vector predictor, and the motion vector of the current block can be signaled by encoding the motion vector difference and an indicator for the motion vector predictor. The motion vector difference can represent the difference between the motion vector of the current block and the motion vector predictor.
[0056] The prediction unit can generate a prediction signal based on various prediction methods and / or techniques described later. For example, the prediction unit can apply intra-prediction or inter-prediction to predict the current block, and can also apply intra-prediction and inter-prediction simultaneously. A prediction method that applies intra-prediction and inter-prediction simultaneously to predict the current block can be called CIIP (combined inter and intra prediction). The prediction unit can also perform intra-block copy (IBC) to predict the current block. Intra-block copy can be used for content image / video coding such as in games, for example, in SCC (screen content coding). IBC is a method of predicting the current block using a reference block that has already been restored in the current picture at a predetermined distance from the current block. When IBC is applied, the position of the reference block in the current picture can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC basically performs prediction within the current picture, but it can be performed similarly to inter-prediction in that it derives the reference block within the current picture. In other words, IBC can use at least one of the interpretation techniques described in this disclosure.
[0057] The predicted signal generated by the prediction unit can be used to generate a reconstructed signal or a residual signal. The subtraction unit 115 can generate a residual signal (residual block, residual sample array) by subtracting the predicted signal output from the prediction unit (predicted block, predicted sample array) from the input image signal (original block, original sample array). The generated residual signal can be transmitted to the conversion unit 120.
[0058] The transformation unit 120 can generate transformation coefficients by applying transformation techniques to the residual signal. For example, the transformation techniques may include at least one of the following: DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented by this graph. CNT refers to a transformation obtained by generating a prediction signal using all previously reconstructed pixels. The transformation process can be applied to pixel blocks of the same size and square shape, or to non-square, variable-sized blocks.
[0059] The quantization unit 130 can quantize the conversion coefficients and transmit them to the entropy coding unit 190. The entropy coding unit 190 can encode the quantized signal (information about the quantized conversion coefficients) and output it in bitstream format. The information about the quantized conversion coefficients can be called residual information. The quantization unit 130 can rearrange the block-form quantized conversion coefficients into a one-dimensional vector format based on the coefficient scan order, and can also generate information about the quantized conversion coefficients based on the one-dimensional vector format of the quantized conversion coefficients.
[0060] The entropy coding unit 190 can perform various coding methods, such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). In addition to the quantized conversion coefficients, the entropy coding unit 190 can also encode information necessary for video / image restoration (e.g., the values of syntax elements) together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream format in units of NAL (network abstraction layer) units. The video / image information may further include information about various parameter sets, such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). The video / image information may also further include general constraint information. The signaling information, transmitted information and / or syntax elements referred to in this disclosure may be encoded via the encoding procedure described above and included in the bitstream.
[0061] The bitstream can be transmitted over a network or stored on a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmission unit (not shown) for transmitting the signal output from the entropy encoding unit 190 and / or a storage unit (not shown) for storing it may be provided as an internal / external element of the image encoding device 100, or the transmission unit may be provided as a component of the entropy encoding unit 190.
[0062] The quantized conversion coefficients output from the quantization unit 130 can be used to generate a residual signal. For example, by applying inverse quantization and inverse transformation to the quantized conversion coefficients via the inverse quantization unit 140 and the inverse transformation unit 150, a residual signal (residual block or residual sample) can be reconstructed.
[0063] The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter-prediction unit 180 or the intra-prediction unit 185. If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The adder 155 may be called the reconstruction unit or the reconstructed block generation unit. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, or, as described later, for inter-prediction of the next picture after filtering.
[0064] The filtering unit 160 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 160 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 170, specifically in the DPB of the memory 170. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter. The filtering unit 160 can generate various filtering-related information, as will be described later in the explanation of each filtering method, and transmit it to the entropy coding unit 190. The filtering-related information can be encoded by the entropy coding unit 190 and output in bitstream format.
[0065] The corrected restored picture transmitted to memory 170 can be used as a reference picture in the interpretation unit 180. When interpretation is applied via this, the image encoding device 100 can avoid prediction mismatches between the image encoding device 100 and the image decoding device, and can also improve encoding efficiency.
[0066] The DPB in memory 170 can store the modified restored picture for use as a reference picture in the inter-prediction unit 180. Memory 170 can store motion information of blocks from which motion information in the current picture has been derived (or encoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 180 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 170 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 185.
[0067] Overview of the image decoding device
[0068] Figure 3 is a schematic diagram showing an image decoding apparatus to which the embodiments of this disclosure can be applied.
[0069] As shown in Figure 3, the image decoding device 200 can be configured to include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an additive unit 235, a filtering unit 240, a memory 250, an inter-prediction unit 260, and an intra-prediction unit 265. The inter-prediction unit 260 and the intra-prediction unit 265 can together be called the "prediction unit". The inverse quantization unit 220 and the inverse transform unit 230 can be included in the residual processing unit.
[0070] All or at least some of the multiple components constituting the image decoding device 200 can be implemented by a single hardware component (e.g., a decoder or processor) depending on the embodiment. Furthermore, the memory 170 may include a DPB and can be implemented by a digital storage medium.
[0071] An image decoding device 200, having received a bitstream containing video / image information, can restore the image by executing a process corresponding to the process performed in the image encoding device 100 in Figure 2. For example, the image decoding device 200 can perform decoding using the processing unit applied in the image encoding device. Therefore, the decoding processing unit can be, for example, a coding unit. The coding unit can be obtained by dividing a coding tree unit or a maximum coding unit. The restored image signal decoded and output via the image decoding device 200 can then be reproduced via a playback device (not shown).
[0072] The image decoding device 200 can receive the signal output from the image encoding device 2 in bitstream format. The received signal can be decoded via the entropy decoding unit 210. For example, the entropy decoding unit 210 can parse the bitstream to derive information necessary for image restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as adaptive parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The image decoding device may further use the parameter set information and / or the general constraint information to decode the image. The signaling information, received information, and / or syntax elements referred to in this disclosure can be obtained from the bitstream by decoding via the decoding procedure. For example, the entropy decoding unit 210 can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements necessary for image reconstruction and the quantized values of conversion coefficients related to the residual. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element from the bitstream, determines a context model using the syntax element information to be decoded, the decoding information of the surrounding blocks and the blocks to be decoded, or the symbol / bin information decoded in a previous step, predicts the probability of bin occurrence based on the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin.Of the information decoded by the entropy decoding unit 210, information related to prediction is provided to the prediction unit (inter-prediction unit 260 and intra-prediction unit 265), and the residual values that have undergone entropy decoding in the entropy decoding unit 210, i.e., quantized conversion coefficients and related parameter information, can be input to the inverse quantization unit 220. In addition, of the information decoded by the entropy decoding unit 210, information related to filtering can be provided to the filtering unit 240. On the other hand, a receiving unit (not shown) that receives signals output from the image coding device may be further provided as an internal / external element of the image decoding device 200, or the receiving unit may be provided as a component of the entropy decoding unit 210.
[0073] On the other hand, the image decoding device according to this disclosure may be called a video / image / picture decoding device. The image decoding device may also include an information decoder (video / image / picture information decoder) and / or a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoding unit 210, and the sample decoder may include at least one of an inverse quantization unit 220, an inverse transform unit 230, an adder unit 235, a filtering unit 240, a memory 250, an inter-prediction unit 260, and an intra-prediction unit 265.
[0074] The inverse quantization unit 220 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 220 can rearrange the quantized transformation coefficients in a two-dimensional block format. In this case, the rearrangement can be performed based on the coefficient scan order performed by the image encoding device. The inverse quantization unit 220 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) to obtain the transformation coefficients.
[0075] The inverse conversion unit 230 can inversely convert the conversion coefficients to obtain residual signals (residual blocks, residual sample arrays).
[0076] The prediction unit can make predictions for the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit 210, the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block and can determine a specific intra / inter-prediction mode (prediction technique).
[0077] The prediction unit can generate prediction signals based on various prediction methods (techniques) described later, as described in the explanation of the prediction unit of the image coding device 100.
[0078] The intra-prediction unit 265 can predict the current block by referring to the samples in the current picture. The description of the intra-prediction unit 185 can also be applied to the intra-prediction unit 265.
[0079] The interprediction unit 260 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on a reference picture. In this case, to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in block, sub-block, or sample units based on the correlation of motion information between surrounding blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In interprediction, surrounding blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the interprediction unit 260 can construct a motion information candidate list based on surrounding blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction can be performed based on various prediction modes (techniques), and the prediction information may include information indicating the mode (technique) of interprediction for the current block.
[0080] The adder 235 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit (including the inter-prediction unit 260 and / or intra-prediction unit 265). If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the reconstructed block. The description of the adder 155 also applies to the adder 235. The adder 235 is sometimes called the reconstruction unit or reconstructed block generation unit. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, or for inter-prediction of the next picture via filtering, as described later.
[0081] The filtering unit 240 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 240 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 250, specifically in the DPB of the memory 250. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter.
[0082] The restored picture stored (modified) in the DPB of memory 250 can be used as a reference picture in the inter-prediction unit 260. Memory 250 can store motion information of blocks from which motion information in the current picture has been derived (or decoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 260 for use as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. Memory 250 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 265.
[0083] In this specification, the embodiments described for the filtering unit 160, inter-prediction unit 180, and intra-prediction unit 185 of the image coding device 100 can be applied similarly or in a corresponding manner to the filtering unit 240, inter-prediction unit 260, and intra-prediction unit 265 of the image decoding device 200, respectively.
[0084] Overview of Interpretation
[0085] The following explains the inter-prediction based on this disclosure.
[0086] The prediction unit of the image encoding / decoding device according to this disclosure can perform interpretation on a block-by-block basis to derive predicted samples. Interpretation can indicate a prediction derived in a manner dependent on data elements of pictures other than the current picture (e.g., sample values or motion information). When interpretation is applied to the current block, a predicted block (predicted block or predicted sample array) for the current block can be derived based on a reference block (reference sample array) identified by motion vectors on the reference picture pointed to by the reference picture index. At this time, in order to reduce the amount of motion information transmitted in interpretation mode, the motion information of the current block can be predicted on a block, subblock, or sample basis based on the correlation of motion information between the surrounding blocks and the current block. The motion information may include motion vectors and reference picture indexes. The motion information may further include interpretation type information (L0 prediction, L1 prediction, Bi prediction, etc.). When interpretation is applied, the surrounding blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture containing the reference block and the reference picture containing the time-peripheral block may be the same or different. The time-peripheral block may be called by names such as collocated reference block, colCU, or colBlock, and the reference picture containing the time-peripheral block may be called by names such as collocated picture (colPic) or colPicture. For example, a list of motion information candidates may be constructed based on the surrounding blocks of the current block, and flags or index information indicating which candidate is selected (used) may be signaled to derive the motion vector and / or reference picture index of the current block.
[0087] Interpretation can be performed based on various prediction modes. For example, in skip mode and merge mode, the motion information of the current block may be identical to the motion information of the selected surrounding block. In skip mode, unlike merge mode, residual signals may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of the selected surrounding block can be used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference. In this disclosure, MVP mode may be used interchangeably with AMVP (Advanced Motion Vector Prediction).
[0088] The motion information may include L0 motion information and / or L1 motion information based on the interpretation type (L0 prediction, L1 prediction, Bi prediction, etc.). A motion vector in the L0 direction may be called an L0 motion vector or MVL0, and a motion vector in the L1 direction may be called an L1 motion vector or MVL1. A prediction based on the L0 motion vector may be called an L0 prediction, a prediction based on the L1 motion vector may be called an L1 prediction, and a prediction based on both the L0 motion vector and the L1 motion vector may be called a bi (Bi) prediction. Here, the L0 motion vector may represent a motion vector associated with the reference picture list L0 (L0), and the L1 motion vector may represent a motion vector associated with the reference picture list L1 (L1). The reference picture list L0 may include pictures earlier in the output order than the current picture as reference pictures, and the reference picture list L1 may include pictures later in the output order than the current picture. The aforementioned earlier pictures can be called forward (reference) pictures, and the aforementioned later pictures can be called backward (reference) pictures. The reference picture list L0 may further include pictures that are later in the output order than the current picture as reference pictures. In this case, the earlier picture may be indexed first in the reference picture list L0, and the later picture may be indexed next. The reference picture list L1 may further include pictures that are earlier in the output order than the current picture as reference pictures. In this case, the later picture may be indexed first in the reference picture list L1, and the earlier picture may be indexed next. Here, the output order may correspond to the POC (picture order count) order.
[0089] Figure 4 is a flowchart of an interpretation-based video / image coding method.
[0090] Figure 5 is a diagram illustrating the configuration of the interpretation unit 180 according to this disclosure.
[0091] The encoding method in Figure 4 can be performed by the image encoding device in Figure 2. Specifically, step S410 can be performed by the interprediction unit 180, and step S420 can be performed by the residual processing unit. Specifically, step S420 can be performed by the subtraction unit 115. Step S430 can be performed by the entropy encoding unit 190. The prediction information in step S430 is derived by the interprediction unit 180, and the residual information in step S430 can be derived by the residual processing unit. The residual information is information about the residual sample. The residual information may include information about the quantized conversion coefficients for the residual sample. As described above, the residual sample is derived as a conversion coefficient via the conversion unit 120 of the image encoding device, and the conversion coefficient can be derived as a quantized conversion coefficient via the quantization unit 130. Information about the quantized conversion coefficients can be encoded by the entropy encoding unit 190 via the residual coding procedure.
[0092] Referring together to Figures 4 and 5, the image coding device can perform inter prediction for the current block (S410). The image coding device can derive the inter prediction mode and motion information for the current block and generate prediction samples for the current block. Here, the inter prediction mode determination, motion information derivation, and prediction sample generation procedures may be performed simultaneously, or one of the procedures may be performed before the others. For example, as shown in Figure 5, the inter prediction unit 180 of the image coding device may include a prediction mode determination unit 181, a motion information derivation unit 182, and a prediction sample derivation unit 183. The prediction mode determination unit 181 can determine the prediction mode for the current block, the motion information derivation unit 182 can derive the motion information for the current block, and the prediction sample derivation unit 183 can derive prediction samples for the current block. For example, the interpretation unit 180 of the image coding device can search for blocks similar to the current block within a certain area (search area) of the reference picture via motion estimation, and derive a reference block whose difference from the current block is the minimum or below a certain standard. Based on this, it can derive a reference picture index that points to the reference picture where the reference block is located, and derive a motion vector based on the positional difference between the reference block and the current block. The image coding device can determine which of various prediction modes to apply to the current block. The image coding device can compare the rate-distortion (RD) cost for the various prediction modes and determine the optimal prediction mode for the current block. However, the method by which the image coding device determines the prediction mode for the current block is not limited to the above example, and various methods can be used.
[0093] For example, when skip mode or merge mode is applied to the current block, the image encoding device can derive merge candidates from surrounding blocks of the current block and construct a merge candidate list using the deriveted merge candidates. The image encoding device can also derive a reference block from among the reference blocks pointed to by the merge candidates included in the merge candidate list whose difference from the current block is the minimum or below a certain standard. In this case, a merge candidate associated with the derived reference block is selected, and merge index information indicating the selected merge candidate is generated and signaled to the image decoding device. The motion information of the current block can be derived using the motion information of the selected merge candidate.
[0094] As another example, when the MVP mode is applied to the current block, the image encoding device can derive motion vector predictor (mvp) candidates from the surrounding blocks of the current block and construct an mvp candidate list using the deriveted mvp candidates. The image encoding device can also use the motion vector of a selected mvp candidate from among the mvp candidates included in the mvp candidate list as the mvp of the current block. In this case, for example, the motion vector pointing to the reference block derived by the motion estimation described above can be used as the motion vector of the current block, and the mvp candidate with the smallest difference between its motion vector and the motion vector of the current block can become the selected mvp candidate. The motion vector difference (MVD), which is the difference obtained by subtracting the mvp from the motion vector of the current block, can be derived. In this case, index information indicating the selected mvp candidate and information regarding the MVD can be signaled to the image decoding device. Furthermore, when MVP mode is applied, the value of the reference picture index can be composed of reference picture index information and separately signaled to the image decoding device.
[0095] The image coding device can derive a residual sample based on the predicted sample (S420). The image coding device can derive the residual sample by comparing the original sample of the current block with the predicted sample. For example, the residual sample can be derived by subtracting the corresponding predicted sample from the original sample.
[0096] The image encoding device can encode image information including prediction information and residual information (S430). The image encoding device can output the encoded image information in bitstream format. The prediction information is information related to the prediction procedure and may include prediction mode information (e.g., skip flag, merge flag, or mode index) and motion information. Of the prediction mode information, the skip flag indicates whether or not the skip mode is applied to the current block, and the merge flag indicates whether or not the merge mode is applied to the current block. Alternatively, the prediction mode information may be information that indicates one of several prediction modes, such as the mode index. If the skip flag and merge flag are both 0, it can be determined that the MVP mode is applied to the current block. The motion information may include candidate selection information (e.g., merge index, mvp flag, or mvp index) which is information for deriving a motion vector. Of the candidate selection information, the merge index can be signaled when the merge mode is applied to the current block and may be information for selecting one of the merge candidates included in the merge candidate list. Of the candidate selection information, the mvp flag or mvp index can be signaled when MVP mode is applied to the current block, and may be information for selecting one of the mvp candidates included in the mvp candidate list. The motion information may also include the MVD information and / or reference picture index information described above. The motion information may also include information indicating whether L0 prediction, L1 prediction, or bi (Bi) prediction is applied. The residual information is information about the residual sample. The residual information may also include information about the quantized transformation coefficients for the residual sample.
[0097] The output bitstream can be stored on a (digital) storage medium and transmitted to an image decoding device, or it can be transmitted to an image decoding device via a network.
[0098] On the other hand, as mentioned above, the image coding device can generate a reconstructed picture (a picture including the reconstructed sample and the reconstructed block) based on the reference sample and the residual sample. This is because the image coding device can derive the same prediction results as the image decoding device, thereby improving coding efficiency. Therefore, the image coding device can store the reconstructed picture (or reconstructed sample, reconstructed block) in memory and use it as a picture for interpretation. As mentioned above, in-loop filtering procedures and the like can be further applied to the reconstructed picture.
[0099] Figure 6 is a flowchart illustrating an interpretation-based video / image decoding method, and Figure 7 is a diagram illustrating the configuration of the interpretation unit 260 according to this disclosure.
[0100] The image decoding device can perform operations corresponding to those performed by the image encoding device. The image decoding device can make predictions for the current block based on the received prediction information and derive prediction samples.
[0101] The decoding method in Figure 6 can be performed by the image decoding device in Figure 3. Steps S610 to S630 can be performed by the interpretation unit 260, and the prediction information in step S610 and the residual information in step S640 can be obtained from the bitstream by the entropy decoding unit 210. The residual processing unit of the image decoding device can derive a residual sample for the current block based on the residual information (S640). Specifically, the inverse quantization unit 220 of the residual processing unit derives a conversion coefficient by performing inverse quantization based on the quantized conversion coefficient derived from the residual information, and the inverse transformation unit 230 of the residual processing unit can derive a residual sample for the current block by performing an inverse transformation on the conversion coefficient. Step S650 can be performed by the addition unit 235 or the reconstruction unit.
[0102] Referring together to Figures 6 and 7, the image decoding device can determine the prediction mode for the current block based on the received prediction information (S610). Based on the prediction mode information in the prediction information, the image decoding device can determine which inter-prediction mode is applied to the current block.
[0103] For example, based on the skip flag, it can be determined whether the skip mode is applied to the current block. Alternatively, based on the merge flag, it can be determined whether the merge mode is applied to the current block or whether the MVP mode is determined. Or, based on the mode index, one of several inter-prediction mode candidates can be selected. The inter-prediction mode candidates may include the skip mode, merge mode, and / or the MVP mode, or may include various inter-prediction modes as described later.
[0104] The image decoding device can derive motion information for the current block based on the determined interprediction mode (S620). For example, if a skip mode or merge mode is applied to the current block, the image decoding device can configure a merge candidate list, which will be described later, and select one of the merge candidates included in the merge candidate list. This selection can be made based on the candidate selection information (merge index) described above. The motion information for the current block can be derived using the motion information for the selected merge candidate. For example, the motion information for the selected merge candidate can be used as the motion information for the current block.
[0105] As another example, when the MVP mode is applied to the current block, the image decoding device can configure an MVP candidate list and use the motion vector of an MVP candidate selected from the MVP candidates included in the MVP candidate list as the MVP of the current block. The selection can be made based on the candidate selection information (MVP flag or MVP index) described above. In this case, the MVD of the current block can be derived based on the information regarding the MVD, and the motion vector of the current block can be derived based on the MVP of the current block and the MVD. Furthermore, the reference picture index of the current block can be derived based on the reference picture index information. The picture pointed to by the reference picture index in the associated reference picture list for the current block can be derived as the reference picture referenced for interpretation of the current block.
[0106] The image decoding device can generate predicted samples for the current block based on the motion information of the current block (S630). In this case, the reference picture can be derived based on the reference picture index of the current block, and the predicted samples for the current block can be derived using the sample of the reference block pointed to on the reference picture by the motion vector of the current block. Depending on the case, a prediction sample filtering procedure can be further performed on all or some of the predicted samples for the current block.
[0107] For example, as shown in Figure 7, the interpretation unit 260 of the image decoding device may include a prediction mode determination unit 261, a motion information derivation unit 262, and a prediction sample derivation unit 263. The interpretation unit 260 of the image decoding device can determine a prediction mode for the current block based on prediction mode information received from the prediction mode determination unit 261, derive motion information (such as motion vectors and / or reference picture indices) for the current block based on motion information received from the motion information derivation unit 262, and derive prediction samples for the current block using the prediction sample derivation unit 263.
[0108] The image decoding device can generate a residual sample for the current block based on the received residual information (S640). The image decoding device can generate a restored sample for the current block based on the predicted sample and the residual sample, and generate a restored picture based on this (S650). As previously mentioned, further procedures such as in-loop filtering can be applied to the restored picture thereafter.
[0109] As described above, the interpretation procedure may include an interpretation mode determination step, a motion information derivation step based on the determined prediction mode, and a prediction execution step (generation of prediction samples) based on the derived motion information. The interpretation procedure can be performed by an image encoding device and an image decoding device as described above.
[0110] Sub-picture overview
[0111] The subpictures in this disclosure are described below.
[0112] A single picture can be divided into tiles, and each tile can be further divided into subpictures. Each subpicture can contain one or more slices and can form rectangular regions within the picture.
[0113] Figure 8 shows an example of a subpicture.
[0114] Referring to Figure 8, one picture can be divided into 18 tiles. Twelve tiles can be placed to the left of the picture, and each of these tiles can contain one subpicture / slice consisting of 16 CTUs. Six tiles can be placed to the right of the picture, and each of these tiles can contain two subpictures / slice consisting of 4 CTUs. As a result, the picture can be divided into 24 subpictures, and each of these subpictures can contain one slice.
[0115] Information about subpictures (e.g., the number and size of subpictures) can be encoded / signaled via higher-level syntax, such as SPS, PPS, and / or slice headers.
[0116] Figure 9 shows an example of an SPS containing information about subpictures.
[0117] Referring to Figure 9, an SPS can include a syntax element subpic_info_present_flag that indicates the presence or absence of subpicture information for a CLVS (coded layer video sequence). For example, a subpic_info_present_flag with a first value (e.g., 0) can indicate that no subpicture information exists for the CLVS, and that only one subpicture exists within each picture of the CLVS. Conversely, a subpic_info_present_flag with a second value (e.g., 1) can indicate that subpicture information exists for the CLVS, and that one or more subpictures exist within each picture of the CLVS. In one example, if the picture spatial resolution can be changed within a CLVS referencing an SPS (e.g., res_change_in_clvs_allowed_flag==1), the value of subpic_info_present_flag can be restricted to the first value (e.g., 0). On the other hand, if the bitstream contains only a subset of the subpictures of the input bitstream to the sub-bitstream extraction process as a result of the sub-bitstream extraction process, the value of subpic_info_present_flag may be restricted to a second value (e.g., 1).
[0118] Furthermore, SPS can include a syntax element sps_num_subpics_minus1 that indicates the number of subpictures. For example, the value of sps_num_subpics_minus1 plus 1 can indicate the number of subpictures contained in each picture within CLVS. In one example, the value of sps_num_subpics_minus1 can be restricted to a range of 0 or greater and less than or equal to Ceil(pic_width_max_in_luma_samplesχCtbSizeY)xCeil(pic_height_max_in_luma_samples / CtbSizeY), where Ceil(x) can be a sailing function that outputs the smallest integer value equal to or greater than x. Furthermore, pic_width_max_in_luma_samples can represent the maximum width in luma sample units for each picture, pic_height_max_in_luma_samples can represent the maximum height in luma sample units for each picture, and CtbSizeY can represent the array size of each luma component CTB (coding tree block) in both width and height. On the other hand, if sps_num_subpics_minus1 does not exist, the value of sps_num_subpics_minus1 can be inferred to be the first value (e.g., 0).
[0119] Furthermore, SPS may include a syntax element sps_independent_subpics_flag that indicates whether or not to treat subpicture boundaries as picture boundaries. For example, sps_independent_subpics_flag with a secondary value (e.g., 1) can indicate that all subpicture boundaries within CLVS are treated as picture boundaries and that loop filtering across the subpicture boundaries is not performed. Conversely, sps_independent_subpics_flag with a primary value (e.g., 0) can indicate that the above constraints do not apply. On the other hand, if sps_independent_subpics_flag is not present, its value can be inferred to be the primary value (e.g., 0).
[0120] Additionally, SPS can include syntax elements subpic_ctu_top_left_x[i], subpic_ctu_top_left_y[i], subpic_width_minus1[i], and subpic_height_minus1[i] that indicate the position and size of the subpicture.
[0121] subpic_ctu_top_left_x[i] can represent the horizontal position of the top-left CTU of the i-th subpicture in units of CtbSizeY. For example, the length of subpic_ctu_top_left_x[i] may be Ceil(Log2((pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)) bits. On the other hand, if subpic_ctu_top_left_x[i] does not exist, its value can be inferred to be the first value (e.g., 0).
[0122] subpic_ctu_top_left_y[i] can represent the vertical position of the top-left CTU of the i-th subpicture in units of CtbSizeY. For example, the length of subpic_ctu_top_left_y[i] may be Ceil(Log2((pic_height_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)) bits. On the other hand, if subpic_ctu_top_left_y[i] does not exist, its value can be inferred to be the first value (e.g., 0).
[0123] The value of subpic_width_minus1[i] plus 1 can represent the width of the i-th subpicture in units of CtbSizeY. For example, the length of subpic_width_minus1[i] may be Ceil(Log2((pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)) bits. On the other hand, if subpic_width_minus1[i] does not exist, its value can be inferred as ((pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)-subpic_ctu_top_left_x[i]-1.
[0124] The value of subpic_height_minus1[i] plus 1 can represent the height of the i-th subpicture in units of CtbSizeY. For example, the length of subpic_height_minus1[i] may be Ceil(Log2((pic_height_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)) bits. On the other hand, if subpic_height_minus1[i] does not exist, its value can be inferred as ((pic_height_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)-subpic_ctu_top_left_y[i]-1.
[0125] Furthermore, SPS may include a subpic_treated_as_pic_flag[i] indicating whether a subpicture is treated as a single picture. For example, a subpic_treated_as_pic_flag[i] with a first value (e.g., 0) may indicate that the i-th subpicture in each encoded picture within the CLVS is not treated as a single picture during the decoding process, excluding in-loop filtering. Conversely, a subpic_treated_as_pic_flag[i] with a second value (e.g., 1) may indicate that the i-th subpicture in each encoded picture within the CLVS is treated as a single picture during the decoding process, excluding in-loop filtering. If subpic_treated_as_pic_flag[i] is not present, its value can be inferred to be the same as the sps_independent_subpics_flag described above. In one example, subpic_treated_as_pic_flag[i] can only be encoded / signaled if the aforementioned sps_independent_subpics_flag has a first value (e.g., 0) (i.e., the subpicture boundary is not treated as a picture boundary).
[0126] On the other hand, if subpic_treated_as_pic_flag[i] has a second value (e.g., 1), then for each output layer and its reference layer in the output layer set (OLS) that includes the layer containing the i-th subpicture as an output layer, all of the following conditions must be true:
[0127] -(Condition 1) All pictures in the output layer and its reference layers must have the same pic_width_in_luma_samples and the same pic_height_in_luma_samples.
[0128] -(Condition 2) All SPS referenced by the output layer and its reference layer must have the same value for sps_num_subpics_minus1, and the same values for subpic_ctu_top_left_x[j], subpic_ctu_top_left_y[j], subpic_width_minus1[j], subpic_height_minus1[j], and loop_filter_across_subpic_enabled_flag[j], respectively, where j is in the range of 0 or greater and less than or equal to sps_num_subpics_minus1.
[0129] Furthermore, SPS may include a syntax element loop_filter_across_subpic_enabled_flag[i] indicating whether in-loop filtering across subpicture boundaries is possible. For example, loop_filter_across_subpic_enabled_flag[i] with a first value (e.g., 0) may indicate that in-loop filtering across the boundary of the i-th subpicture within each encoded picture in the CLVS is not possible. Conversely, loop_filter_across_subpic_enabled_flag[i] with a second value (e.g., 1) may indicate that in-loop filtering across the boundary of the i-th subpicture within each encoded picture in the CLVS is possible. If loop_filter_across_subpic_enabled_flag[i] is not present, its value can be inferred to be the same as 1-sps_independent_subpics_flag. In one example, loop_filter_across_subpic_enabled_flag[i] can only be encoded / signaled if the aforementioned sps_independent_subpics_flag has a first value (e.g., 0) (i.e., the subpicture boundary is not treated as a picture boundary). On the other hand, as a requirement for bitstream consistency, the form of the subpictures must be such that when each subpicture is decoded, the overall left boundary and overall top boundary of each subpicture are composed of the picture boundary or the boundary of a previously decoded subpicture.
[0130] Figure 10 shows a method by which an image encoding device according to one embodiment of the present disclosure encodes an image using a subpicture.
[0131] An image encoding device can encode the current picture based on its subpicture structure. Alternatively, an image encoding device can encode at least one subpicture that constitutes the current picture and output a (sub)bitstream containing (encoded) information about the (encoded) at least one subpicture.
[0132] Referring to Figure 10, the image encoding device can divide the input picture into multiple subpictures (S1010). The image encoding device can then generate information about the subpictures (S1020). Here, the information about the subpictures may include, for example, information about the area of the subpictures and / or information about the grid spacing to be used for the subpictures. The information about the subpictures may also include information about whether each subpicture can be treated as a single picture and / or information about whether in-loop filtering can be performed across the boundaries of each subpicture (boundary).
[0133] The image encoding device can encode at least one subpicture based on information about the subpictures. For example, each subpicture can be encoded independently based on information about the subpicture. The image encoding device can then encode image information containing information about the subpictures and output a bitstream (S1030). Here, the bitstream for the subpictures is sometimes called a substream or subbitstream.
[0134] Figure 11 shows a method by which an image decoding device according to one embodiment of the present disclosure decodes an image using a subpicture.
[0135] The image decoding device can decode at least one subpicture currently contained in the picture using (encoded) information relating to at least one (encoded) subpicture obtained from the (sub)bitstream.
[0136] Referring to Figure 11, the image decoding device can obtain information about subpictures from the bitstream (S1110). Here, the bitstream may include substreams or subbitstreams for subpictures. The information about subpictures can be configured in the higher-level syntax of the bitstream. The image decoding device can then derive at least one subpicture based on the information about subpictures (S1120).
[0137] The image decoding device can decode at least one subpicture based on information about the subpicture (S1130). For example, if the current subpicture containing the current block is treated as a single picture, the current subpicture can be decoded independently. Also, if in-loop filtering can be performed across the boundary of the current subpicture, in-loop filtering (e.g., deblocking filtering) can be performed on the boundary of the current subpicture and the boundary of adjacent subpictures adjacent to the boundary. Also, if the boundary of the current subpicture coincides with the picture boundary, in-loop filtering across the boundary of the current subpicture may not be performed. The image decoding device can decode subpictures based on the CABAC method, prediction method, residual processing method (transformation, quantization), in-loop filtering scheme, etc. The image decoding device can then output at least one decoded subpicture, or output the current picture containing at least one subpicture. The decoded subpictures can be output in the form of an OPS (output sub-picture set). For example, in relation to a 360-degree or omnidirectional image, if only a portion of the picture is currently being rendered, only some of the sub-pictures within the current picture can be decoded, and all or some of the decoded sub-pictures can be rendered according to the user's viewport.
[0138] Wrap-around overview
[0139] When interpretation is applied to the current block, the predicted block of the current block can be guided based on the reference block identified by the motion vector of the current block. In this case, if at least one reference sample within the reference block falls outside the boundary of the reference picture, the sample value of the reference sample can be replaced with the sample value of an adjacent sample located at the boundary or outermost edge of the reference picture. This is called padding, and the boundary of the reference picture can be extended through the padding.
[0140] On the other hand, when the reference picture is obtained from a 360-degree image, a continuity can exist between the left and right boundaries of the reference picture. This allows a sample adjacent to the left (or right) boundary of the reference picture to have the same / similar sample values and / or motion information as a sample adjacent to the right (or left) boundary of the picture. Based on this characteristic, at least one reference sample that falls outside the boundary of the reference picture within the reference block can be replaced by an adjacent sample within the reference picture that corresponds to that reference sample. This is called (horizontal) wrap-around motion compensation, and the motion vector of the current block can be adjusted to point inside the reference picture via the wrap-around motion compensation.
[0141] Wrap-around motion compensation refers to a coding tool designed to improve the visual quality of restored images / videos, such as 360-degree images / videos projected in ERP format. According to existing motion compensation processes, if the motion vector of a current block points to a sample that falls outside the boundary of a reference picture, the sample value of the sample that falls outside the boundary can be induced by copying the sample value of the nearest adjacent sample to the boundary via repeating padding. However, since 360-degree images / videos are captured spherically and inherently have no image boundaries, a reference sample that falls outside the boundary of a reference picture on the projected domain (2D domain) can always be induced from an adjacent sample adjacent to the reference sample on the old domain (3D domain). Therefore, repeating padding is unsuitable for 360-degree images / videos and can induce a visual artifact called a seam artifact in the restored viewport image / video.
[0142] On the other hand, when general projection methods are applied, 2D-to-3D and 3D-to-2D coordinate transformations are performed along with sample interpolation for fractional sample positions, making it difficult to obtain adjacent samples for wrap-around motion compensation on the old domain. However, when the ERP projection method is applied, spherical adjacent samples that fall outside the left (or right) boundary of the reference picture can be obtained relatively easily from samples within the right (or left) boundary of the reference picture. Therefore, considering the relative ease of implementation and widespread use of the ERP projection method, wrap-around motion compensation may be more effective for 360-degree images / videos encoded in the ERP format.
[0143] Figure 12 shows an example of a 360-degree image converted into a 2D picture.
[0144] Referring to Figure 12, the 360-degree image 1210 can be converted into a 2D picture 1230 through a projection process. Depending on the projection method applied to the 360-degree image 1210, the 2D picture 1230 can have various projection formats, such as ERP (equi-rectangular projection) format or PERP (Padded ERP) format.
[0145] The 360-degree image 1210, due to its image characteristics acquired from all directions, does not have an image boundary. However, the 2D picture 1230 acquired from the 360-degree image 1210 has an image boundary due to the projection process. In this case, the left boundary LBd and the right boundary RBd of the 2D picture 1230 may form a single line RL within the 360-degree image 1210 and be in contact with each other. Therefore, the similarity between samples adjacent to the left boundary LBd and the right boundary RBd within the 2D picture 1230 may be relatively high.
[0146] On the other hand, a predetermined region within the 360-degree image 1210 can correspond to either an internal or external region of the 2D picture 1230, depending on the reference image boundary. For example, when the left boundary LBd of the 2D picture 1230 is used as the reference, region A within the 360-degree image 1210 can correspond to region A1, which is located outside the 2D picture 1230. Conversely, when the right boundary RBd of the 2D picture 1230 is used as the reference, region A within the 360-degree image 1210 can correspond to region A2, which is located inside the 2D picture 1230. Regions A1 and A2 can have identical / similar sample attributes to each other, in that they correspond to the same region A with respect to the 360-degree image 1210.
[0147] Based on these characteristics, external samples of a 2D picture 1230 that are outside the left boundary LBd can be replaced by internal samples of the 2D picture 1230 located at a predetermined distance in the first direction DIR1 via wrap-around motion compensation. For example, external samples of a 2D picture 1230 contained in region A1 can be replaced by internal samples of the 2D picture 1230 contained in region A2. Similarly, external samples of a 2D picture 1230 that are outside the right boundary RBd can be replaced by internal samples of the 2D picture 1230 located at a predetermined distance in the second direction DIR2 via wrap-around motion compensation.
[0148] Figure 13 shows an example of a wrap-around motion compensation process.
[0149] Referring to Figure 13, if interpretation is applied to the current block 1310, the predicted block for the current block 1310 can be induced based on the reference block 1330.
[0150] Reference block 1330 can be identified by the motion vector 1320 of current block 1310. In one example, the motion vector 1320 may point to the upper left position of reference block 1330 relative to the upper left position of the same-position block 1315, which is located at the same position as current block 1310 in the reference picture.
[0151] Reference block 1330 may include a first region 1335 that extends beyond the left boundary of the reference picture, as shown in Figure 13. Since the first region 1335 is not currently available for interpretation of block 1310, it can be replaced by a second region 1340 in the reference picture via wrap-around motion compensation. The second region 1340 may correspond to the same region as the first region 1335 on the old domain (3D domain), and the position of the second region 1340 can be determined by adding a wrap-around offset to a given position of the first region 1335 (e.g., upper left position).
[0152] The wraparound offset can be set to the ERP width of the picture before padding. Here, the ERP width can mean the width of the original picture in ERP format obtained from the 360-degree image (i.e., the ERP picture). Horizontal padding can be performed on the left and right borders of the ERP picture. This allows the current picture width (PicWidth) to be determined as the sum of the ERP width, the amount of padding on the left border of the ERP picture, and the amount of padding on the right border of the ERP picture. On the other hand, the wraparound offset can be encoded / signaled using a predetermined syntax element (e.g., pps_ref_wraparound_offset) in the higher-level syntax. This syntax element is not affected by the amount of padding on the left and right borders of the ERP picture, and as a result, asymmetric padding on the original picture can be supported. In other words, the amount of padding on the left boundary of an ERP picture (left padding) and the amount of padding on the right boundary (right padding) can be different from each other.
[0153] The aforementioned information regarding wrap-around motion compensation (e.g., activation status, wrap-around offset, etc.) can be encoded / signaled via higher-level syntax, such as SPS and / or PPS.
[0154] Figure 14a shows an example of SPS including information about wrap-around motion compensation.
[0155] Referring to Figure 14a, SPS can include a syntax element sps_ref_wraparound_enabled_flag that indicates whether or not wrap-around motion compensation is applied at the sequence level. For example, sps_ref_wraparound_enabled_flag with a first value (e.g., 0) can indicate that wrap-around motion compensation is not applied to the current video sequence, including the current block. Conversely, sps_ref_wraparound_enabled_flag with a second value (e.g., 1) can indicate that wrap-around motion compensation is applied to the current video sequence, including the current block. In one example, wrap-around motion compensation for the current video sequence can only be applied if the picture width (e.g., pic_width_in_luma_samples) and CTB width (CtbSizeY) satisfy the following conditions.
[0156] -(Condition):(CtbSizeY / MinCbSizeY+1)≧(pic_width_in_luma_samples / MinCbSizeY-1)
[0157] If the above conditions are not met, for example, if the value of (CtbSizeY / MinCbSizeY+1) is greater than the value of (pic_width_in_luma_samples / MinCbSizeY-1), sps_ref_wraparound_enabled_flag may be restricted to a first value (e.g., 0). Here, CtbSizeY can mean the width or height of the luma component CTB, and MinCbSizeY can mean the minimum width or minimum height of the luma component CB (coding block). Also, pic_width_max_in_luma_samples can mean the maximum width in luma samples for each picture.
[0158] Figure 14b shows an example of PPS including information about wrap-around motion compensation.
[0159] Referring to Figure 14b, PPS can include a syntax element pps_ref_wraparound_enabled_flag that indicates whether or not to apply wraparound motion compensation at the picture level.
[0160] Referring to Figure 14b, PPS can include a syntax element pps_ref_wraparound_enabled_flag that indicates whether wrap-around motion compensation is applied at the sequence level. For example, pps_ref_wraparound_enabled_flag with a first value (e.g., 0) can indicate that wrap-around motion compensation is not applied to the current picture, including the current block. Conversely, pps_ref_wraparound_enabled_flag with a second value (e.g., 1) can indicate that wrap-around motion compensation is applied to the current picture, including the current block. In one example, wrap-around motion compensation for the current picture can only be applied if the picture width (e.g., pic_width_in_luma_samples) is greater than the CTB width (CtbSizeY). For example, if the value of (CtbSizeY / MinCbSizeY+1) is greater than the value of (pic_width_in_luma_samples / MinCbSizeY-1), then pps_ref_wraparound_enabled_flag can be restricted to a first value (e.g., 0). In other examples, if sps_ref_wraparound_enabled_flag has a first value (e.g., 0), then the value of pps_ref_wraparound_enabled_flag can be restricted to a first value (e.g., 0).
[0161] Furthermore, PPS can include the syntax element pps_ref_wraparound_offset, which indicates the offset for wraparound motion compensation. For example, the value of pps_ref_wraparound_offset plus ((CtbSizeY / MinCbSizeY)+2) can indicate the wraparound offset for calculating the wraparound position in luma samples. The value of pps_ref_wraparound_offset can be greater than or equal to 0 and less than or equal to ((pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)-2). On the other hand, the variable PpsRefWraparoundOffset can be set to the same value as (pps_ref_wraparound_offset+(CtbSizeY / MinCbSizeY)+2). The aforementioned variable PpsRefWraparoundOffset can be used in the process of clipping reference samples that fall outside the boundaries of the reference picture.
[0162] On the other hand, if a picture is currently divided into multiple subpictures, wrap-around motion compensation can be performed selectively based on the attributes of each subpicture.
[0163] Figure 15 is a flowchart showing how an image encoding device performs wrap-around motion compensation.
[0164] Referring to Figure 15, the image encoding device can now determine whether or not the subpicture is being coded independently (S1510).
[0165] If the subpicture is currently coded independently (YES in S1510), the image encoding device may decide not to perform wrap-around motion compensation for the current block (S1530). In this case, the image encoding device can clip the reference sample position of the current block with respect to the subpicture boundary and perform motion compensation using the reference sample at the clipped position. This operation can be performed, for example, using a luma sample bilinear interpolation process, a luma sample interpolation filtering process, a luma integer sample fetching process, or a chroma sample interpolation process.
[0166] If subpictures are not currently coded independently (NO in S1510), the image encoder can determine whether wrap-around motion compensation is available (S1520). The image encoder can also determine whether wrap-around motion compensation is available at the sequence level. For example, if all subpictures in a video sequence currently have discontinuous subpicture boundaries, wrap-around motion compensation can be restricted to not be available for the video sequence at the time. If wrap-around motion compensation is not available at the sequence level, it can be enforced that wrap-around motion compensation is not available at the picture level. Conversely, if wrap-around motion compensation is available at the sequence level, the image encoder can determine whether wrap-around motion compensation is available at the picture level.
[0167] If wrap-around motion compensation is available (YES in S1520), the image encoding device can perform wrap-around motion compensation on the current block (S1540). In this case, the image encoding device can perform wrap-around motion compensation by shifting the reference sample position of the current block by the wrap-around offset, and then clipping the shifted position relative to the boundary of the reference picture.
[0168] In contrast, if wrap-around motion compensation is not available (NO in S1520), the image encoding device may choose not to perform wrap-around motion compensation for the current block (S1550). In this case, the image encoding device may clip the reference sample position of the current block relative to the boundary of the reference picture and perform motion compensation using the reference sample at the clipped position.
[0169] On the other hand, the image decoding device can determine whether or not a subpicture is currently coded independently based on subpicture-related information obtained from the bitstream (e.g., subpic_treated_as_pic_flag). For example, if subpic_treated_as_pic_flag has a first value (e.g., 0), the subpicture may not be currently coded independently. Conversely, if subpic_treated_as_pic_flag has a second value (e.g., 1), the subpicture may be currently coded independently.
[0170] The image decoding device then determines whether wraparound motion compensation is available based on wraparound-related information (e.g., pps_ref_wraparound_enabled_flag) obtained from the bitstream, and can perform wraparound motion compensation on the current block based on this determination. For example, if pps_ref_wraparound_enabled_flag has a first value (e.g., 0), the image decoding device may determine that wraparound motion compensation is not available for the current picture and may not perform wraparound motion compensation on the current block. Conversely, if pps_ref_wraparound_enabled_flag has a second value (e.g., 1), the image decoding device may determine that wraparound motion compensation is available for the current picture and can perform wraparound motion compensation on the current block.
[0171] As described above, according to the method in Figure 15, wrap-around motion compensation can only be performed if the subpicture is not currently treated as a single picture. As a result, a problem arises in that wrap-around related coding tools cannot be used together with various subpicture related coding tools that assume independent coding of subpictures. This can act as a factor that degrades encoding / decoding performance for pictures where inter-boundary continuity exists, such as ERP pictures or PERP pictures.
[0172] To solve these problems, according to the embodiments of this disclosure, wrap-around motion compensation can be performed under predetermined conditions even when subpictures are currently coded independently. The embodiments of this disclosure will be described in detail below.
[0173] Example 1
[0174] According to Embodiment 1 of the present disclosure, if there are one or more subpictures in the current video sequence that are coded independently and have a width different from the picture width, wrap-around motion compensation can be forced not to be available for the current video sequence. In this case, flag information indicating whether or not wrap-around motion compensation is available for the current video sequence (e.g., sps_ref_wraparound_enabled_flag) can be restricted to having a first value (e.g., 0).
[0175] Figure 16 is a flowchart showing how an image encoding device according to one embodiment of the present disclosure determines whether wrap-around motion compensation can be used.
[0176] Referring to Figure 16, the image encoding device can determine whether there are currently one or more subpictures in the video sequence that are being coded independently (S1610).
[0177] If the determination determines that there are no subpictures currently coded independently within the video sequence ("NO" in S1610), the image encoding device can determine whether wrap-around motion compensation is available for the current video sequence based on predetermined wrap-around constraints (S1640). In this case, based on the determination, the image encoding device can encode flag information (e.g., sps_wraparound_enabled_flag) indicating whether wrap-around motion compensation is available for the current video sequence into a first value (e.g., 0) or a second value (e.g., 1).
[0178] As an example of the aforementioned wrap-around constraint, if wrap-around motion compensation is restricted to one or more output layer sets (OLSs) identified by a video parameter set (VPS), then wrap-around motion compensation may be restricted to be unavailable for the current video sequence. As another example of the aforementioned wrap-around constraint, if all subpictures in the current video sequence have discontinuous subpicture boundaries, then wrap-around motion compensation may be restricted to be unavailable for the current video sequence.
[0179] In contrast, if there is currently one or more subpictures that are coded independently within the video sequence (YES in S1610), the image encoding device can determine whether at least one of the independently coded subpictures has a width different from the picture width (S1620).
[0180] In one embodiment, the picture width can be derived from Equation 1 based on the maximum width that the picture can currently have in the video sequence.
[0181]
number
[0182] Here, pic_width_max_in_luma_samples represents the maximum width of the picture in luma samples, CtbSizeY represents the width of the CTB (coding tree block) within the picture in luma samples, and CtbLog2SizeY can represent the log-scale value of CtbSizeY.
[0183] If the determination finds that at least one of the independently coded subpictures has a width different from the picture width (YES in S1620), the image encoding device can determine that wrap-around motion compensation is not currently available for the video sequence (S1630). In this case, based on the determination, the image encoding device can encode sps_ref_wraparound_enabled_flag to a first value (e.g., 0).
[0184] In contrast, if all independently coded subpictures have the same width as the picture width (NO in S1620), the image encoding device can determine whether wrap-around motion compensation is currently available for the video sequence based on the wrap-around constraints described above (S1640). In this case, the image encoding device can encode sps_ref_wraparound_enabled_flag to a first value (e.g., 0) or a second value (e.g., 1) based on the determination.
[0185] Although Figure 16 illustrates that steps S1610 and S1620 are performed sequentially, this is merely illustrative and does not limit the embodiments of this disclosure. For example, step S1620 may be performed simultaneously with step S1610, or it may be performed before step S1610.
[0186] On the other hand, the sps_ref_wraparound_enabled_flag encoded by the image encoding device can be stored in the bitstream and signaled to the image decoding device. In this case, the image decoding device can determine whether wrap-around motion compensation is currently available for the video sequence based on the sps_ref_wraparound_enabled_flag obtained from the bitstream.
[0187] For example, if sps_ref_wraparound_enabled_flag has a first value (e.g., 0), the image decoder may determine that wraparound motion compensation is not currently available for the video sequence and may not perform wraparound motion compensation for the current block. In this case, the reference sample position of the current block may be clipped relative to the reference picture boundary or sub-picture boundary, and motion compensation may be performed using the reference sample at the clipped position.
[0188] In other words, the image decoding device can perform correct motion compensation according to this disclosure without having to separately determine whether there are one or more subpictures currently coded independently in the video sequence and having a width different from the picture width. However, the operation of the image decoding device is not limited to this. For example, the image decoding device can determine whether there are one or more subpictures currently coded independently in the video sequence and having a width different from the picture width, and then perform motion compensation based on that determination. More specifically, the image decoding device can determine whether there are one or more subpictures currently decoded independently in the video sequence and having a width different from the picture width, and if such subpictures exist, it can consider sps_ref_wraparound_enabled_flag to a first value (e.g., 0) and not perform wrap-around motion compensation.
[0189] In contrast, if sps_ref_wraparound_enabled_flag has a second value (e.g., 1), the image decoder can determine that wrap-around motion compensation is currently available for the video sequence. In this case, the image decoder can obtain additional flag information (e.g., pps_ref_wraparound_enabled_flag) from the bitstream indicating whether wrap-around motion compensation is currently available for the picture, and based on the obtained flag information, can determine whether or not to perform wrap-around motion compensation for the current block.
[0190] For example, if pps_ref_wraparound_enabled_flag has a first value (e.g., 0), the image decoder may choose not to perform wrap-around motion compensation for the current block. In this case, the reference sample position of the current block may be clipped relative to the reference picture boundary or sub-picture boundary, and motion compensation may be performed using the reference sample at the clipped position. Conversely, if pps_ref_wraparound_enabled_flag has a second value (e.g., 1), the image decoder may perform wrap-around motion compensation for the current block.
[0191] As described above, according to Embodiment 1 of this disclosure, if there is one or more subpictures in the current video sequence that are independently coded and have a width different from the picture width, wrap-around motion compensation can be forced not to be available for the current video sequence. This implies that if all independently coded subpictures in the current video sequence have the same width as the picture width, wrap-around motion compensation can be applied to each subpicture in the current video sequence regardless of whether they are independently coded or not. As a result, the coding tools related to subpictures and the coding tools related to wrap-around motion compensation can be used together, thereby improving the efficiency of encoding / decoding.
[0192] Example 2
[0193] According to Embodiment 2 of the present disclosure, if there are one or more subpictures in the current video sequence that have a width different from the picture width, wrap-around motion compensation can be forced not to be available for the current video sequence. In this case, flag information indicating whether or not wrap-around motion compensation is available for the current video sequence (e.g., sps_ref_wraparound_enabled_flag) can be restricted to having a first value (e.g., 0).
[0194] Figure 17 is a flowchart showing how an image encoding device according to one embodiment of the present disclosure determines whether wrap-around motion compensation can be used.
[0195] Referring to Figure 17, the image encoding device can determine whether there are currently one or more sub-pictures in the video sequence that have a width different from the picture width (S1710).
[0196] If the determination indicates that there is currently one or more sub-pictures in the video sequence with a width different from the picture width (YES in S1710), the image encoding device can determine that wrap-around motion compensation is not currently available for the video sequence (S1720). In this case, the image encoding device can encode sps_ref_wraparound_enabled_flag to a first value (e.g., 0).
[0197] In contrast, if there are currently no sub-pictures in the video sequence with a width different from the picture width (i.e., all sub-pictures have a width equal to the picture width) ("NO" in S1710), the image encoding device can determine whether wrap-around motion compensation is currently available for the video sequence based on predetermined wrap-around constraints (S1730). An example of the wrap-around constraints is as described above with reference to Figure 16. In this case, the image encoding device can encode sps_ref_wraparound_enabled_flag to a first value (e.g., 0) or a second value (e.g., 1) based on the determination.
[0198] On the other hand, the sps_ref_wraparound_enabled_flag encoded by the image encoding device can be stored in the bitstream and signaled to the image decoding device. In this case, the image decoding device can determine whether wrap-around motion compensation is currently available for the video sequence based on the sps_ref_wraparound_enabled_flag obtained from the bitstream.
[0199] For example, if sps_ref_wraparound_enabled_flag has a first value (e.g., 0), the image decoder can determine that wrap-around motion compensation is not currently available for the video sequence and may not perform wrap-around motion compensation for the current block. Conversely, if sps_ref_wraparound_enabled_flag has a second value (e.g., 1), the image decoder can determine that wrap-around motion compensation is currently available for the video sequence. In this case, the image decoder can obtain additional flag information (e.g., pps_ref_wraparound_enabled_flag) from the bitstream indicating whether wrap-around motion compensation is currently available for the picture, and determine whether to perform wrap-around motion compensation for the current block based on the obtained flag information.
[0200] As described above, according to Embodiment 2 of this disclosure, if there is one or more subpictures in the current video sequence that have a width different from the picture width, wrap-around motion compensation can be forced not to be available for the current video sequence. This implies that if all subpictures in the current video sequence have the same width as the picture width, wrap-around motion compensation can be applied to each subpicture in the current video sequence regardless of whether it is independently coded. This allows the subpicture-related coding tool and the wrap-around motion compensation-related coding tool to be used together, thereby improving the efficiency of encoding / decoding.
[0201] Example 3
[0202] According to Embodiment 3 of the present disclosure, wrap-around motion compensation for the current block can be performed if the current subpicture has the same width as the current picture even if it is coded independently, or if the current subpicture is not coded independently.
[0203] Figure 18 is a flowchart showing a method by which an image encoding device according to one embodiment of the present disclosure performs wrap-around motion compensation.
[0204] Referring to Figure 18, the image encoding device can now determine whether or not the subpicture is being coded independently (S1810).
[0205] If the subpicture is currently coded independently (YES in S1810), the image encoding device can determine whether the width of the subpicture is equal to the width of the picture (S1820).
[0206] If the width of the subpicture is currently equal to the width of the current picture (YES in S1820), the image encoding device can determine whether wrap-around motion compensation is available (S1830).
[0207] For example, if at least one of the independently coded subpictures in the current video sequence has a width different from the picture width, the image encoder can determine that wrap-around motion compensation is not available. In this case, the image encoder can encode a flag (e.g., sps_ref_wraparound_enabled_flag) indicating whether or not wrap-around motion compensation is available for the current video sequence into a first value (e.g., 0) based on the determination.
[0208] In contrast, if all independently coded subpictures in a video sequence currently have the same width as the picture width, the image encoding device can determine whether wrap-around motion compensation is available based on predetermined wrap-around constraints.
[0209] As an example of the wrap-around constraint, if the CTB width (e.g., CtbSizeY) is greater than the picture width (e.g., pic_width_in_luma_samples), wrap-around motion compensation may be restricted from being available for the current picture. As another example of the wrap-around constraint, if wrap-around motion compensation is restricted for one or more output layer sets (OLSs) identified by the video parameter set (VPS), wrap-around motion compensation may be restricted from being available for the current picture. As yet another example of the wrap-around constraint, if all subpictures within the current picture have discontinuous subpicture boundaries, wrap-around motion compensation may be restricted from being available for the current picture.
[0210] If the determination indicates that wrap-around motion compensation is available (YES in S1830), the image encoding device can perform wrap-around motion compensation for the current block based on the boundary of the current sub-picture (S1840). For example, if the reference sample position of the current block is outside the left boundary of the reference picture (e.g., xInti<0), the image encoding device can perform wrap-around motion compensation by shifting the x-coordinate of the reference sample in the positive direction by the wrap-around offset (e.g., PpsRefWraparoundOffset×MinCbSizeY) and then clipping it based on the left and right boundaries of the current sub-picture. Alternatively, if the reference sample position of the current block is outside the right boundary of the reference picture (e.g., xInti>picW-1), the image encoding device can perform wrap-around motion compensation by shifting the x-coordinate of the reference sample in the negative direction by the wrap-around offset and then clipping it based on the left and right boundaries of the current sub-picture.
[0211] In contrast, if wrap-around motion compensation is not available ("NO" in S1830), the image encoding device may choose not to perform wrap-around motion compensation for the current block (S1850). In this case, the image encoding device may clip the reference sample position of the current block relative to the boundary of the current subpicture, and perform motion compensation using the reference sample at the clipped position.
[0212] Returning to step S1810, if the subpicture is not currently coded independently (NO in S1810), the image coding device can determine whether wrap-around motion compensation is available (S1860). The specific method for making this determination is as described above in step S1830.
[0213] If the determination indicates that wrap-around motion compensation is available (YES in S1860), the image encoding device can perform wrap-around motion compensation on the current block based on the boundary of the reference picture (S1870). In this case, the image encoding device can perform wrap-around motion compensation by shifting the reference sample position of the current block by the wrap-around offset, and then clipping the shifted position based on the boundary of the reference picture.
[0214] In contrast, if wrap-around motion compensation is not available ("NO" in S1860), the image encoding device may choose not to perform wrap-around motion compensation for the current block (S1880). In this case, the image encoding device may clip the reference sample position of the current block relative to the boundary of the reference picture, and perform motion compensation using the reference sample at the clipped position.
[0215] On the other hand, the image decoding device can determine whether or not a subpicture is currently coded independently based on subpicture-related information obtained from the bitstream (e.g., subpic_treated_as_pic_flag). For example, if subpic_treated_as_pic_flag has a first value (e.g., 0), the subpicture may not be coded independently. Conversely, if subpic_treated_as_pic_flag has a second value (e.g., 1), the subpicture may be coded independently.
[0216] The image decoding device then determines whether wraparound motion compensation is available based on wraparound-related information (e.g., pps_ref_wraparound_enabled_flag) obtained from the bitstream, and can perform wraparound motion compensation on the current block based on this determination. For example, if pps_ref_wraparound_enabled_flag has a first value (e.g., 0), the image decoding device may determine that wraparound motion compensation is not available for the current picture and may not perform wraparound motion compensation on the current block. Conversely, if pps_ref_wraparound_enabled_flag has a second value (e.g., 1), the image decoding device may determine that wraparound motion compensation is available for the current picture and can perform wraparound motion compensation on the current block.
[0217] The image decoding device can perform correct motion compensation according to this disclosure without having to separately determine whether the width of the current subpicture is equal to the width of the current picture. However, the operation of the image decoding device is not limited to this. For example, as shown in Figure 18, the image decoding device can determine whether the width of the independently coded current subpicture is equal to the width of the current picture, and then, based on the result of that determination, determine whether wrap-around motion compensation is available.
[0218] As described above, according to Embodiment 3 of this disclosure, if the subpicture currently has the same width as the current picture even when coded independently, or if the subpicture currently is not coded independently, wrap-around motion compensation for the current block can be performed. This allows the subpicture-related coding tool and the wrap-around motion compensation-related coding tool to be used together, thereby improving the efficiency of encoding / decoding.
[0219] Hereinafter, an image encoding / decoding method according to one embodiment of the present disclosure will be described in detail with reference to Figures 19 and 20.
[0220] Figure 19 is a flowchart showing an image encoding method according to one embodiment of the present disclosure.
[0221] The image encoding method shown in Figure 19 can be performed by the image encoding device shown in Figure 2. For example, steps S1910 and S1920 can be performed by the interpretation unit 180, and step S1930 can be performed by the entropy encoding unit 190.
[0222] Referring to Figure 19, the image encoding device can decide whether or not to apply wrap-around motion compensation to the current block (S1910).
[0223] In one embodiment, the image encoding device can determine whether wrap-around motion compensation is available based on whether there are one or more sub-pictures that are independently coded within the current video sequence, including the current block, and have a width different from the picture width. For example, if there are one or more sub-pictures that are independently coded within the current video sequence and have a width different from the picture width, the image encoding device can determine that wrap-around motion compensation is not available. Conversely, if all independently coded sub-pictures within the current video sequence have the same width as the picture width, the image encoding device can determine that wrap-around motion compensation is available based on predetermined wrap-around constraints. Here, an example of the wrap-around constraints is as described above with reference to Figures 16 to 18.
[0224] Based on the above decision, the image encoding device can then decide whether or not to apply wrap-around motion compensation to the current block. For example, if wrap-around motion compensation is not available for the current picture, the image encoding device can decide not to apply wrap-around motion compensation to the current block. Conversely, if wrap-around motion compensation is available for the current picture, the image encoding device can decide to apply wrap-around motion compensation to the current block.
[0225] The image encoding device can generate a predicted block for the current block by performing interpretation based on the decision result of step S1910 (S1920). For example, when applying wrap-around motion compensation to the current block, the image encoding device can shift the reference sample position of the current block by the wrap-around offset and clip the shifted position relative to the reference picture boundary or sub-picture boundary. Then, the image encoding device can generate a predicted block for the current block by performing motion compensation using the reference sample at the clipped position. Conversely, when not applying wrap-around motion compensation to the current block, the image encoding device can generate a predicted block for the current block by clipping the reference sample position of the current block relative to the reference picture boundary or sub-picture boundary and then performing motion compensation using the reference sample at the clipped position.
[0226] The image encoding device can currently encode interprediction information and wrap-around information related to wrap-around motion compensation for the block to generate a bitstream (S1930).
[0227] In one embodiment, the wraparound information may include a first flag (e.g., sps_ref_wraparound_enabled_flag) indicating whether wraparound motion compensation is available for the video sequence containing the current block. The first flag may be independently coded and have a first value (e.g., 0) indicating that wraparound motion compensation is not available for the current video sequence, based on the presence of one or more subpictures in the current video sequence having widths different from the width of the current picture containing the current block. In this case, the width of the current picture may be derived based on information about the maximum width that a picture can have in the current video sequence (e.g., pic_width_max_in_luma_samples), as described above with reference to Equation 1.
[0228] In one embodiment, the wraparound information may further include a second flag (e.g., pps_ref_wraparound_enabled_flag) indicating whether wraparound motion compensation is currently available for the picture. The second flag may have a first value (e.g., 0) indicating that wraparound motion compensation is not currently available for the picture, based on the first flag having a first value (e.g., 0). The second flag may also have a first value (e.g., 0) indicating that wraparound motion compensation is not available for the current picture, based on predetermined conditions relating to the width of the coding tree block (CTB) in the current picture and the width of the current picture. For example, if the width of the CTB in the current picture (e.g., CtbSizeY) is greater than the width of the picture (e.g., pic_width_in_luma_samples), sps_ref_wraparound_enabled_flag may be limited to a first value (e.g., 0).
[0229] In one embodiment, the wraparound information may further include a wraparound offset (e.g., pps_ref_wraparound_offset) based on the fact that wraparound motion compensation is currently available for the picture. The image encoding device can perform wraparound motion compensation based on the wraparound offset.
[0230] Figure 20 is a flowchart showing an image decoding method according to one embodiment of the present disclosure.
[0231] The image decoding method shown in Figure 20 can be performed by the image decoding device shown in Figure 3. For example, steps S2010 and S2020 can be performed by the interpretation unit 260.
[0232] Referring to Figure 20, the image decoding device can obtain interprediction information and wrap-around information for the current block from the bitstream (S2010). Here, the interprediction information for the current block may include motion information for the current block, such as a reference picture index and differential motion vector information. The wrap-around information is information regarding wrap-around motion compensation and may include a first flag (e.g., sps_ref_wraparound_enabled_flag) indicating whether wrap-around motion compensation is available for the current video sequence containing the current block. The first flag may have a first value (e.g., 0) that is independently coded and indicates that wrap-around motion compensation is not available based on the presence of one or more subpictures in the current video sequence having widths different from the width of the current picture containing the current block. In this case, the width of the current picture may be derived based on information regarding the maximum width that a picture can have in the current video sequence (e.g., pic_width_max_in_luma_samples), as described above with reference to Equation 1.
[0233] In one embodiment, the wraparound information may further include a second flag (e.g., pps_ref_wraparound_enabled_flag) indicating whether wraparound motion compensation is available for the current picture. The second flag may have a first value (e.g., 0) indicating that wraparound motion compensation is not available for the current picture, based on the first flag having a first value (e.g., 0). The second flag may also have a first value (e.g., 0) indicating that wraparound motion compensation is not available for the current picture, based on predetermined conditions relating to the width of the coding tree block (CTB) in the current picture and the width of the current picture. For example, if the width of the CTB in the current picture (e.g., CtbSizeY) is greater than the width of the picture (e.g., pic_width_in_luma_samples), sps_ref_wraparound_enabled_flag may be limited to a first value (e.g., 0).
[0234] The image decoding device can generate a predicted block for the current block based on inter-prediction information and wrap-around information obtained from the bitstream (S2020).
[0235] For example, if wrap-around motion compensation is available for the current picture (e.g., pps_ref_wraparound_enabled_flag==1), wrap-around motion compensation can be performed on the current block. In this case, the image decoder can shift the reference sample position of the current block by the wrap-around offset and clip the shifted position relative to the reference picture boundary or sub-picture boundary. The image decoder can then generate a predicted block for the current block by performing motion compensation using the reference sample at the clipped position.
[0236] In contrast, if wraparound motion compensation is not currently available for the picture (e.g., pps_ref_wraparound_enabled_flag==0), the image encoder can generate a predicted block for the current block by clipping the reference sample position of the current block relative to the reference picture boundary or subpicture boundary, and then performing motion compensation using the reference sample at the clipped position.
[0237] As described above, according to the image encoding / decoding method of one embodiment of the present disclosure, if all independently coded subpictures in the current video sequence have the same width as the picture width, then wrap-around motion compensation can be used for all subpictures in the current video sequence. This allows the subpicture-related coding tool and the wrap-around motion compensation-related coding tool to be used together, thereby improving the efficiency of encoding / decoding.
[0238] The names of the syntax elements described in this disclosure may include information about the location where the syntax element is signaled. For example, a syntax element beginning with "sps_" may mean that the syntax element is signaled in a sequence parameter set (SPS). Furthermore, syntax elements beginning with "pps_", "ph_", "sh_", etc., may mean that the syntax element is signaled in a picture parameter set (PPS), picture header, slice header, etc., respectively.
[0239] The exemplary methods in this disclosure are presented as a series of actions for clarity of explanation, but this is not intended to restrict the order in which the steps are performed, and each step may be performed simultaneously or in a different order, if necessary. To implement the methods according to this disclosure, the exemplary steps may be further expanded to include other steps, or some steps may be expanded to include the remaining steps, or some steps may be expanded to include additional steps.
[0240] In this disclosure, an image encoding device or image decoding device that performs a predetermined operation (step) may perform an operation (step) to confirm the conditions or status of the execution of said operation (step). For example, if it is stated that a predetermined operation is performed when a predetermined condition is satisfied, the image encoding device or image decoding device may perform an operation to confirm whether or not the predetermined condition is satisfied, and then perform the predetermined operation.
[0241] The various embodiments of this disclosure are not intended to list all possible combinations, but rather to illustrate representative aspects of this disclosure. The matters described in the various embodiments may be applied independently or in combination of two or more.
[0242] Furthermore, various embodiments of this disclosure can be implemented by hardware, firmware, software, or a combination thereof. In the case of hardware implementation, it can be implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), general processors, controllers, microcontrollers, microprocessors, etc.
[0243] Furthermore, the image decoding and image encoding devices to which the embodiments of this disclosure are applied can be included in multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video conferencing equipment, real-time communication equipment such as video communications, mobile streaming equipment, storage media, camcorders, video-on-demand (VoD) service providers, over-the-top (OTT) video equipment, internet streaming service providers, 3D video equipment, image-phone video equipment, and medical video equipment, and can be used to process video signals or data signals. For example, over-the-top (OTT) video equipment can include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).
[0244] Figure 21 illustrates a content streaming system to which the embodiments of this disclosure can be applied.
[0245] As shown in Figure 21, a content streaming system to which an embodiment of the present disclosure is applied may broadly include an encoding server, a streaming server, a web server, media storage, user equipment, and multimedia input devices.
[0246] The encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and transmitting this bitstream to the streaming server. In other cases, if a multimedia input device such as a smartphone, camera, or video camera directly generates the bitstream, the encoding server can be omitted.
[0247] The bitstream can be generated by an image encoding method and / or image encoding apparatus to which an embodiment of the present disclosure is applied, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0248] The streaming server transmits multimedia data to the user's device based on the user's request via a web server, and the web server can act as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server can transmit multimedia data to the user. In this case, the content streaming system may include a separate control server, in which case the control server can play a role in controlling the commands and responses between the devices within the content streaming system.
[0249] The streaming server can receive content from media storage and / or encoding servers. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0250] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices such as smartwatches, smart glasses, HMDs (head-mounted displays), digital TVs, desktop computers, and digital signage.
[0251] Each server within the aforementioned content streaming system can be operated as a distributed server, in which case the data received from each server can be processed in a distributed manner.
[0252] Figure 22 is a schematic diagram illustrating an architecture for providing a 3D image / video service that can be utilized in the embodiments of this disclosure.
[0253] Figure 22 can illustrate a 360-degree or omnidirectional video / image processing system. Furthermore, the system in Figure 22 can be implemented, for example, with an Extended Reality (XR) assisted device. That is, the system can provide a means of delivering virtual reality to the user.
[0254] Extended reality is a general term encompassing virtual reality (VR), augmented reality (AR), and mixed reality (MR). VR technology provides real-world objects and backgrounds solely as computer graphics (CG) images, AR technology provides virtual CG images on top of real-world images, and MR technology is a computer graphics technology that blends and combines virtual objects with the real world.
[0255] Mixed Reality (MR) technology is similar to augmented reality (AR) technology in that it displays real and virtual objects together. However, there is a difference in that in AR technology, virtual objects are used to complement real objects, whereas in MR technology, virtual and real objects are used with equal importance.
[0256] XR technology can be applied to HMDs (Head-Mount Displays), HUDs (Head-Up Displays), mobile phones, tablet PCs, laptop computers, desktops, TVs, digital signage, etc., and devices to which XR technology is applied can be called XR devices. An XR device may include the first digital device and / or second digital device described later.
[0257] 360-degree content refers to all content used to realize and deliver VR, and may include 360-degree video and / or 360-degree audio. 360-degree video can mean video or image content that is captured or played back simultaneously in all directions (360 degrees or less) necessary to deliver VR. Hereinafter, 360-degree video can mean 360-degree video. 360-degree audio can also mean audio content used to deliver VR, and may mean spatial audio content where the sound source can be perceived as being located in a specific three-dimensional space. 360-degree content can be generated, processed, and transmitted to users, who can consume VR experiences using 360-degree content. 360-degree video is sometimes called omnidirectional video, and 360-degree images are sometimes called omnidirectional images. Furthermore, the following explanation is based on 360-degree video, and the examples in this document are not limited to VR but may include processing of video / image content such as AR and MR. 360-degree video can refer to a video or image displayed on a 3D space in various forms depending on the 3D model. For example, 360-degree video can be displayed on a spherical surface.
[0258] This method proposes a particularly effective way to provide 360-degree video. To provide 360-degree video, first, 360-degree video can be captured via one or more cameras. The captured 360-degree video is transmitted through a series of processes, and the receiving end can process and render the received data back into the original 360-degree video. This allows the 360-degree video to be provided to the user.
[0259] Specifically, the overall process for providing 360-degree video may include a capture process, a preparation process, a transmission process, a processing process, a rendering process, and / or a feedback process.
[0260] The capture process can refer to the process of capturing images or videos for each of multiple viewpoints via one or more cameras. The capture process can generate image / video data like that shown in Figure 22, 2210. Each plane in Figure 22, 2210 can represent an image / video for each viewpoint. These captured images / videos can also be called raw data. Metadata related to the capture can be generated during the capture process.
[0261] For this capture, a special camera for VR can be used. In some embodiments, when attempting to provide 360-degree video of a computer-generated virtual space, capture via an actual camera may not be performed. In this case, the capture process can be replaced simply by the process of generating the relevant data.
[0262] The preparation process may involve processing the captured images / videos and metadata generated during the capture process. The captured images / videos may undergo processes such as stitching, projection, region-wise packing, and / or encoding during the preparation process.
[0263] First, each image / video can undergo a stitching process. The stitching process can involve combining the captured images / videos to create a single panoramic or spherical image / video.
[0264] Subsequently, the stitched images / videos can undergo a projection process. During the projection process, the stitched images / videos can be projected onto a 2D image. This 2D image is sometimes referred to as a 2D image frame, depending on the context. Projecting as a 2D image can also be described as mapping to a 2D image. The projected image / video data can take the form of a 2D image, such as 2220 in Figure 22.
[0265] Video data projected onto a 2D image can undergo region-wise packing to improve video coding efficiency. Region-wise packing refers to the process of dividing the video data projected onto the 2D image into regions for processing. Here, a region refers to an area of the 2D image onto which 360-degree video data is projected. Depending on the implementation, these regions can be divided evenly across the 2D image or arbitrarily. Depending on the implementation, these regions can also be distinguished by the projection scheme. Region-wise packing is an optional process and can be omitted during the preparation process.
[0266] Depending on the embodiment, this processing step may include rotating or rearranging each region on a 2D image to improve video coding efficiency. For example, by rotating these regions so that certain edges of the regions are located close to each other, coding efficiency can be increased.
[0267] Depending on the embodiment, this processing may include increasing or decreasing the resolution for specific regions in order to differentiate the resolution across different areas of the 360-degree video. For example, regions that are relatively more important in the 360-degree video may have a higher resolution than other regions. Video data projected onto a 2D image or video data packed region by region can undergo an encoding process via a video codec.
[0268] Depending on the embodiment, the preparation process may further include an editing process. This editing process may involve further editing of the image / video data before and after projection. Similarly, metadata related to stitching, projection, encoding, and editing can be generated during the preparation process. Furthermore, metadata related to the initial viewpoint or ROI (Region of Interest) of the video data projected onto the 2D image can also be generated.
[0269] The transmission process may involve processing and transmitting image / video data and metadata that have undergone a preparation process. Processing can be performed using any transmission protocol for transmission. The processed data can be transmitted via broadcast networks and / or broadband. This data can also be transmitted to the recipient on demand. The recipient can receive the data via various routes.
[0270] The processing process can be defined as the process of decoding the received data and reprojecting the projected image / video data onto a 3D model. In this process, image / video data projected onto a 2D image can be reprojected onto a 3D space. This process may also be called mapping or projection depending on the context. At this time, the 3D space to which it is mapped can have different forms depending on the 3D model. For example, a 3D model could be a sphere, cube, cylinder, or pyramid.
[0271] Depending on the embodiment, the processing steps may further include editing and upscaling processes. The editing process may involve further editing of the image / video data before and after reprojection. If the image / video data is reduced in size, the upscaling process can enlarge its size through sample upscaling. If necessary, downscaling can also be performed to reduce the size.
[0272] The rendering process can refer to the process of rendering and displaying image / video data that has been reprojected onto a 3D space. In some expressions, the combination of reprojection and rendering can be described as rendering onto a 3D model. An image / video reprojected onto (or rendered onto) a 3D model can have a form like 2230 in Figure 22. Figure 2230 shows the case where it is reprojected onto a spherical 3D model. The user can view a portion of the rendered image / video via a VR display or the like. In this case, the area viewed by the user may have a form like 2240 in Figure 22.
[0273] The feedback process may refer to a process of transmitting various feedback information that can be obtained in the display process to the transmitting side. Interactivity can be provided in 360-degree video consumption via the feedback process. According to some embodiments, head orientation information, viewport information indicating the area currently viewed by the user, and the like can be transmitted to the transmitting side in the feedback process. According to some embodiments, a user can interact with content implemented in a VR environment, and in this case, information related to the interaction can also be transmitted to the transmitting side or the service provider side in the feedback process. According to some embodiments, the feedback process may be omitted.
[0274] Head orientation information may refer to information about the position, angle, movement and the like of a user's head. Based on this information, information about the area that the user is currently viewing in the 360-degree video, that is, viewport information can be calculated.
[0275] Viewport information may be information about the area that a user is currently viewing in a 360-degree video. Gaze analysis performed based on the viewport information allows confirmation of how a user consumes 360-degree video and which area of 360-degree video the user gazes at and for how long. Gaze analysis may be performed at the receiving side and then transmitted to the transmitting side via a feedback channel. Devices such as VR displays can extract the viewport area based on the position / orientation of the user's head, vertical or horizontal field of view (FOV) information supported by the device, and the like.
[0276] On the other hand, 360-degree video / images can be processed based on subpictures. A projected picture or packed picture including a 2D image can be divided into subpictures, and processing can be performed in units of subpictures. For example, a higher resolution can be provided for a specific subpicture according to a user viewport, or only a specific subpicture can be encoded and signaled to a receiving apparatus (decoding apparatus side). In this case, the decoding apparatus can receive the subpicture bitstream, restore / decode the specific subpicture, and perform rendering according to the user viewport.
[0277] According to an embodiment, the above-described feedback information can be not only transmitted to the transmitting side but also consumed at the receiving side. That is, decoding, reprojection, rendering and other processes at the receiving side can be performed using the above-described feedback information. For example, using head orientation information and / or viewport information, only the 360-degree video corresponding to the area currently viewed by a user can be preferentially decoded and rendered.
[0278] Here, a viewport or viewport area may refer to an area that a user is viewing in 360-degree video. A viewpoint is a point that the user is viewing in 360-degree video, and may refer to the center point of the viewport area. That is, the viewport is an area centered on the viewpoint, and the size, shape and the like occupied by the area can be determined by FOV (Field Of View).
[0279] In the overall architecture for providing the aforementioned 360-degree video, image / video data that undergoes a series of processes of capture / projection / encoding / transmission / decoding / reprojection / rendering can be referred to as 360-degree video data. Also, the term 360-degree video data is sometimes used as a concept including metadata or signaling information related to such image / video data.
[0280] A standardized media file format can be defined for storing and transmitting the aforementioned audio or video media data. In some embodiments, the media file may have a file format based on ISO BMFF (ISO base media file format).
[0281] The scope of this disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) that enable the operation of various embodiments to be performed on a device or computer, and non-transitory computer-readable medium on which such software or commands etc. are stored and can be executed on a device or computer. [Industrial applicability]
[0282] The embodiments described herein can be used for encoding / decoding images.
Claims
1. An image decoding method performed by an image decoding device, Steps include obtaining interpretation information and wrap-around information for the current block within the current picture from the bitstream. The step of generating a predicted block for the current block based on the inter prediction information and the wrap-around information, The aforementioned wrap-around information includes a first flag and a second flag, The first flag indicates whether wrap-around motion compensation is available for the current video sequence, which includes the current picture. The second flag indicates whether the wrap-around motion compensation is available for the current picture. Based on the fact that the current video sequence includes at least one sub-picture that is treated as a picture and has a width different from a value derived based on information regarding the maximum width of pictures in the current video sequence, the first flag obtained from the bitstream is constrained to have a first value indicating that the wrap-around motion compensation is not available for the current video sequence. Image decoding method wherein the second flag has a first value indicating that wrap-around motion compensation is not available for the current picture, based on predetermined conditions relating to the width of the coding tree block (CTB) in the current picture and the width of the current picture.
2. An image encoding method performed by an image encoding device, The steps include determining whether wrap-around motion compensation is applied to the current block within the current picture, The steps include generating a predicted block for the current block by performing an inter-prediction based on the aforementioned decision, The step includes encoding the interprediction information of the current block and the wrap-around information of the wrap-around motion compensation, The aforementioned wrap-around information includes a first flag and a second flag, The first flag indicates whether wrap-around motion compensation is available for the current video sequence, which includes the current picture. The second flag indicates whether the wrap-around motion compensation is available for the current picture. Based on the fact that the current video sequence includes at least one sub-picture which is treated as a picture and has a width different from a value derived based on information regarding the maximum width of pictures in the current video sequence, the encoded first flag has a first value indicating that the wrap-around motion compensation is not available for the current video sequence. An image encoding method wherein the encoded second flag has a first value indicating that wrap-around motion compensation is not available for the current picture, based on predetermined conditions relating to the width of the coding tree block (CTB) in the current picture and the width of the current picture.
3. A method for transmitting a bitstream generated by an image encoding method, The aforementioned image encoding method is The steps include determining whether wrap-around motion compensation is applied to the current block within the current picture, The steps include generating a predicted block for the current block by performing an inter-prediction based on the aforementioned decision, The step includes encoding the interprediction information of the current block and the wrap-around information of the wrap-around motion compensation, The aforementioned wrap-around information includes a first flag and a second flag, The first flag indicates whether wrap-around motion compensation is available for the current video sequence, which includes the current picture. The second flag indicates whether the wrap-around motion compensation is available for the current picture. Based on the fact that the current video sequence includes at least one sub-picture which is treated as a picture and has a width different from a value derived based on information regarding the maximum width of pictures in the current video sequence, the encoded first flag has a first value indicating that the wrap-around motion compensation is not available for the current video sequence. The method wherein the encoded second flag has a first value indicating that wrap-around motion compensation is not available for the current picture, based on predetermined conditions relating to the width of the coding tree block (CTB) in the current picture and the width of the current picture.
Citation Information
Patent Citations
Sample derivation for 360-degree video coding
WO2020069058A1
Methods for performing wrap-around motion compensation
WO2021127118A1
Wraparound offsets for reference picture resampling in video coding
WO2021133979A1