Image encoding / decoding method and apparatus based on winding motion compensation, and recording medium storing bitstream
Patent Information
- Application Number
- CN202180037615.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-14
- Filing Date
- 2021-03-26
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2041-03-26
AI Technical Summary
传输信息量或比特量的增加导致传输成本和存储成本的增加
Smart Images

Figure CN115699755B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to image encoding / decoding methods and apparatus, and more specifically, to an image encoding and decoding method and apparatus based on winding motion compensation, and a recording medium for storing bitstreams generated by the image encoding method / apparatus of this disclosure. Background Technology
[0002] Recently, there has been an increasing demand for high-resolution and high-quality images, such as high-definition (HD) and ultra-high-definition (UHD) images, across various fields. With the improvement in image data resolution and quality, the amount of information or bits transmitted increases relative to existing image data. This increase in the amount of information or bits transmitted leads to increased transmission and storage costs.
[0003] Therefore, efficient image compression techniques are needed to effectively transmit, store, and reproduce information about high-resolution and high-quality images. Summary of the Invention
[0004] Technical issues
[0005] The purpose of this disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0006] Another object of this disclosure is to provide an image encoding / decoding method and apparatus based on winding motion compensation.
[0007] Another object of this disclosure is to provide an image encoding / decoding method and apparatus based on scrolling motion compensation for independently encoded sub-pictures.
[0008] Another object of this disclosure is to provide a method for transmitting a bitstream generated by an image encoding method or device according to this disclosure.
[0009] Another object of this disclosure is to provide a recording medium for storing bitstreams generated by an image encoding method or device according to this disclosure.
[0010] Another object of this disclosure is to provide a recording medium that stores a bitstream received, decoded and used to reconstruct an image by an image decoding device according to this disclosure.
[0011] The technical problems solved by this disclosure are not limited to those described above. Other technical problems not described herein will become clear to those skilled in the art through the following description.
[0012] Technical solution
[0013] An image decoding method performed by an image decoding device according to one aspect of this disclosure includes the following steps: obtaining inter-frame prediction information and winding information of a current block from a bitstream; and generating a prediction block of the current block based on the inter-frame prediction information and winding information. The winding information may include a first flag specifying whether winding motion compensation is enabled for a current frame including the current block. Based on the first flag having a predetermined value specifying whether winding motion compensation is enabled for the current frame, a prediction block can be generated by performing winding motion compensation, and based on whether a current sub-frame including the current block is independently encoded, winding motion compensation can be performed based on the boundary of the current sub-frame or the boundary of a reference frame of the current block.
[0014] An image decoding apparatus according to another aspect of this disclosure includes a memory and at least one processor. The at least one processor can be configured to obtain inter-frame prediction information and winding information of a current block from a bitstream and to generate a prediction block of the current block based on the inter-frame prediction information and winding information. The winding information may include a first flag specifying whether winding motion compensation is enabled for a current frame including the current block. Based on the first flag having a predetermined value specifying that winding motion compensation is enabled for the current frame, a prediction block can be generated by performing winding motion compensation, and based on whether the current sub-frame including the current block is independently encoded, winding motion compensation can be performed based on the boundary of the current sub-frame or the boundary of a reference frame of the current block.
[0015] An image coding method according to another aspect of this disclosure includes the following steps: determining whether rollover motion compensation is applied to the current block; generating a prediction block for the current block by performing inter-frame prediction based on whether rollover motion compensation is applied to the current block; and encoding inter-frame prediction information for the current block and rollover information for rollover motion compensation. The rollover information may include a first flag specifying whether rollover motion compensation is enabled for the current frame including the current block, and rollover motion compensation may be performed based on the boundary of the current sub-frame or the boundary of a reference frame of the current block, depending on whether the current sub-frame including the current block is encoded independently.
[0016] Additionally, according to another aspect of this disclosure, a computer-readable recording medium can store a bitstream generated by the image encoding device or image encoding method of this disclosure.
[0017] Alternatively, according to another aspect of the transmission method of this disclosure, a bit stream generated by the image encoding device or image encoding method of this disclosure can be transmitted.
[0018] The features described above in this brief overview are merely exemplary aspects of the following detailed description of this disclosure and do not limit the scope of this disclosure.
[0019] Beneficial effects
[0020] According to this disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.
[0021] Furthermore, according to this disclosure, an image encoding / decoding method and apparatus based on winding motion compensation can be provided.
[0022] Furthermore, according to this disclosure, an image encoding / decoding method and apparatus based on scrolling motion compensation for independently encoded sub-pictures can be provided.
[0023] Furthermore, according to this disclosure, a method for transmitting a bitstream generated by an image encoding method or device according to this disclosure can be provided.
[0024] Furthermore, according to this disclosure, it is possible to provide a recording medium for storing a bitstream generated by an image encoding method or apparatus according to this disclosure.
[0025] Furthermore, according to this disclosure, a recording medium can be provided that stores a bitstream received, decoded, and used to reconstruct an image by an image decoding device according to this disclosure.
[0026] Those skilled in the art will understand that the effects achievable through this disclosure are not limited to those specifically described above, and that other advantages of this disclosure will become clearer from the detailed description. Attached Figure Description
[0027] Figure 1 This is a view that schematically illustrates a video coding system to which embodiments of the present disclosure are applicable.
[0028] Figure 2 This is a view schematically illustrating an image encoding device to which embodiments of the present disclosure are applicable.
[0029] Figure 3 This is a view schematically illustrating an image decoding device to which embodiments of the present disclosure are applicable.
[0030] Figure 4 This is a schematic flowchart illustrating the decoding process to which the embodiments of this disclosure apply.
[0031] Figure 5 This is a schematic flowchart illustrating the coding process to which the embodiments of this disclosure apply.
[0032] Figure 6 This is a flowchart illustrating a video / image decoding method based on inter-frame prediction.
[0033] Figure 7 This is a view illustrating the configuration of the inter-frame predictor 260 according to this disclosure.
[0034] Figure 8This is a view that shows an example of a sub-screen.
[0035] Figure 9 This is a view showing an example of an SPS that includes information about the sub-screen.
[0036] Figure 10 This is a view illustrating a method of encoding an image using sub-screens in an image encoding apparatus according to an embodiment of the present disclosure.
[0037] Figure 11 This is a view illustrating a method of decoding an image using a sub-screen according to an embodiment of the present disclosure.
[0038] Figure 12 This is a view showing an example of a 360-degree image converted to a two-dimensional screen.
[0039] Figure 13 This is a view illustrating an example of the horizontal winding motion compensation process.
[0040] Figure 14a This is a view illustrating an example of an SPS that includes information about winding motion compensation.
[0041] Figure 14b This is a view illustrating an example of a PPS that includes information about winding motion compensation.
[0042] Figure 15 This is a flowchart illustrating a method for an image decoding device to perform roll motion compensation based on sub-picture attributes.
[0043] Figure 16 This is a flowchart illustrating a method for an image encoding apparatus according to an embodiment of the present disclosure to determine whether to enable winding motion compensation.
[0044] Figure 17 This is a flowchart illustrating a method for performing scrolling motion compensation by an image decoding device based on sub-picture attributes according to an embodiment of the present disclosure.
[0045] Figure 18 This is a view illustrating an example of an SPS according to an embodiment of this disclosure.
[0046] Figure 19 This is a flowchart illustrating an image encoding method according to an embodiment of the present disclosure.
[0047] Figure 20 This is a flowchart illustrating an image decoding method according to an embodiment of the present disclosure.
[0048] Figure 21 This is a view illustrating the content streaming system to which embodiments of this disclosure are applicable.
[0049] Figure 22 This is a schematic illustration of an architecture for providing three-dimensional image / video services that can utilize embodiments of the present disclosure. Detailed Implementation
[0050] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings to facilitate implementation by those skilled in the art. However, this disclosure can be implemented in various different forms and is not limited to the embodiments described herein.
[0051] In describing this disclosure, detailed descriptions of relevant known functions or constructions will be omitted if they unnecessarily obscure the scope of this disclosure. In the accompanying drawings, portions irrelevant to the description of this disclosure are omitted, and similar reference numerals are assigned to similar portions.
[0052] In this disclosure, when a component is "connected," "coupled," or "linked" to another component, it may include not only direct connections but also indirect connections where intermediate components exist. Furthermore, when a component "comprises" or "has" other components, unless otherwise stated, it means that other components may be included, not excluded.
[0053] In this disclosure, the terms first, second, etc., are used only for the purpose of distinguishing one component from other components and do not limit the order or importance of the components, unless otherwise stated. Accordingly, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.
[0054] In this disclosure, the components are distinguished from each other to clearly describe each feature, but this does not mean that the components must be separate. That is, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed and implemented across multiple hardware or software units. Therefore, unless otherwise specified, implementations of these integrated or distributed components are included within the scope of this disclosure.
[0055] In this disclosure, the components described in the various embodiments are not necessarily essential components, and some components may be optional. Therefore, embodiments consisting of a subset of the components described in the embodiments are also included within the scope of this disclosure. Furthermore, embodiments that include other components besides those described in the various embodiments are also included within the scope of this disclosure.
[0056] This disclosure relates to the encoding and decoding of images. Unless redefined in this disclosure, the terms used herein may have the general meaning commonly used in the art to which this disclosure pertains.
[0057] In this disclosure, a "picture" generally refers to a unit representing an image within a specific time period, while a slice / tile is a coding unit that constitutes part of a picture. A picture can be composed of one or more slices / tiles. Furthermore, a slice / tile may include one or more coding tree units (CTUs).
[0058] In this disclosure, "pixel" or "pixel" can refer to the smallest unit that constitutes a frame (or image). Furthermore, "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or a pixel value, or it can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0059] In this disclosure, "unit" can refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information associated with that region. In some cases, the term "unit" may be used interchangeably with terms such as "sample array," "block," or "region." Generally, an M×N block may include a set (or array) of samples (or transform coefficients) with M columns and N rows.
[0060] In this disclosure, "current block" can mean one of "current coding block," "current coding unit," "coding target block," "decoding target block," or "processing target block." When performing prediction, "current block" can mean "current prediction block" or "prediction target block." When performing transform (inverse transform) / quantization (dequantization), "current block" can mean "current transform block" or "transform target block." When performing filtering, "current block" can mean "filter target block."
[0061] Furthermore, in this disclosure, unless explicitly stated as a chroma block, "current block" may mean a block that includes both luma component blocks and chroma component blocks, or "the luma block of the current block." The luma component block of the current block can be represented by an explicit description including terms such as "luma block" or "current luma block." Similarly, "the chroma component block of the current block" can be represented by an explicit description including terms such as "chroma block" or "current chroma block."
[0062] In this disclosure, the terms “ / ” or “,” can be interpreted as indicating “and / or”. For example, “A / B” and “A, B” can mean “A and / or B”. Furthermore, “A / B / C” and “A / B / C” can mean “at least one of A, B and / or C”.
[0063] In this disclosure, the term "or" should be interpreted to indicate "and / or". For example, the expression "A or B" may include 1) only "A", 2) only "B", or 3) both "A and B". In other words, in this disclosure, "or" should be interpreted to indicate "additionally or alternatively".
[0064] Overview of Video Encoding Systems
[0065] Figure 1 This is a schematic diagram illustrating a video coding system to which the embodiments of this disclosure are applicable.
[0066] The video encoding system according to the embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may deliver encoded video and / or image information or data to the decoding device 20 in the form of a file or stream via a digital storage medium or network.
[0067] The encoding device 10 according to an embodiment may include a video source generator 11, an encoding unit 12, and a transmitter 13. The decoding device 20 according to an embodiment may include a receiver 21, a decoding unit 22, and a renderer 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmitter 13 may be included in the encoding unit 12. The receiver 21 may be included in the decoding unit 22. The renderer 23 may include a display, and the display may be configured as a separate device or an external component.
[0068] The video source generator 11 can acquire video / images through a process of capturing, compositing, or generating video / images. The video source generator 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device may include, for example, a computer, tablet computer, and smartphone, and can generate video / images (electronically). For example, virtual video / images can be generated by a computer, etc. In this case, the video / image capture process can be replaced by a process of generating related data.
[0069] The encoding unit 12 can encode the input video / image. For compression and encoding efficiency, the encoding unit 12 can perform a series of processes, such as prediction, transformation, and quantization. The encoding unit 12 can output encoded data (encoded video / image information) in the form of a bitstream.
[0070] Transmitter 13 can transmit encoded video / image information or data, output in bitstream form, to receiver 21 of decoding device 20 in the form of a file or stream via digital storage medium or network. Digital storage medium can include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. Transmitter 13 can include elements for generating media files according to a predetermined file format and may include elements for transmission via broadcast / communication networks. Receiver 21 can extract / receive bitstreams from storage medium or network and transmit the bitstreams to decoding unit 22.
[0071] The decoding unit 22 can decode video / images by performing a series of processes corresponding to the operations of the encoding unit 12, such as dequantization, inverse transform, and prediction.
[0072] Renderer 23 can render decoded video / images. The rendered video / images can be displayed on a monitor.
[0073] Overview of Image Encoding Devices
[0074] Figure 2 This is a schematic view illustrating an image encoding device to which embodiments of this disclosure may be applied.
[0075] like Figure 2 As shown, the image encoding device 100 may include an image segmenter 110, a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter-frame predictor 180, an intra-frame predictor 185, and an entropy encoder 190. The inter-frame predictor 180 and the intra-frame predictor 185 may be collectively referred to as "predictors". The transformer 120, quantizer 130, dequantizer 140, and inverse transformer 150 may be included in a residual processor. The residual processor may also include a subtractor 115.
[0076] In some implementations, all or at least some of the components configuring the image encoding device 100 may be configured by a single hardware component (e.g., an encoder or a processor). Furthermore, the memory 170 may include a decoded screen buffer (DPB) and may be configured by a digital storage medium.
[0077] Image segmenter 110 can segment an input image (or picture or frame) input to image encoding device 100 into one or more processing units. For example, a processing unit may be called an encoding unit (CU). Encoding units can be obtained by recursively segmenting encoding tree units (CTUs) or maximum encoding units (LCUs) according to a quadtree / binary tree / tritree (QT / BT / TT) structure. For example, an encoding unit can be segmented into multiple encoding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the segmentation of encoding units, a quadtree structure can be applied first, followed by a binary tree structure and / or a ternary tree structure. The encoding process according to this disclosure can be performed based on the final encoding unit that is no longer segmented. The maximum encoding unit can be used as the final encoding unit, or a deeper encoding unit obtained by segmenting the maximum encoding unit can be used as the final encoding unit. Here, the encoding process may include prediction, transformation, and reconstruction processes, which will be described later. As another example, the processing unit of the encoding process may be a prediction unit (PU) or a transformation unit (TU). Prediction units and transform units can be partitioned or segmented from the final coding unit. Prediction units can be sample prediction units, and transform units can be units used to derive transform coefficients and / or units used to derive residual signals from transform coefficients.
[0078] The predictor (inter-frame predictor 180 or intra-frame predictor 185) can perform prediction on the block to be processed (the current block) and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. The predictor can generate various information related to the prediction of the current block and transmit the generated information to the entropy encoder 190. The information about the prediction can be encoded in the entropy encoder 190 and output as a bitstream.
[0079] Intra-predictor 185 can predict the current block by referencing samples in the current frame. Depending on the intra-prediction mode and / or intra-prediction technique, the reference samples may be located among the neighbors of the current block or may be placed separately. Intra-prediction modes may include multiple non-directional modes and multiple directional modes. Non-directional modes may include, for example, DC mode and planar mode. Depending on the level of detail in the prediction direction, directional modes may include, for example, 33 or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. Intra-predictor 185 can determine the prediction mode to be applied to the current block by using prediction modes applied to neighboring blocks.
[0080] Inter-frame predictor 180 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference frame. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, dual prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. The reference frame including the reference block and the reference frame including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a juxtaposed reference block, a juxtaposed CU (colCU), etc. The reference frame including the temporally neighboring block may be referred to as a juxtaposed frame (colPic). For example, inter-frame predictor 180 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate to use to deduce the motion vector and / or reference frame index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 180 can use motion information from neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, residual signals may not be transmitted. In motion vector prediction (MVP) mode, motion vectors from neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be signaled by encoding motion vector differences and indicators of the motion vector predictors. The motion vector difference can refer to the difference between the motion vector of the current block and the motion vector predictor.
[0081] The predictor can generate a prediction signal based on various prediction methods and techniques described below. For example, the predictor can apply not only intra-frame prediction or inter-frame prediction, but also both intra-frame prediction and inter-frame prediction simultaneously to predict the current block. A prediction method that simultaneously applies both intra-frame prediction and inter-frame prediction to predict the current block can be called Combined Intra-Frame and Inter-Frame Prediction (CIIP). Furthermore, the predictor can perform Intra-Frame Block Copy (IBC) to predict the current block. Intra-Frame Block Copy can be used for content image / video coding in games, such as Screen Content Coding (SCC). IBC is a method of predicting the current frame using a previously reconstructed reference block in the current frame at a predetermined distance from the current block. When IBC is applied, the position of the reference block in the current frame can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC essentially performs prediction in the current frame, but can be performed similarly to inter-frame prediction because the reference block is derived within the current frame. That is, IBC can use at least one inter-frame prediction technique described in this disclosure.
[0082] The predicted signal generated by the predictor can be used to generate a reconstructed signal or a residual signal. Subtractor 115 generates a residual signal (residual block or residual sample array) by subtracting the predicted signal (predicted block or predicted sample array) output from the predictor from the input image signal (original block or original sample array). The generated residual signal can be transmitted to converter 120.
[0083] Transformer 120 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform techniques may include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loève Transform (KLT), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented graphically. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. Furthermore, the transform processing can be applied to square pixel blocks of the same size or to blocks of variable size instead of square.
[0084] Quantizer 130 quantizes the transform coefficients and transmits them to entropy encoder 190. Entropy encoder 190 encodes the quantized signal (information about the quantized transform coefficients) and outputs a bitstream. This information about the quantized transform coefficients can be referred to as residual information. Quantizer 130 can rearrange the block-type quantized transform coefficients into a one-dimensional vector based on the coefficient scan order and generate information about the quantized transform coefficients based on this one-dimensional vector form.
[0085] The entropy encoder 190 can perform various encoding methods, such as exponential Columbus coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 190 can encode, either together or separately, the information required for video / image reconstruction other than the quantization transform coefficients (e.g., values of syntax elements). The encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form at the Network Abstraction Layer (NAL) level. The video / image information may also include information about various parameter sets, such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Furthermore, the video / image information may also include general constraint information. The signaled information, transmitted information, and / or syntax elements described in this disclosure can be encoded and included in the bitstream through the above encoding process.
[0086] The bitstream can be transmitted over a network or stored in a digital storage medium. The network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not shown) for transmitting the signal output from the entropy encoder 190 and / or a storage unit (not shown) for storing the signal may be included as internal / external components of the image encoding device 100. Alternatively, a transmitter may be provided as a component of the entropy encoder 190.
[0087] The quantization transform coefficients output from quantizer 130 can be used to generate residual signals. For example, the residual signals (residual blocks or residual samples) can be reconstructed by applying dequantization and inverse transform to the quantization transform coefficients through dequantizer 140 and inverse transformer 150.
[0088] Adder 155 adds the reconstructed residual signal to the prediction signal output from inter-frame predictor 180 or intra-frame predictor 185 to generate a reconstructed signal (reconstructed frame, reconstructed block, reconstructed sample array). If the block to be processed has no residual, such as in the case of applying skip mode, the prediction block can be used as a reconstructed block. Adder 155 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current frame, and can be used for inter-frame prediction of the next frame by filtering as described below.
[0089] Filter 160 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 160 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 170, specifically in the DPB of memory 170. Various filtering methods can include, for example, deblocking filtering, sample adaptive offsetting, adaptive loop filtering, bilateral filtering, etc. Filter 160 can generate various filtering-related information and transmit the generated information to entropy encoder 190, as described later in the description of each filtering method. The filtering-related information can be encoded by entropy encoder 190 and output as a bitstream.
[0090] The modified reconstructed frame transmitted to memory 170 can be used as a reference frame in inter-frame predictor 180. When inter-frame prediction is applied by image encoding device 100, prediction mismatch between image encoding device 100 and image decoding device can be avoided and coding efficiency can be improved.
[0091] The DPB of memory 170 can store modified reconstructed frames for use as reference frames in inter-frame predictor 180. Memory 170 can store motion information of blocks from which motion information in the current frame is derived (or encoded) and / or motion information of already reconstructed blocks in the frame. The stored motion information can be transmitted to inter-frame predictor 180 and used as motion information for spatially or temporally neighboring blocks. Memory 170 can store reconstructed samples of reconstructed blocks in the current frame and can transmit the reconstructed samples to intra-frame predictor 185.
[0092] Overview of image decoding devices
[0093] Figure 3 This is a schematic view illustrating an image decoding device to which embodiments of the present disclosure may be applied.
[0094] like Figure 3 As shown, the image decoding device 200 may include an entropy decoder 210, a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame predictor 260, and an intra-frame predictor 265. The inter-frame predictor 260 and the intra-frame predictor 265 may be collectively referred to as "predictors". The dequantizer 220 and the inverse transformer 230 may be included in a residual processor.
[0095] According to an implementation, all or at least some of the components of the image decoding device 200 can be configured by hardware components (e.g., a decoder or a processor). Furthermore, the memory 250 may include a decoded screen buffer (DPB) or may be configured by a digital storage medium.
[0096] The image decoding device 200, having received a bitstream including video / image information, can perform operations related to... Figure 2 The image is reconstructed by processing corresponding to the processing performed by the image encoding device 100. For example, the image decoding device 200 can perform decoding using a processing unit applied in the image encoding device. Therefore, the decoding processing unit can be, for example, an encoding unit. The encoding unit can be obtained by segmenting a coding tree unit or a maximum coding unit. The reconstructed image signal decoded and output by the image decoding device 200 can be reproduced by a reproduction device (not shown).
[0097] Image decoding device 200 can receive data in bitstream form from... Figure 2The signal output by the image encoding device. The received signal can be decoded by the entropy decoder 210. For example, the entropy decoder 210 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets, such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). In addition, the video / image information may also include general constraint information. The image decoding device can also decode the picture based on the information about the parameter sets and / or general constraint information. The information and / or syntax elements notified / received by signals described in this disclosure can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 210 decodes the information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC, and outputs the values of the syntax elements required for image reconstruction and the quantized values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine the context model using information about the target syntax element, decoding information of neighboring blocks and the target block, or information about symbols / bins decoded in the previous stage, perform arithmetic decoding on the bins based on the determined context model by predicting the occurrence probability of the bins, and generate symbols corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method can update the context model after determining the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin. The prediction-related information in the information decoded by the entropy decoder 210 can be provided to the predictors (inter-frame predictor 260 and intra-frame predictor 265), and the residual value of entropy decoding performed in the entropy decoder 210, i.e., the quantization transform coefficients and related parameter information, can be input to the dequantizer 220. In addition, the filtering information in the information decoded by the entropy decoder 210 can be provided to the filter 240. Furthermore, the receiver (not shown) for receiving signals output from the image encoding device may be further configured as an internal / external element of the image decoding device 200, or the receiver may be a component of the entropy decoder 210.
[0098] Furthermore, the image decoding apparatus according to this disclosure can be referred to as a video / image / screen decoding apparatus. The image decoding apparatus can be divided into an information decoder (video / image / screen information decoder) and a sample decoder (video / image / screen sample decoder). The information decoder may include an entropy decoder 210. The sample decoder may include at least one of a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame predictor 260, or an intra-frame predictor 265.
[0099] Dequantizer 220 can dequantize the quantized transform coefficients and output transform coefficients. Dequantizer 220 can rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement can be performed based on the coefficient scan order performed in the image encoding device. Dequantizer 220 can obtain transform coefficients by performing dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information).
[0100] The inverse transformer 230 can perform inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).
[0101] The predictor can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on the prediction information output from the entropy decoder 210, and can determine a specific intra-frame / inter-frame prediction mode (prediction technique).
[0102] Similar to that described in the predictor of the image coding device 100, the predictor can generate a predictive signal based on various prediction methods (techniques) described later.
[0103] Intra-predictor 265 can predict the current block by referring to samples in the current frame. The description of intra-predictor 185 also applies to intra-predictor 265.
[0104] Inter-frame predictor 260 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference frame. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, dual prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. For example, inter-frame predictor 260 can configure a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference frame index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the inter-frame prediction mode of the current block.
[0105] Adder 235 generates a reconstruction signal (reconstructed frame, reconstruction block, reconstruction sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including inter-frame predictor 260 and / or intra-frame predictor 265). If the block to be processed has no residual (e.g., in the case of applying skip mode), the prediction block can be used as a reconstruction block. The description of adder 155 also applies to adder 235. Adder 235 may be referred to as a reconstructor or reconstruction block generator. The generated reconstruction signal can be used for intra-frame prediction of the next block to be processed in the current frame, and can be used for inter-frame prediction of the next frame by filtering as described below.
[0106] Filter 240 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 240 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 250, specifically in the DPB of memory 250. Various filtering methods may include, for example, deblocking filtering, adaptive sample shifting, adaptive loop filtering, bilateral filtering, etc.
[0107] The (modified) reconstructed frame stored in the DPB of memory 250 can be used as a reference frame in inter-frame predictor 260. Memory 250 can store motion information of blocks from which motion information in the current frame is derived (or decoded) and / or motion information of already reconstructed blocks in the frame. The stored motion information can be transmitted to inter-frame predictor 260 to be used as motion information for spatially or temporally neighboring blocks. Memory 250 can store reconstructed samples of reconstructed blocks in the current frame and transmit the reconstructed samples to intra-frame predictor 265.
[0108] In this disclosure, the embodiments described in the filter 160, inter-frame predictor 180 and intra-frame predictor 185 of the image encoding device 100 can be equally or correspondingly applied to the filter 240, inter-frame predictor 260 and intra-frame predictor 265 of the image decoding device 200.
[0109] Overview of inter-frame prediction
[0110] The inter-frame prediction according to this disclosure will be described below.
[0111] According to this disclosure, the predictor of the image coding / decoding apparatus can perform inter-frame prediction on a block-by-block basis to derive prediction samples. Inter-frame prediction can be a prediction derived in a manner that depends on data elements (e.g., sample values, motion information, etc.) of frames other than the current frame. When inter-frame prediction is applied to the current block, the predicted block (prediction block or prediction sample array) of the current block can be derived based on a reference block (reference sample array) specified by a motion vector on a reference frame indicated by a reference frame index. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information of the current block can be predicted on a block, sub-block, or sample-by-sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction type (L0 prediction, L1 prediction, dual prediction, etc.) information. When inter-frame prediction is applied, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. The reference frame including the reference block and the reference frame including the temporally neighboring block may be the same or different. A temporally neighboring block can be referred to as a collated reference block, a collated CU (colCU), or a colBlock, and a reference picture that includes a temporally neighboring block can be referred to as a collated picture (colPic) or a colPicture. For example, a candidate list of motion information can be constructed based on the neighboring blocks of the current block, and flags or index information of which candidate can be selected (used) can be signaled to derive the motion vector and / or reference picture index of the current block.
[0112] Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the motion information of the current block can be equal to the motion information of the selected neighboring blocks. In skip mode, unlike merge mode, residual signals may not be sent. In motion information prediction (MVP) mode, the motion vectors of the selected neighboring blocks can be used as motion vector predictors, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictors and the motion vector difference. In this disclosure, MVP mode can have the same meaning as Advanced Motion Vector Prediction (AMVP).
[0113] Depending on the inter-frame prediction type (L0 prediction, L1 prediction, dual prediction, etc.), motion information can include L0 motion information and / or L1 motion information. A motion vector in the L0 direction can be called an L0 motion vector or MVL0, and a motion vector in the L1 direction can be called an L1 motion vector or MVL1. Prediction based on the L0 motion vector can be called L0 prediction, prediction based on the L1 motion vector can be called L1 prediction, and prediction based on both L0 and L1 motion vectors can be called dual prediction. Here, the L0 motion vector can indicate the motion vector (L0) associated with the reference frame list L0, and the L1 motion vector can indicate the motion vector (L1) associated with the reference frame list L1. The reference frame list L0 can include frames that precede the current frame in output order as reference frames, and the reference frame list L1 can include frames that follow the current frame in output order. Previous frames can be called forward (reference) frames, and subsequent frames can be called backward (reference) frames. The reference frame list L0 can also include frames that follow the current frame in output order as reference frames. In this scenario, within the reference screen list L0, previous screens can be indexed first, followed by subsequent screens. The reference screen list L1 can also include screens that precede the current screen in the output order. In this case, within the reference screen list L1, subsequent screens can be indexed first, followed by previous screens. Here, the output order can correspond to the screen order count (POC) order.
[0114] Figure 4 This is a flowchart illustrating a video / image coding method based on inter-frame prediction.
[0115] Figure 5 This is a view illustrating the configuration of the inter-frame predictor 180 according to this disclosure.
[0116] Figure 4 The encoding method can be determined by Figure 2 The image encoding device performs the following steps: Specifically, step S410 can be performed by the inter-frame predictor 180, and step S420 can be performed by the residual processor. Specifically, step S420 can be performed by the subtractor 115. Step S430 can be performed by the entropy encoder 190. The prediction information of step S430 can be derived by the inter-frame predictor 180, and the residual information of step S430 can be derived by the residual processor. The residual information is information about the residual samples. The residual information may include information about the quantization transform coefficients used for the residual samples. As described above, the residual samples can be derived into transform coefficients by the transformer 120 of the image encoding device, and the transform coefficients can be derived into quantization transform coefficients by the quantizer 130. The information about the quantization transform coefficients can be encoded by the entropy encoder 190 through the residual encoding process.
[0117] Refer to together Figure 4 and Figure 5 The image coding device can perform inter-frame prediction for the current block (S410). The image coding device can deduce the inter-frame prediction mode and motion information for the current block and generate prediction samples for the current block. Here, the inter-frame prediction mode determination, motion information deduction, and prediction sample generation processes can be performed simultaneously, or any one of them can be performed before other processes. For example, as... Figure 5 As shown, the inter-frame predictor 180 of the image coding apparatus may include a prediction mode determination unit 181, a motion information derivation unit 182, and a prediction sample derivation unit 183. The prediction mode determination unit 181 determines the prediction mode of the current block, the motion information derivation unit 182 derives the motion information of the current block, and the prediction sample derivation unit 183 derives the prediction samples of the current block. For example, the inter-frame predictor 180 of the image coding apparatus can search for blocks similar to the current block within a predetermined region (search region) of a reference frame using motion estimation, and derive a reference block whose difference from the current block is equal to or less than a predetermined criterion or minimum value. Based on this, a reference frame index indicating the reference frame in which the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The image coding apparatus can determine the mode applicable to the current block among various prediction modes. The image coding apparatus can compare rate distortion (RD) costs for various prediction modes and determine the optimal prediction mode for the current block. However, the method by which the image coding apparatus determines the prediction mode of the current block is not limited to the above example, and various methods can be used.
[0118] For example, when a skip mode or merge mode is applied to the current block, the image encoding device can deduce merge candidates from neighboring blocks of the current block and use the deduced merge candidates to construct a merge candidate list. Alternatively, the image encoding device can deduce reference blocks from among the reference blocks indicated by the merge candidates included in the merge candidate list that have a difference from the current block equal to or less than a predetermined criterion or minimum value. In this case, a merge candidate associated with the deduced reference block can be selected, and merge index information indicating the selected merge candidate can be generated and signaled to the image decoding device. Motion information of the current block can be deduced using the motion information of the selected merge candidate.
[0119] As another example, when the MVP mode is applied to the current block, the image encoding device can derive motion vector predictor (MVP) candidates from the neighboring blocks of the current block and construct an MVP candidate list using the derived MVP candidates. Alternatively, the image encoding device can use the motion vector of an MVP candidate selected from the MVP candidate list as the MVP of the current block. In this case, for example, the motion vector of a reference block derived through the above motion estimation can be used as the motion vector of the current block, and the MVP candidate with the motion vector having the smallest difference from the motion vector of the current block can be the selected MVP candidate. The motion vector difference (MVD) can be derived as the difference obtained by subtracting the MVP from the motion vector of the current block. In this case, the index information indicating the selected MVP candidate and information about the MVD can be signaled to the image decoding device. Additionally, when the MVP mode is applied, the value of the reference frame index can be constructed as reference frame index information and signaled separately to the image decoding device.
[0120] The image coding device can derive residual samples based on the predicted samples (S420). The image coding device can derive residual samples by comparing the original samples of the current block with the predicted samples. For example, residual samples can be derived by subtracting the corresponding predicted samples from the original samples.
[0121] The image encoding device can encode image information including prediction information and residual information (S430). The image encoding device can output the encoded image information in the form of a bitstream. The prediction information can include prediction mode information (e.g., skip flag, merge flag, or mode index, etc.) and motion information as information related to the prediction process. In the prediction mode information, the skip flag indicates whether the skip mode is applied to the current block, and the merge flag indicates whether the merge mode is applied to the current block. Alternatively, the prediction mode information can indicate one of a variety of prediction modes, similar to a mode index. When the skip flag and the merge flag are 0, it can be determined that the MVP mode is applied to the current block. The motion information can include candidate selection information (e.g., merge index, MVP flag, or MVP index) as information for deriving motion vectors. In the candidate selection information, the merge index can be signaled when the merge mode is applied to the current block, and can be information for selecting one of the merge candidates included in the merge candidate list. In the candidate selection information, the MVP flag or MVP index can be signaled when the MVP mode is applied to the current block, and can be information for selecting one of the MVP candidates in the MVP candidate list. Additionally, motion information may include information about the aforementioned MVD and / or reference frame index information. Furthermore, motion information may include information indicating whether L0 prediction, L1 prediction, or dual prediction is applied. Residual information is information about the residual samples. Residual information may include information about the quantization transformation coefficients used for the residual samples.
[0122] The output bitstream can be stored in (digital) storage media and sent to an image decoding device or it can be sent to an image decoding device via a network.
[0123] As described above, the image coding device can generate a reconstructed frame (including a frame of reconstructed samples and reconstructed blocks) based on reference samples and residual samples. This is to allow the image coding device to derive the same prediction result as that performed by the image decoding device, thereby improving coding efficiency. Therefore, the image coding device can store the reconstructed frame (or reconstructed samples and reconstructed blocks) in memory and use it as a reference frame for inter-frame prediction. As mentioned above, the in-loop filtering process is also applied to the reconstructed frame.
[0124] Figure 6 This is a flowchart illustrating a video / image decoding method based on inter-frame prediction. Figure 7 This is a view illustrating the configuration of the inter-frame predictor 260 according to this disclosure.
[0125] An image decoding device can perform operations corresponding to those performed by an image encoding device. The image decoding device can perform predictions for the current block and derive prediction samples based on received prediction information.
[0126] Figure 6 The decoding method can be derived from Figure 3 The image decoding device performs the steps S610 to S630. Steps S610 to S630 can be performed by the inter-frame predictor 260, and the prediction information of step S610 and the residual information of step S640 can be obtained from the bitstream by the entropy decoder 210. The residual processor of the image decoding device can derive the residual samples of the current block based on the residual information (S640). Specifically, the dequantizer 220 of the residual processor can perform dequantization based on the dequantized transform coefficients derived from the residual information to derive the transform coefficients, and the inverse transformer 230 of the residual processor can perform an inverse transform on the transform coefficients to derive the residual samples of the current block. Step S650 can be performed by the adder 235 or the reconstructor.
[0127] Refer to together Figure 6 and Figure 7 The image decoding device can determine the prediction mode of the current block based on the received prediction information (S610). The image decoding device can determine which inter-frame prediction mode is applied to the current block based on the prediction mode information in the prediction information.
[0128] For example, a skip flag can be used to determine whether a skip mode applies to the current block. Alternatively, a merge flag can be used to determine whether a merge mode or an MVP mode applies to the current block. Alternatively, a mode index can be used to select one of various inter-frame prediction mode candidates. Inter-frame prediction mode candidates may include skip mode, merge mode, and / or MVP mode, or may include various inter-frame prediction modes as described below.
[0129] The image decoding device can deduce motion information for the current block based on the determined inter-frame prediction mode (S620). For example, when a skip mode or merge mode is applied to the current block, the image decoding device can construct a merge candidate list (described below) and select one of the merge candidates included in the merge candidate list. The selection can be performed based on the aforementioned candidate selection information (merge index). The motion information of the current block can be deduced using the motion information of the selected merge candidate. For example, the motion information of the selected merge candidate can be used as the motion information of the current block.
[0130] As another example, when the MVP mode is applied to the current block, the image decoding device can construct an MVP candidate list and use the motion vector of the MVP candidate selected from the MVP candidates included in the MVP candidate list as the MVP of the current block. Selection can be performed based on the aforementioned candidate selection information (MVP flag or MVP index). In this case, the MVD of the current block can be derived based on information about the MVD, and the motion vector of the current block can be derived based on the MVP and MVD of the current block. Additionally, the reference frame index of the current block can be derived based on reference frame index information. The frame indicated by the reference frame index in the reference frame list of the current block can be derived as the reference frame to be referenced for inter-frame prediction of the current block.
[0131] The image decoding device can generate a prediction sample for the current block based on the motion information of the current block (S630). In this case, a reference frame can be derived based on the reference frame index of the current block, and the prediction sample for the current block can be derived using samples of the reference block indicated by the motion vector of the current block on the reference frame. In some cases, a prediction sample filtering process can also be performed on all or some of the prediction samples of the current block.
[0132] For example, such as Figure 7 As shown, the inter-frame predictor 260 of the image decoding device may include a prediction mode determination unit 261, a motion information derivation unit 262, and a prediction sample derivation unit 263. In the inter-frame predictor 260 of the image decoding device, the prediction mode determination unit 261 can determine the prediction mode of the current block based on the received prediction mode information, the motion information derivation unit 262 can derive the motion information (motion vector and / or reference frame index, etc.) of the current block based on the received motion information, and the prediction sample derivation unit 263 can derive the prediction samples of the current block.
[0133] The image decoding device can generate residual samples for the current block based on the received residual information (S640). The image decoding device can generate reconstructed samples for the current block based on the predicted samples and residual samples, and generate a reconstructed image based on this (S650). Thereafter, the in-loop filtering process is applied to the reconstructed image as described above.
[0134] As described above, the inter-frame prediction process may include the steps of determining an inter-frame prediction mode, deriving motion information based on the determined prediction mode, and performing prediction (generating prediction samples) based on the derived motion information. As described above, the inter-frame prediction process may be performed by an image encoding device and an image decoding device.
[0135] Overview of sub-screen
[0136] The sub-screens according to this disclosure will be described below.
[0137] A frame can be divided into tiles, and each tile can be further divided into sub-frames. Each sub-frame can include one or more slices and forms a rectangular area within the frame.
[0138] Figure 8 This is a view that shows an example of a sub-screen.
[0139] Reference Figure 8 A single frame can be divided into 18 tiles. Twelve tiles can be positioned on the left side of the frame, and each tile can include a sub-frame / slice consisting of 16 CTUs. Additionally, six tiles can be positioned on the right side of the frame, and each tile can include two sub-frames / slices consisting of four CTUs. As a result, the frame can be divided into 24 sub-frames, and each sub-frame can include a slice.
[0140] Information about subframes (e.g., the number and size of subframes) can be encoded or signaled using advanced syntax such as SPS, PPS, and / or slice headers.
[0141] Figure 9 This is an example view of SPS that includes information about the sub-screens.
[0142] Reference Figure 9 The SPS can include the syntax element `subpic_info_present_flag` specifying whether subpicture information exists for a Coding Layer Video Sequence (CLVS). For example, a `subpic_info_present_flag` with a first value (e.g., 0) can specify that no subpicture information exists for the CLVS and that only one subpicture exists in each frame of the CLVS. A `subpic_info_present_flag` with a second value (e.g., 1) can specify that subpicture information exists for the CLVS and that one or more subpictures exist in each frame of the CLVS. In the example, when the frame spatial resolution within the CLVS of the reference SPS can be changed (e.g., `res_change_in_clvs_allowed_flag == 1`), the value of `subpic_info_present_flag` should be equal to the first value (e.g., 0). Furthermore, when the bitstream is the result of a sub-bitstream extraction process and only contains a subset of subpictures from the input bitstream of the sub-bitstream extraction process, the value of `subpic_info_present_flag` should be the second value (e.g., 1).
[0143] Additionally, SPS can include the syntax element `sps_num_subpics_minus1` indicating the number of subpicks. For example, adding 1 to `sps_num_subpics_minus1` can specify the number of subpicks included in each picture in CLVS. In the example, the value of `sps_num_subpics_minus1` should be in the range of 0 to `Ceil(pic_width_max_in_luma_samples / CtbSizeY)*Ceil(pic_height_max_in_luma_samples / CtbSizeY)` (inclusive). Here, `Ceil(x)` can be the `ceiling` function used to output the smallest integer value greater than or equal to `x`. Furthermore, `pic_width_max_in_luma_samples` can refer to the maximum width of the luminance sample units in each picture, `pic_height_max_in_luma_samples` can refer to the maximum height of the luminance sample units in each picture, and `CtbSizeY` can refer to the array size of each Luminance Component Coded Tree Block (CTB) in both width and height. Furthermore, when sps_num_subpics_minus1 does not exist, the value of sps_num_subpics_minus1 can be inferred to be equal to the first value (e.g., 0).
[0144] Additionally, SPS can include the syntax element `sps_independent_subpics_flag` specifying whether subpick boundaries are considered picture boundaries. For example, `sps_independent_subpics_flag` with a second value (e.g., 1) can specify that all subpick boundaries in CLVS are considered picture boundaries and that there is no loop filtering across subpick boundaries. In contrast, `sps_independent_subpics_flag` with a first value (e.g., 0) can specify that the above constraint is not applied. Furthermore, when `sps_independent_subpics_flag` is not present, its value can be inferred to be equal to the first value (e.g., 0).
[0145] Additionally, SPS may include syntax elements subpic_ctu_top_left_x[i], subpic_ctu_top_left_y[i], subpic_width_minus1[i], and subpic_height_minus1[i] that specify the position and size of subpicts.
[0146] `subpic_ctu_top_left_x[i]` specifies the horizontal position of the top-left CTU of the i-th subpicture in units of `CtbSizeY`. In the example, the length of `subpic_ctu_top_left_x[i]` can be `Ceil(Log2((pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY))` bits. Furthermore, when `subpic_ctu_top_left_x[i]` does not exist, its value can be inferred to be equal to the first value (e.g., 0).
[0147] `subpic_ctu_top_left_y[i]` specifies the vertical position of the top-left CTU of the i-th subpicture in units of `CtbSizeY`. In the example, the length of `subpic_ctu_top_left_y[i]` can be `Ceil(Log2((pic_height_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY))` bits. Furthermore, when `subpic_ctu_top_left_y[i]` does not exist, its value can be inferred to be equal to the first value (e.g., 0).
[0148] Increasing `subpic_width_minus1[i]` by 1 specifies the width of the i-th subpic in units of `CtbSizeY`. In the example, the length of `subpic_width_minus1[i]` can be `Ceil(Log2((pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY))` bits. Furthermore, when `subpic_width_minus1[i]` does not exist, its value can be inferred to be equal to `((pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)-subpic_ctu_top_left_x[i]-1`.
[0149] Increasing `subpic_height_minus1[i]` by 1 specifies the height of the i-th subpic in units of `CtbSizeY`. In the example, the length of `subpic_height_minus1[i]` can be `Ceil(Log2((pic_height_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY))` bits. Furthermore, when `subpic_height_minus1[i]` does not exist, its value can be inferred to be equal to `((pic_height_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)-subpic_ctu_top_left_y[i]-1`.
[0150] Additionally, SPS can include `subpic_treated_as_pic_flag[i]` specifying whether a subpic is considered a picture. For example, `subpic_treated_as_pic_flag[i]` with a first value (e.g., 0) can specify that the i-th subpic of each encoded picture in the CLVS is not considered a picture during decoding, except for in-loop filtering operations. In contrast, `subpic_treated_as_pic_flag[i]` with a second value (e.g., 1) can specify that the i-th subpic of each encoded picture in the CLVS is considered a picture during decoding, except for in-loop filtering operations. When `subpic_treated_as_pic_flag[i]` is not present, the value of `subpic_treated_as_pic_flag[i]` can be inferred to be equal to the aforementioned `sps_independent_subpics_flag`. In the example, `subpic_treated_as_pic_flag[i]` can only be encoded / signaled if the aforementioned `sps_independent_subpics_flag` has a first value (e.g., 0) (i.e., when the subpic boundary is not considered a picture boundary).
[0151] Furthermore, when subpic_treated_as_pic_flag[i] has a second value (e.g., 1), the requirement for bitstream consistency is that for each output layer in the set of output layers (OLS) that includes the layer containing the i-th subpic as the output layer and its reference layer, all of the following conditions are true.
[0152] (Condition 1) All images in the output layer and its reference layer should have the same value for pic_width_in_luma_samples and the same value for pic_height_in_luma_samples.
[0153] (Condition 2) All SPS referenced by the output layer and its reference layer should have the same value of sps_num_subpics_minus1, and should also have the same values of subpic_ctu_top_left_x[j], subpic_ctu_top_left_y[j], subpic_width_minus1[j], subpic_height_minus1[j], and loop_filter_across_subpic_enabled_flag[j]. Here, j is in the range from 0 to sps_num_subpics_minus1 (inclusive).
[0154] Additionally, SPS can include the syntax element `loop_filter_across_subpic_enabled_flag[i]` specifying whether in-loop filtering is performed across subpic boundaries. For example, `loop_filter_across_subpic_enabled_flag[i]` with a first value (e.g., 0) can specify that in-loop filtering is not performed across the boundaries of the i-th subpic in each encoded picture within the CLVS. `loop_filter_across_subpic_enabled_flag[i]` with a second value (e.g., 1) can specify that in-loop filtering is performed across the boundaries of the i-th subpic in each encoded picture within the CLVS. When `loop_filter_across_subpic_enabled_flag[i]` is not present, its value can be inferred to be equal to 1 - `sps_independent_subpics_flag`. In the example, loop_filter_across_subpic_enabled_flag[i] can only be encoded / signaled if the aforementioned sps_independent_subpics_flag has a first value (e.g., 0) (i.e., when subpiction boundaries are not considered picture boundaries). Furthermore, the requirement for bitstream consistency is that the shape of the subpictions should be such that when each subpiction is decoded, its entire left and top boundaries should consist of picture boundaries or the boundaries of previously decoded subpictions.
[0155] Figure 10This is a view illustrating a method of encoding an image using sub-screens in an image encoding apparatus according to an embodiment of the present disclosure.
[0156] The image encoding device can encode the current frame based on the sub-frame structure. Alternatively, the image encoding device can encode at least one sub-frame that constitutes the current frame and output a (sub)bitstream that includes (encoded) information of at least one (encoded) sub-frame.
[0157] Reference Figure 10 The image encoding device can divide the input image into multiple sub-images (S1010). Additionally, the image encoding device can generate information about the sub-images (S1020). This information may include, for example, information about the area of the sub-image and / or information about the grid spacing within the sub-image. Furthermore, the information about the sub-images may include information indicating whether each sub-image is considered an image and / or information indicating whether in-loop filtering can be performed across the boundaries of each sub-image.
[0158] The image encoding device can encode at least one sub-frame based on information about the sub-frame. For example, each sub-frame can be encoded independently based on information about the sub-frame. Furthermore, the image encoding device can encode image information including information about the sub-frames and output a bitstream (S1030). Here, the bitstream of the sub-frames can be referred to as a sub-stream or sub-bitstream.
[0159] Figure 11 This is a view illustrating a method of decoding an image using a sub-screen according to an embodiment of the present disclosure.
[0160] An image decoding device can use the encoding information of at least one (encoded) sub-picture obtained from the (sub)bitstream to decode at least one sub-picture included in the current picture.
[0161] Reference Figure 11 The image decoding device can obtain information about a sub-picture from the bitstream (S1110). Here, the bitstream may include a sub-stream of the sub-picture or a sub-bitstream. The information about the sub-picture can be constructed according to the high-level syntax of the bitstream. In addition, the image decoding device can deduce at least one sub-picture based on the information about the sub-picture (S1120).
[0162] The image decoding device can decode at least one sub-picture based on information about the sub-picture (S1130). For example, when the current sub-picture, including the current block, is considered a picture, the current sub-picture can be decoded independently. Additionally, when in-loop filtering can be performed across the boundary of the current sub-picture, in-loop filtering (e.g., deblocking filtering) can be performed on the boundary of the current sub-picture and the boundaries of neighboring sub-pictures adjacent to the aforementioned boundary. Furthermore, when the boundary of the current sub-picture matches the picture boundary, in-loop filtering is not performed across the boundary of the current sub-picture. The image decoding device can decode the sub-picture based on methods such as CABAC, prediction methods, residual processing methods (transformation and quantization), and in-loop filtering methods. Furthermore, the image decoding device can output at least one decoded sub-picture or output a current picture including at least one sub-picture. The decoded sub-pictures can be output in the form of an output sub-picture set (OPS). For example, regarding a 360-degree image or an omnidirectional image, when only a portion of the current picture is rendered, only some of all the sub-pictures in the current picture can be decoded, and all or some of the decoded sub-pictures can be rendered according to the user's viewport.
[0163] Overview of winding
[0164] When inter-frame prediction is applied to the current block, the predicted block for the current block can be derived based on the reference block specified by the motion vector of the current block. In this case, when at least one reference sample in the reference block is outside the boundary of the reference frame, the sample value of the reference sample can be replaced by the sample value of a neighboring sample existing at the outermost edge or boundary of the reference frame. This can be called padding, and the boundary of the reference frame can be extended by padding.
[0165] Furthermore, when obtaining a reference frame from a 360-degree image, there may be continuity between the left and right boundaries of the reference frame. Therefore, a sample adjacent to the left (or right) boundary of the reference frame may have sample values and / or similar motion information to the sample adjacent to the right (or left) boundary of the frame. Based on these characteristics, at least one reference sample in the reference block outside the reference frame can be replaced by a neighboring sample in the reference frame corresponding to the reference sample. This can be called (horizontal) rollover motion compensation, and the motion vector of the current block can be adjusted through rollover motion compensation to indicate the interior of the reference frame.
[0166] Scrolling motion compensation refers to coding tools designed to improve the visual quality of reconstructed images / videos (e.g., 360-degree images / videos projected in an equal rectangular projection (ERP) format). According to existing motion compensation processes, the motion vector of the current block references samples outside the boundary of a reference frame. The sample values of these samples outside the boundary are derived by repeatedly padding to copy the sample values of the nearest neighbor samples on the boundary. However, since 360-degree images / videos are captured on a sphere and inherently lack image boundaries, reference samples outside the boundary of the reference frame in the projection domain (two-dimensional domain) can always be obtained from neighboring samples adjacent to the reference sample in the spherical domain (three-dimensional domain). Therefore, repeated padding is unsuitable for 360-degree images / videos and results in a visual artifact called "seam artifact" in the reconstructed viewport image / video.
[0167] When applying general projection formats, obtaining neighboring samples for roll motion compensation in the spherical domain can be challenging due to the 2D-to-3D and 3D-to-2D coordinate transformations and sample interpolation for fractional sample positions. However, when applying the ERP projection format, spherical neighboring samples outside the left (or right) boundary of the reference frame can be relatively easily obtained from samples within the right (or left) boundary of the reference frame. Given the widespread use and relative ease of implementation of the ERP projection format, horizontal roll motion compensation may be more effective for 360-degree images / videos encoded in the ERP projection format.
[0168] Figure 12 This is a view showing an example of a 360-degree image converted to a two-dimensional screen.
[0169] Reference Figure 12 A 360-degree image 1210 can be converted into a two-dimensional image 1230 through a projection process. Depending on the projection method applied to the 360-degree image 1210, the two-dimensional image 1230 can have various projection formats such as equal rectangular projection (ERP) or filled ERP (PERP).
[0170] Due to the characteristics of images acquired in all directions, the 360-degree image 1210 does not have image boundaries. However, due to the projection process, the two-dimensional image 1230 obtained from the 360-degree image 1210 has image boundaries. In this case, the left boundary LBd and right boundary RBd of the two-dimensional image 1230 form a line RL within the 360-degree image 1210 and may touch each other. Therefore, the similarity between samples adjacent to the left boundary LBd and right boundary RBd within the two-dimensional image 1230 may be relatively high.
[0171] Furthermore, based on the reference image boundary, a predetermined region in the 360-degree image 1210 can correspond to an inner or outer region of the two-dimensional image 1230. For example, based on the left boundary LBd of the two-dimensional image 1230, region A in the 360-degree image 1210 can correspond to region A1 outside the two-dimensional image 1230. In contrast, based on the right boundary RBd of the two-dimensional image 1230, region A in the 360-degree image 1210 can correspond to region A2 inside the two-dimensional image 1230. Regions A1 and A2 correspond to the same region A based on the 360-degree image 1210, and therefore have the same / similar sample properties.
[0172] Based on these characteristics, external samples outside the left boundary LBd of the two-dimensional image 1230 can be replaced by internal samples located at a predetermined distance away on the first direction DIR 1 of the two-dimensional image 1230 through a winding motion compensation. For example, external samples included in region A1 of the two-dimensional image 1230 can be replaced by internal samples included in region A2 of the two-dimensional image 1230. Similarly, external samples outside the right boundary RBd of the two-dimensional image 1230 can be replaced by internal samples located at a predetermined distance away on the second direction DIR 2 of the two-dimensional image 1230 through a winding motion compensation.
[0173] Figure 13 This is a view illustrating an example of the horizontal winding motion compensation process.
[0174] Reference Figure 13 When inter-frame prediction is applied to the current block 1310, the prediction block of the current block 1310 can be derived based on the reference block 1330.
[0175] Reference block 1330 can be specified by motion vector 1320 of current block 1310. In the example, motion vector 1320 can indicate the upper left position of reference block 1330 relative to the upper left position of the same position block 1315 that exists in the same position as current block 1310 in the reference frame.
[0176] like Figure 13 As shown, reference block 1330 may include a first region 1335 outside the left boundary of the reference frame. The first region 1335 cannot be used for inter-frame prediction of the current block 1310, and therefore can be replaced by a second region 1340 in the reference frame through rollover motion compensation. The second region 1340 may correspond to the same region as the first region 1335 in the spherical domain (three-dimensional domain), and the position of the second region 1340 can be specified by adding a rollover offset to a predetermined position (e.g., the upper left position) of the first region 1335.
[0177] A wraparound offset can be set to the ERP width before the current frame is filled. Here, the ERP width can refer to the width of the original frame (i.e., the ERP frame) in ERP format obtained from a 360-degree image. A horizontal fill process can be performed relative to the left and right boundaries of the ERP frame. Therefore, the width of the current frame, PicWidth, can be determined as a value obtained by adding the ERP width, the left fill of the left boundary of the ERP frame, and the right fill of the right boundary of the ERP frame. Furthermore, predefined syntax elements in the high-level syntax (e.g., pps_ref_wraparound_offset) can be used to encode / signal the wraparound offset. These syntax elements are unaffected by the fill of the left and right boundaries of the ERP frame, resulting in support for asymmetrical fill of the original frame. That is, the left fill of the left boundary of the ERP frame and the right fill of the right boundary can be different from each other.
[0178] Information regarding the aforementioned winding motion compensation (e.g., enable, winding offset, etc.) can be encoded / signed using advanced syntax such as SPS and / or PPS.
[0179] Figure 14a This is a view illustrating an example of an SPS that includes information about winding motion compensation.
[0180] Reference Figure 14a SPS can include the syntax element `sps_ref_wraparound_enabled_flag` that specifies whether wraparound motion compensation is applied at the sequence level. For example, `sps_ref_wraparound_enabled_flag` with a first value (e.g., 0) can specify that wraparound motion compensation is not applied to the current video sequence including the current block. In contrast, `sps_ref_wraparound_enabled_flag` with a second value (e.g., 1) can specify that wraparound motion compensation is applied to the current video sequence including the current block. In the example, wraparound motion compensation for the current video sequence is applied only if the frame width (e.g., `pic_width_in_luma_samples`) and CTB width `CtbSizeY` meet the following conditions.
[0181] -(Condition): (CtbSizeY / MinCbSizeY+1)≥(pic_width_in_luma_samples / MinCbSizeY-1)
[0182] When the above conditions are not met, for example, when the value of (CtbSizeY / MinCbSizeY+1) is greater than the value of (pic_width_in_luma_samples / MinCbSizeY-1), sps_ref_wraparound_enabled_flag should be equal to the first value (e.g., 0). Here, CtbSizeY can refer to the width or height of the luma component CTB, and MinCbSizeY can refer to the minimum width or height of the luma component coded block (CB). Additionally, pic_width_max_in_luma_samples can refer to the maximum width of the luma sample unit in each frame.
[0183] Figure 14b This is a view illustrating an example of a PPS that includes information about winding motion compensation.
[0184] Reference Figure 14b PPS can include the syntax element pps_ref_wraparound_enabled_flag that specifies whether to apply wraparound motion information at the screen level.
[0185] Reference Figure 14b The PPS can include the syntax element `pps_ref_wraparound_enabled_flag` that specifies whether wraparound motion information is applied at the sequence level. For example, a `pps_ref_wraparound_enabled_flag` with a first value (e.g., 0) can specify that wraparound motion compensation is not applied to the current frame including the current block. In contrast, a `pps_ref_wraparound_enabled_flag` with a second value (e.g., 1) can specify that wraparound motion compensation is applied to the current frame including the current block. In the example, wraparound motion compensation for the current frame is applied only if the frame width (e.g., `pic_width_in_luma_samples`) is greater than the CTB width `CtbSizeY`. For example, when the value of `(CtbSizeY / MinCbSizeY+1)` is greater than the value of `(pic_width_in_luma_samples / MinCbSizeY-1)`, `pps_ref_wraparound_enabled_flag` should be equal to the first value (e.g., 0). In another example, when sps_ref_wraparound_enabled_flag has a first value (e.g., 0), the value of pps_ref_wraparound_enabled_flag should be equal to the first value (e.g., 0).
[0186] Additionally, PPS can include the syntax element `pps_ref_wraparound_offset` to specify the offset for wrapping motion compensation. For example, `pps_ref_wraparound_offset` plus `((CtbSizeY / MinCbSizeY)+2)` can specify the wrapping offset used to calculate the wrapping position in units of luminance samples. The value of `pps_ref_wraparound_offset` can be in the range of 0 to `((pic_width_in_luma_samples / MinCbSizeY)-(CtbSizeY / MinCbSizeY)-2)` (inclusive). Furthermore, the variable `PpsRefWraparoundOffset` can be set to equal to `(pps_ref_wraparound_offset+(CtbSizeY / MinCbSizeY)+2)`. The variable `PpsRefWraparoundOffset` can be used in the process of cropping reference samples outside the boundaries of the reference image.
[0187] In addition, when the current screen is divided into multiple sub-screens, scrolling motion compensation can be performed based on the attributes of each sub-screen.
[0188] Figure 15 This is a flowchart illustrating a method for an image decoding device to perform roll motion compensation based on sub-picture attributes.
[0189] Reference Figure 15 The image decoding device can determine whether the current image has been independently encoded (S1510).
[0190] When the current sub-frame is independently encoded (S1510 is "Yes"), the image decoding device can crop the position of the reference sample based on the sub-frame boundary used for motion compensation of the current block (S1520). The above operation can be performed, for example, using any of the following processes: luminance sample bilinear interpolation, luminance sample interpolation, luminance integer sample extraction, or chrominance sample interpolation.
[0191] When the current sub-picture is not independently encoded (S1510 is "No"), the image decoding device can determine whether to enable roll motion compensation for the current block (S1530).
[0192] Whether to enable wraparound motion compensation for the current block can be determined based on a predefined variable (e.g., refWraparoundEnabledFlag). For example, when refWraparoundEnabledFlag has a first value (e.g., 0), wraparound motion compensation may not be enabled for the current block. In contrast, when refWraparoundEnabledFlag has a second value (e.g., 1), wraparound motion compensation may be enabled for the current block. In the example, the value of refWraparoundEnabledFlag can be derived based on predefined flags (e.g., pps_ref_wraparound_enabled_flag) obtained from high-level syntax (e.g., the set of screen parameters).
[0193] When wrap motion compensation is enabled for the current block (S1530 is "Yes"), the image decoding device can use a wrap offset to modify the position of the reference sample (S1540). For example, the image decoding device can modify the position of the reference sample by shifting the x-coordinate of the reference sample by a wrap offset in the positive or negative direction (e.g., PpsRefWraparoundOffset*MinCbSizeY). Additionally, the image decoding device can crop the modified reference sample position based on the boundary of the reference frame used for motion compensation of the current block (S1550).
[0194] In contrast, when roll motion compensation is not enabled for the current block (S1530 is "No"), the image decoding device can crop the position of the reference sample based on the boundary of the reference frame used for motion compensation of the current block (S1560).
[0195] Furthermore, the above cropping operation can be performed using a bilinear interpolation process for brightness samples. Detailed examples are shown in Table 1 below.
[0196] [Table 1]
[0197]
[0198] Referring to Table 1, the predefined clipping functions (Clip3, ClipH) can be used to adjust the brightness position (xInti, yInti) of the reference sample in integer sample units within the reference frame boundary or sub-frame boundary. Here, Clip3(x, y, z) means a function that outputs x when z is less than x, outputs y when z is greater than y, and outputs z otherwise. Similarly, ClipH(x, y, z) means a function that outputs z+x when z is less than 0, outputs zx when z is greater than y-1, and outputs z otherwise.
[0199] When the current sub-picture is encoded independently, process 1 can be executed. Specifically, a cropping operation based on the sub-picture boundaries (A110 and A120) can be performed on the x and y coordinates of the reference sample. In process 1, SubpicLeftBoundaryPos indicates the left boundary of the sub-picture, SubpicRightBoundaryPos indicates the right boundary of the sub-picture, SubpicTopBoundaryPos indicates the top boundary of the sub-picture, and SubpicBotBoundaryPos indicates the bottom boundary of the sub-picture.
[0200] When the current sub-picture is not independently encoded, process 2 can be executed. Specifically, for the x-coordinate of the reference sample, a wrap-around offset (e.g., PpsRefWraparoundOffset) can be selectively applied depending on whether wrap-around motion compensation is enabled for the current block, and a clipping operation based on the reference picture boundary can be performed for the selectively applied x-coordinate (A130). In contrast, for the y-coordinate of the reference sample, a clipping operation based on the reference picture boundary can be performed regardless of whether wrap-around motion compensation is enabled (A140). In process 2, picW can indicate the reference picture width, and picH can indicate the reference picture height.
[0201] Alternatively, the above cropping operation can be performed using a luminance sample interpolation filtering process. Detailed examples are shown in Table 2 below.
[0202] [Table 2]
[0203]
[0204] Referring to Table 2, the brightness position (xInti, yInti) of the reference sample in integer sample units can be adjusted within the reference frame boundary or sub-frame boundary using predetermined clipping functions (Clip3, ClipH). Repeated descriptions from Table 1 will be omitted.
[0205] Process 1 or 2 can be selectively executed depending on whether the current sub-screen is independently encoded. Alternatively, when the current sub-screen is not independently encoded (process 2), winding motion compensation can be performed only for the x-coordinate of the reference sample (A230).
[0206] Alternatively, the above cropping operation can be performed using the brightness integer sample extraction process. Detailed examples are shown in Table 3 below.
[0207] [Table 3]
[0208]
[0209] Referring to Table 3, the brightness position (xInt, yInt) of the reference sample in integer sample units can be adjusted within the reference frame boundary or sub-frame boundary using predetermined clipping functions (Clip3, ClipH). Repeated descriptions from Table 1 will be omitted.
[0210] Process 1 or 2 can be selectively executed depending on whether the current sub-screen is independently encoded. Alternatively, when the current sub-screen is not independently encoded (process 2), winding motion compensation can be performed only for the x-coordinate of the reference sample (A330).
[0211] Alternatively, the above cropping operation can be performed using the chroma sample interpolation process. Detailed examples are shown in Table 4 below.
[0212] [Table 4]
[0213]
[0214] Referring to Table 4, the luminance position (xInt, yInt) of the reference sample in integer sample units can be adjusted within the reference frame boundary or sub-frame boundary using predetermined clipping functions (Clip3, ClipH). In this case, unlike the methods described above with reference to Tables 1 to 3, the reference frame boundary or sub-frame boundary can be determined based on the chroma sample. For example, SubWidthC and SubHeightC can indicate the width ratio and height ratio between the luminance sample and the chroma sample. Furthermore, based on the chroma sample, the left boundary of the sub-frame can be determined as (SubpicLeftBoundaryPos / SubWidthC), the right boundary as (SubpicRightBoundaryPos / SubWidthC), the top boundary as (SubpicTopBoundaryPos / SubWidthC), and the bottom boundary as (SubpicBotBoundaryPos / SubWidthC).
[0215] Process 1 or 2 can be selectively executed depending on whether the current sub-screen is independently encoded. Alternatively, when the current sub-screen is not independently encoded (process 2), winding motion compensation can be performed only for the x-coordinate of the reference sample (A430).
[0216] Furthermore, it is obvious to those skilled in the art that Figure 15 The method can also be performed by an image encoding device.
[0217] according to Figure 15The method only performs roll motion compensation for the current block when the current sub-picture is not independently encoded. As a result, roll-related encoding tools cannot be used with various sub-picture-related encoding tools that are based on independent sub-picture encoding. This can degrade encoding / decoding performance for pictures with continuity between boundaries (e.g., ERP pictures or PEP pictures).
[0218] To address this problem, according to embodiments of this disclosure, even when the current sub-screen is encoded independently, winding motion compensation can be performed based on predetermined conditions. Embodiments of this disclosure will be described in detail below.
[0219] According to an embodiment of this disclosure, when all independently coded sub-pictures in the current video sequence have a width equal to the picture width, scrolling motion compensation can be enabled for all sub-pictures in the current video sequence.
[0220] Figure 16 This is a flowchart illustrating a method for an image encoding apparatus according to an embodiment of the present disclosure to determine whether to enable winding motion compensation.
[0221] Reference Figure 16 The image encoding device can determine whether there are one or more independently encoded sub-pictures in the current video sequence (S1610).
[0222] When it is determined that there are no one or more independently encoded sub-frames in the current video sequence (S1610 is "No"), the image encoding device may determine whether to enable wraparound motion compensation for the current video sequence based on a predetermined wraparound constraint (S1640). In this case, the image encoding device may encode the flag information specifying whether wraparound motion compensation is enabled for the current video sequence (e.g., sps_wraparound_enabled_flag) into a first value (e.g., 0) or a second value (e.g., 1) based on this determination.
[0223] As an example of a wrap constraint, when wrap motion compensation is constrained for one or more Output Layer Sets (OLS) specified by a Video Parameter Set (VPS), wrap motion compensation should not be enabled for the current video sequence. As another example of a wrap constraint, when all subframes in the current video sequence have discontinuous subframe boundaries, wrap motion compensation should not be enabled for the current video sequence.
[0224] In contrast, when it is determined that there are one or more independently encoded sub-frames in the current video sequence (S1610 is "yes"), the image encoding device can determine whether there is a sub-frame with a width different from the frame width among the independently encoded sub-frames (S1620).
[0225] In one implementation, the frame width can be derived based on the maximum width of the frames in the current video sequence, as shown in Equation 1.
[0226] [Formula 1]
[0227] (pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY
[0228] Among them, pic_width_max_in_luma_samples can indicate the maximum screen width in units of luminance samples, CtbSizeY can indicate the width of the coding tree block (CTB) in the screen in units of luminance samples, and CtbLog2SizeY can indicate the logarithmic scale value of CtbSizeY.
[0229] When it is determined that at least one of the independently encoded sub-frames has a width different from the frame width (S1620 is "Yes"), the image encoding device can determine that wraparound motion compensation is disabled for the current video sequence (S1630). In this case, the image encoding device can encode sps_ref_wraparound_enabled_flag as a first value (e.g., 0).
[0230] In contrast, when it is determined that at least one of the independently encoded sub-frames has the same width as the frame width (S1620 is "No"), the image encoding device can determine whether to enable wraparound motion compensation for the current video sequence based on the aforementioned wraparound constraints (S1640). In this case, the image encoding device can encode sps_ref_wraparound_enabled_flag as a first value (e.g., 0) or a second value (e.g., 1) based on this determination.
[0231] Despite Figure 16 Steps S1610 and S1620 are shown to be executed sequentially, but this is only an example and the implementation of this disclosure is not limited thereto. For example, step S1620 may be executed simultaneously with step S1610 or before step S1610.
[0232] Furthermore, the `sps_ref_wraparound_enabled_flag` encoded by the image encoding device can be stored in the bitstream and signaled to the image decoding device. In this case, the image decoding device can determine whether wraparound motion information is enabled for the current video sequence based on the `sps_ref_wraparound_enabled_flag` obtained from the bitstream.
[0233] For example, when `sps_ref_wraparound_enabled_flag` has a first value (e.g., 0), the image decoding device can determine that wraparound motion compensation is disabled for the current video sequence and can therefore not perform wraparound motion compensation for the current block. In this case, the reference sample position of the current block can be cropped based on the reference frame boundary or subframe boundary, and motion compensation can be performed using the reference sample at the cropped position.
[0234] That is, the image decoding device can perform correct motion compensation according to this disclosure without separately determining whether there are one or more independently encoded subframes in the current video sequence with a width different from the screen width. However, the operation of the image decoding device is not limited to this. For example, the image decoding device can determine whether there are one or more independently encoded subframes in the current video sequence with a width different from the screen width, and then perform motion compensation based on the determination result. More specifically, the image decoding device can determine whether there are one or more independently encoded subframes in the current video sequence with a width different from the screen width, and when such subframes exist, it can avoid performing wraparound motion compensation by treating sps_ref_wraparound_enabled_flag as a first value (e.g., 0).
[0235] In contrast, when `sps_ref_wraparound_enabled_flag` has a second value (e.g., 1), the image decoding device can determine that wraparound motion compensation is enabled for the current video sequence. In this case, the image decoding device can additionally obtain a wraparound flag (e.g., `pps_ref_wraparound_enabled_flag`) from the bitstream specifying whether wraparound motion compensation is enabled for the current frame, and determine whether to perform wraparound motion compensation for the current block based on the obtained wraparound flag information.
[0236] For example, when pps_ref_wraparound_enabled_flag has a first value (e.g., 0), the image decoding device may not perform wraparound motion compensation for the current block. In this case, the reference sample position of the current block can be cropped based on a reference frame boundary or subframe boundary, and motion compensation can be performed using the reference sample at the cropped position. In contrast, when pps_ref_wraparound_enabled_flag has a second value (e.g., 1), the image decoding device can perform wraparound motion compensation for the current block.
[0237] As described above, when all independently coded subframes in the current video sequence have a width equal to the frame width, rollover motion compensation can be applied to all subframes in the current video sequence. Therefore, since subframe-related coding tools and rollover motion compensation-related coding tools can be used together, encoding / decoding efficiency can be further improved.
[0238] Figure 17 This is a flowchart illustrating a method for performing scrolling motion compensation by an image decoding device based on sub-picture attributes according to an embodiment of the present disclosure.
[0239] Reference Figure 17 The image decoding device can determine whether the current subpicture is encoded independently (S1710). In this example, whether the current subpicture is encoded independently can be determined based on high-level syntax, such as a predetermined flag in the SPS (e.g., subpic_treated_as_pic_flag). For example, when subpic_treated_as_pic_flag has a first value (e.g., 0), the current subpicture may not have been encoded independently. In contrast, when subpic_treated_as_pic_flag has a second value (e.g., 1), the current subpicture can be encoded independently.
[0240] When the current sub-picture is independently encoded (S1710 is "Yes"), the image decoding device can determine whether to enable roll motion compensation for the current block (S1720).
[0241] Whether to enable winding motion compensation for the current block can be determined based on a predetermined variable (e.g., refWraparoundEnabledFlag). For example, when refWraparoundEnabledFlag has a first value (e.g., 0), winding motion compensation may not be enabled for the current block. In contrast, when refWraparoundEnabledFlag has a second value (e.g., 1), winding motion compensation may be enabled for the current block. In the example, refWraparoundEnabledFlag can be set as shown in Equation 2 below.
[0242] [Equation 2]
[0243] refWraparoundEnabledFlag=pps_ref_wraparound_enabled_flag&&! refPicIsScaled
[0244] The `pps_ref_wraparound_enabled_flag` parameter specifies whether wraparound motion compensation is enabled for the current frame, including the current block (enabled when the value is 1, disabled when the value is 0). `pps_ref_wraparound_enabled_flag` can be obtained through advanced syntax (e.g., the Picture Parameter Set (PPS)). Additionally, the variable `refPicIsScaled` specifies whether reference frame scaling is performed (or whether reference frame scaling is required). For example, `refPicIsScaled` has a first value (e.g., 0) when reference frame scaling is performed, and a second value (e.g., 1) when reference frame scaling is not performed.
[0245] Referring to Equation 2, when wraparound motion compensation is not enabled for the current frame (e.g., pps_ref_wraparound_enabled_flag == 0) or when reference frame scaling is performed (e.g., refPicIsScaled == 1), refWraparoundEnabledFlag can have a first value (e.g., 0) specifying that wraparound motion compensation is not enabled for the current block. Conversely, when wraparound motion compensation is enabled for the current frame (e.g., pps_ref_wraparound_enabled_flag == 1) and reference frame scaling is not performed (e.g., refPicIsScaled == 0), refWraparoundEnabledFlag can have a second value (e.g., 1) specifying that wraparound motion compensation is enabled for the current block. Scrollaround motion compensation and reference frame scaling can be performed selectively.
[0246] In the implementation, the variable SliceRefWraparoundEnabledFlag can be defined and derived as follows regarding whether to enable winding motion compensation.
[0247] -SliceRefWraparoundEnabledFlag can be set to be equal to the above pps_ref_wrap_around_enabled_flag.
[0248] - When SliceRefWraparoundEnabledFlag has a second value (e.g., 1), subpic_treated_as_pic_flag[i] has a second value (e.g., 1) for the current subpic (i.e., the subpic to which the current slice belongs) (i.e., when the current subpic is encoded independently), and the width of the current subpic is different from the width of the picture, the value of SliceRefWraparoundEnabledFlag can be set to equal to the first value (e.g., 0).
[0249] - When SliceRefWraparoundEnabledFlag has a second value (e.g., 1) and the reference frame is not scaled, SliceRefWraparoundEnabledFlag can be used to enable the above refWraparoundEnabledFlag (i.e., refWraparoundEnabledFlag = 1).
[0250] The derivation of SliceRefWraparoundEnabledFlag can be as defined in the semantics of the slice header. Detailed examples are shown in Table 5.
[0251] [Table 5]
[0252]
[0253] Referring to Table 5, SliceRefWraparoundEnabledFlag can be set to equal pps_ref_wraparound_enabled_flag. Additionally, when SliceRefWraparoundEnabledFlag has a second value (e.g., 1), subpic_treated_as_pic_flag[CurrSubpicIdx] has a second value (e.g., 1), and subpic_width_minus1[CurrSubpicIdx]+1 is less than Ceil(pic_width_in_luma_samples÷CtbSizeY), the value of SliceRefWraparoundEnabledFlag can be reset to equal the first value (e.g., 0).
[0254] When roll motion compensation is enabled for the current block (S1720 is "Yes"), the image decoding device can use the roll offset to modify the position of the reference sample (S1730).
[0255] A wrap offset can be set to the ERP width before the current frame is filled. Here, the ERP width can refer to the width of the original frame in ERP format obtained from the 360-degree image (i.e., the ERP frame). The wrap offset can be determined based on predefined syntax elements (e.g., pps_ref_wraparound_offset) obtained through high-level syntax (e.g., the Picture Parameter Set (PPS)). Additionally, the x-coordinate of the reference sample can be shifted in either the positive or negative direction using the wrap offset.
[0256] The image decoding device can crop and modify the reference sample position based on the sub-screen boundary (S1740). Therefore, the position of the reference sample can be included within the sub-screen boundary.
[0257] In contrast, when rollover motion compensation is not enabled for the current block (S1720 is "No"), the image decoding device can crop the position of the reference sample based on the sub-picture boundary (S1750). Therefore, the position of the reference sample can be included in the sub-picture.
[0258] As described above, when the current sub-picture is encoded independently, regardless of whether rollover motion compensation is enabled for the current block, the clipping operation of the reference sample position can be performed based on the sub-picture boundary (S1740 and S1750). This is likely because the encoded sub-picture has continuity based on the sub-picture boundary, rather than the reference picture boundary.
[0259] Returning to step S1710, when the current sub-frame is not independently encoded (S1710 is "No"), the image decoding device can determine whether to enable wraparound motion compensation for the current block (S1760). As described above, whether to enable wraparound motion compensation for the current block can be determined based on a predetermined variable (e.g., refWraparoundEnabledFlag).
[0260] When roll motion compensation is enabled for the current block (S1760 is "Yes"), the image decoding device can use roll offset to modify the position of the reference sample (S1770). Additionally, the image decoding device can crop the modified reference sample position based on the reference screen boundary (S1780). Therefore, the position of the reference sample can be included within the reference screen boundary.
[0261] In contrast, when roll motion compensation is not enabled for the current block (S1760 is "No"), the image decoding device can crop the reference sample position based on the reference screen boundary (S1790).
[0262] As described above, when the current sub-frame is not independently encoded, regardless of whether rollover motion compensation is enabled for the current block, a clipping operation of the reference sample position can be performed based on the reference frame boundary (S1780, S1790). This may be because the non-encoded sub-frame is based on the reference frame boundary, rather than on the continuity of the sub-frame boundary.
[0263] Furthermore, the above-described cropping operation can be performed using any one of the following processes: luminance sample bilinear interpolation, luminance sample interpolation, luminance integer sample extraction, or chrominance sample interpolation. Additionally, according to embodiments of this disclosure, even when the current sub-frame is encoded independently, scrolling motion compensation can be performed. Therefore, process 1 described in Tables 1 to 4 can be modified as shown in Tables 6 to 9 below.
[0264] [Table 6]
[0265]
[0266] [Table 7]
[0267]
[0268] [Table 8]
[0269]
[0270] [Table 9]
[0271]
[0272] Referring to Tables 6 to 9, unlike existing cropping operations, even when the current sub-picture is encoded independently, if scroll motion compensation is enabled, the reference sample position (A610 to A910) can be cropped based on the sub-picture boundary.
[0273] Specifically, when the current sub-frame is encoded independently, process 1 can be executed. For the x-coordinate of the reference sample, a wrap-around offset (e.g., PpsRefWraparoundOffset) is selectively applied depending on whether wrap-around motion compensation is enabled for the current block (e.g., refWraparoundEnabledFlag == 1), and a cropping operation based on the sub-frame boundary can be performed for the selectively applied x-coordinate. However, for the y-coordinate of the reference sample, as described above with reference to Tables 1 to 4, a general cropping operation (or padding operation) based on the sub-frame boundary can be performed. This may mean that wrap-around motion compensation is only applied to the left and right boundaries of the reference frame (i.e., horizontal wrap-around motion compensation).
[0274] In Table 9, the winding offset xOffset used for winding motion compensation in the chromaticity sample interpolation process (see Table 8) can be derived as shown in Equation 3 below.
[0275] [Formula 3]
[0276] xOffset=(PpsRefWraparoundOffset)*MinCbSizeY) / SubWidthC
[0277] Referring to Equation 3, the wrap-around offset (xOffset) can be calculated by multiplying the offset information PpsRefWraparoundOffset obtained from the high-level syntax (e.g., the picture parameter set (PPS)) by the minimum width MinCbSizeY of the coded block (CB), and then dividing the multiplied value by the width ratio SubWidthC of the luma-chroma samples.
[0278] In the implementation, as described above, when the variable SliceRefWraparoundEnabledFlag is newly defined, the refWraparoundEnabledFlag in Tables 6 to 9 can be derived as shown in Equation 4 below, without using pps_ref_wraparound_enabled_flag obtained from PPS.
[0279] [Formula 4]
[0280] `refWraparoundEnabledFlag = SliceRefWraparoundEnabledFlag && ! refPicIsScaled` refers to Equation 4. When `SliceRefWraparoundEnabledFlag` has a first value (e.g., 0) or `refPicIsScaled` has a second value (e.g., 1), `refWraparoundEnabledFlags` can be set to equal the first value (e.g., 0). Conversely, when `SliceRefWraparoundEnabledFlag` has a second value (e.g., 1) and `refPicIsScaled` has a first value (e.g., 0), `refWraparoundEnabledFlags` can be set to equal the second value (e.g., 1).
[0281] In implementation, new variables LeftBoundaryPos, RightBoundaryPos, TopBoundaryPos, and / or BottomBoundaryPos can be defined to specify the boundaries of the subpic. These new variables can indicate the position of the subpic's boundaries when subpic_treated_as_flag has a first value (e.g., 0) or a second value (e.g., 1). Detailed examples of how to derive these new variables are shown in Table 10.
[0282] [Table 10]
[0283]
[0284] Referring to Table 10, the initial value of the variable LeftBoundaryPos can be set to 0. Additionally, the initial value of the variable RightBoundaryPos can be set to pic_width_in_luma_samples-1. Furthermore, the initial value of the variable TopBoundaryPos can be set to 0. Additionally, the initial value of BottomBoundaryPos can be set to pic_height_in_luma_samples-1.
[0285] When the current subpic is encoded independently (i.e., subpic_treated_as_pic_flag[CurrSubpicIdx] == 1), the values of LeftBoundaryPos, RightBoundaryPos, TopBoundaryPos, and BottomBoundaryPos can be updated. Specifically, the value of LeftBoundaryPos can be updated to subpic_ctu_top_left_x[CurrSubpicIdx] * CtbSizeY. Additionally, the value of RightBoundaryPos can be updated to Min(pic_width_max_in_luma_samples-1,(subpic_ctu_top_left_x[CurrSubpicIdx]+subpic_width_minus1[CurrSubpicIdx]+1)*CtbSizeY-1). Here, Min(x,y) refers to a function that outputs the smaller value between x and y. Additionally, the value of TopBoundaryPos can be updated to subpic_ctu_top_left_y[CurrSubpicIdx]*CtbSizeY. Furthermore, the value of BottomBoundaryPos can be updated to Min(pic_height_max_in_luma_samples-1,(subpic_ctu_top_left_y[CurrSubpicIdx]+subpic_height_minus1[CurrSubpicIdx]+1)*CtbSizeY-1).
[0286] The processes in Tables 6 to 9 above can be modified using new variables for sub-screen boundaries as shown in Tables 11 to 14.
[0287] [Table 11]
[0288]
[0289] [Table 12]
[0290]
[0291] [Table 13]
[0292]
[0293] [Table 14]
[0294]
[0295] Except for the use of new variables for sub-picture boundaries, the processes in Tables 11 to 14 are the same as those in Tables 6 to 9, and therefore repeated descriptions will be omitted. Furthermore, the Temporal Motion Vector Prediction (TMVP) derivation process, the sub-block-based temporal merging candidate derivation process, and the reconstructed affine control point motion vector merging candidate derivation process can be performed using the new variables for sub-picture boundaries.
[0296] Figure 18 This is a view illustrating an example of an SPS according to an embodiment of this disclosure. References above will be omitted. Figure 9 and Figure 14a The description of SPS is a repetitive description.
[0297] Reference Figure 18 The SPS can include a flag, `sps_ref_wraparound_enabled_flag`, specifying whether wraparound motion compensation is enabled at the sequence level. For example, `sps_ref_wraparound_enabled_flag` equal to a first value (e.g., 0) can specify that (horizontal) wraparound motion compensation is applied in inter-frame prediction. In contrast, `sps_ref_wraparound_enabled_flag` equal to a second value (e.g., 1) can specify that (horizontal) wraparound motion compensation is not applied.
[0298] The requirement for bitstream consistency is that, for all subpics in a frame, when the syntax element `subpic_width_minus1[i]` specifying the width of the i-th subpict plus 1 is different from `(pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)`, the value of `sps_ref_wraparound_enabled_flag` should be equal to the first value (e.g., 0). Here, `i` can be in the range of 0 to the value of the syntax element `sps_num_subpics_minus1` specifying the number of subpicts in the frame (inclusive).
[0299] In the implementation, the `sps_ref_wraparound_enabled_flag` can be signaled before the syntax elements of the subpicture (e.g., `subpic_info_present_flag`, `sps_num_subpics_minus1`, etc.). Alternatively, when `sps_ref_wraparound_enabled_flag` has a second value (e.g., 1), the syntax elements for the position and width of the subpicture (e.g., `subpic_ctu_top_left_x[i]`, `subpic_ctu_top_left_y[i]`, and `subpic_width_minus1[i]`) can be signaled without prior notification.
[0300] In this case, the value of the syntax element `subpic_ctu_top_left_x[i]`, which specifies the horizontal position of the top-left CTU of the i-th sub-picture in units of `CtbSizeY`, can be deduced to be equal to the first value (e.g., 0). Additionally, the syntax element `subpic_ctu_top_left_y[i]`, which specifies the vertical position of the top-left CTU of the i-th sub-picture in units of `CtbSizeY`, can be deduced to be equal to `subpic_ctu_top_left_y[i-1] + subpic_height_minus1[i-1]`. Here, `subpic_height_minus1[i-1]` can specify the height of the (i-1)-th sub-picture minus 1. Furthermore, the syntax element `subpic_width_minus1[i]`, which specifies the width of the sub-picture, can be deduced to be equal to `((pic_width_max_in_luma_samples + CtbSizeY - 1) >> CtbLog2SizeY) - 1`. Here, pic_width_max_in_luma_samples can indicate the maximum image width in luminance sample units, CtbSizeY can indicate the width of the coding tree block (CTB), and CtbLog2SizeY can indicate the logarithmic scale value of the CTB width.
[0301] In another implementation, when sps_ref_wraparound_enabled_flag has a second value (e.g., 1), the syntax elements for the position and width of the sub-screen should be deduced to the values mentioned above.
[0302] Reference Figure 18Only when sps_ref_wraparound_enabled_flag has a first value (e.g., 0) can subpic_ctu_top_left_x[i], subpic_ctu_top_left_y[i], and subpic_width_minus1[i] (1810) be signaled.
[0303] `subpic_ctu_top_left_x[i]` specifies the horizontal position of the top-left CTU of the i-th sub-picture in units of `CtbSizeY`. The length of `subpic_ctu_top_left_x[i]` can be `Ceil(Log2((pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)` bits. When `subpic_ctu_top_left_x[i]` does not exist in the bitstream (i.e., it is not signaled), the value of `subpic_ctu_top_left_x[i]` can be deduced to be equal to 0.
[0304] Additionally, `subpic_ctu_top_left_y[i]` can specify the vertical position of the top-left CTU of the i-th subpicture in units of `CtbSizeY`. The length of `subpic_ctu_top_left_y[i]` can be `Ceil(Log2((pic_height_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY))` bits. The following applies when `subpic_ctu_top_left_y[i]` does not exist in the bitstream (i.e., it is not signaled).
[0305] When the value of i is greater than 0 and sps_ref_wraparound_enabled_flag has a second value (e.g., 1), the value of subpic_ctu_top_left_y[i] can be deduced to be equal to (subpic_ctu_top_left_y[i-1]+subpic_height_minus1[i-1]+1).
[0306] Otherwise, the value of subpic_ctu_top_left_y[i] can be derived as equal to (((pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)-subpic_ctu_top_left_x[i]-1).
[0307] Additionally, incrementing `subpic_width_minus1[i]` by 1 specifies the width of the i-th subpic in units of `CtbSizeY`. The length of `subpic_width_minus1[i]` can be `Ceil(Log2((pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY))` bits. The following applies when `subpic_width_minus1[i]` is not present in the bitstream (i.e., it is not signaled).
[0308] When sps_ref_wraparound_enabled_flag has a second value (e.g., 1), the value of subpic_width_minus1[i] can be deduced to be equal to (((pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)-1).
[0309] Otherwise, the value of subpic_width_minus1[i] can be derived as equal to (((pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)-subpic_ctu_top_left_x[i]-1).
[0310] Furthermore, according to another implementation, the aforementioned inference conditions for subpic_ctu_top_left_x[i], subpic_ctu_top_left_y[i], and subpic_width_minus1[i] should be constrained. For example, the bitstream consistency constraint is that when sps_ref_wraparound_enabled_flag has a second value (e.g., 1), the value of subpic_ctu_top_left_x[i] is inferred to be equal to the first value (e.g., 0). Additionally, the bitstream consistency constraint is that when sps_ref_wraparound_enabled_flag has a second value (e.g., 1), subpic_ctu_top_left_y[i] is inferred to be equal to subpic_ctu_top_left_y[i-1] + subpic_height_minus1[i-1]. Additionally, the constraint on bitstream consistency is that when sps_ref_wraparound_enabled_flag has a second value (e.g., 1), subpic_width_minus1[i] is inferred to be equal to ((pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)-1.
[0311] According to embodiments of this disclosure, scrolling motion compensation applies to the current sub-picture regardless of whether it is independently encoded. In this case, when the current sub-picture is independently encoded, scrolling motion compensation can be performed based on the sub-picture boundary. In contrast, when the current sub-picture is not independently encoded, scrolling motion compensation can be performed based on a reference picture boundary. Therefore, since sub-picture-related encoding tools and scrolling motion compensation-related encoding tools can be used together, encoding / decoding efficiency can be further improved.
[0312] In the following text, reference will be made to Figure 19 and Figure 20 A detailed description of an image encoding / decoding method according to embodiments of the present disclosure is provided.
[0313] Figure 19 This is a flowchart illustrating an image encoding method according to an embodiment of the present disclosure.
[0314] Figure 19 Image encoding methods can be derived from Figure 2 The image encoding device performs the operation. For example, steps S1910 and S1920 can be performed by the inter-frame predictor 180, and step S1930 can be performed by the entropy encoder 190.
[0315] Reference Figure 19The image encoding device can determine whether the winding motion compensation is applied to the current block (S1910).
[0316] In one implementation, the image encoding device can determine whether to enable wrapping motion compensation based on whether there are one or more independently encoded sub-frames in the current video sequence, including the current block, with a width different from the screen width. For example, when there are one or more independently encoded sub-frames in the current video sequence with a width different from the screen width, the image encoding device can determine to disable wrapping motion compensation. Conversely, when all independently encoded sub-frames in the current video sequence have the same width as the screen width, the image encoding device can determine to enable wrapping motion compensation based on a predetermined wrapping constraint. Here, an example of a wrapping constraint is referred to above. Figures 16 to 18 It has been described.
[0317] Furthermore, the image encoding device can use this determination to decide whether to apply scroll motion compensation to the current block. For example, when scroll motion compensation is disabled for the current frame, the image encoding device can determine that scroll motion compensation should not be performed for the current block. In contrast, when scroll motion compensation is enabled for the current frame, the image encoding device can determine that scroll motion compensation should be performed for the current block.
[0318] In this implementation, since the current sub-picture is independently encoded and has a width different from that of the current picture, the rollover motion compensation of the current block can be skipped.
[0319] The image coding device can generate a prediction block for the current block by performing inter-frame prediction based on the determination result of step S1910 (S1920).
[0320] In this implementation, depending on whether the current subpic is independently encoded, motion compensation can be performed based on the boundary of the current subpic including the current block or the boundary of the reference picture of the current block. For example, when the current subpic is independently encoded (e.g., subpic_treated_as_pic_flag == 1), motion compensation for the current block can be performed based on the boundary of the current subpic. Conversely, when the current subpic is not independently encoded (e.g., subpic_treated_as_pic_flag == 0), motion compensation for the current block can be performed based on the boundary of the reference picture.
[0321] In this implementation, roll motion compensation for the current block can be performed by modifying the x-coordinate of a reference block in the reference frame based on a roll offset, where the reference block is specified by the motion vector of the current block. In the example, a roll offset can be set to the ERP width before the current frame is filled. Here, the ERP width can refer to the width of the original frame in ERP format obtained from a 360-degree image (i.e., the ERP frame). Alternatively, roll motion compensation for the current block can be performed by cropping the modified x-coordinate of the reference block to the boundary of the current sub-frame or the reference frame.
[0322] In this implementation, the winding motion compensation for the current block can be performed based on the left and right boundary positions, which are set based on whether the current sub-frame is independently encoded. For example, when the current sub-frame is independently encoded, the left and right boundary positions can be set to the left and right boundary positions of the current sub-frame. In contrast, when the current sub-frame is not independently encoded, the left and right boundary positions can be set to the left and right boundary positions of the reference frame. Therefore, since whether the current sub-frame is independently encoded is considered in the boundary positions used for cropping the reference block, it is not necessary to separately determine whether the current sub-frame is independently encoded. Therefore, the motion compensation process for generating the prediction block of the current block can be further simplified. The image encoding device can encode the inter-frame prediction information and the winding information for winding motion compensation of the current block to generate a bitstream (S1930).
[0323] In implementations, the wrapping information may include a first flag (e.g., pps_ref_wraparound_enabled_flag) specifying whether wrapping motion compensation is enabled for the current frame. Based on enabling wrapping motion compensation for the current video sequence including the current block (e.g., sps_ref_wraparound_enabled_flag == 0), the first flag may have a first value (e.g., 0) specifying that wrapping motion compensation is disabled for the current frame. Additionally, based on predetermined conditions regarding the width of the Coding Tree Block (CTB) in the current frame and the width of the current frame, the first flag may have a first value (e.g., 0) specifying that wrapping motion compensation is disabled for the current frame. For example, when the width of the CTB in the current frame (e.g., CtbSizeY) is greater than the width of the frame (e.g., pic_width_in_luma_samples), pps_ref_wraparound_enabled_flag should have a first value (e.g., 0).
[0324] In this implementation, based on enabling wraparound motion compensation for the current frame, the wraparound information may also include a wraparound offset (e.g., pps_ref_wraparound_offset). The image encoding device may perform wraparound motion compensation based on the wraparound offset.
[0325] Figure 20 This is a flowchart illustrating an image decoding method according to an embodiment of the present disclosure.
[0326] Figure 20 Image decoding methods can be derived from Figure 3 The image decoding device performs these steps. For example, steps S2010 and S2020 can be performed by the inter-frame predictor 260.
[0327] Reference Figure 20 The image decoding device can obtain inter-frame prediction information and wrapping information of the current block from the bitstream (S2010). Here, the inter-frame prediction information of the current block may include motion information of the current block, such as reference frame index, difference motion vector information, etc. The wrapping information may include a first flag specifying whether wrapping motion compensation is enabled for the current frame (e.g., pps_ref_wraparound_enabled_flag). Based on disabling wrapping motion compensation for the current video sequence including the current block (e.g., sps_ref_wraparound_enabled_flag == 0), the first flag may have a first value (e.g., 0) specifying that wrapping motion compensation is disabled for the current frame. In addition, based on predetermined conditions regarding the width of the coded tree block (CTB) in the current frame and the width of the current frame, the first flag may have a first value (e.g., 0) specifying that wrapping motion compensation is disabled for the current frame. For example, when the width of the CTB in the current frame (e.g., CtbSizeY) is greater than the width of the frame (e.g., pic_width_in_luma_samples), sps_ref_wraparound_enabled_flag should have a first value (e.g., 0).
[0328] In this implementation, based on enabling wraparound motion compensation for the current video sequence including the current block (e.g., sps_ref_wraparound_enabled_flag == 1), each of the top-left position and width of the current sub-pic can be set (or inferred) to a predetermined value. For example, in this case, the x-coordinate of the top-left position of the current sub-pic (e.g., subpic_ctu_top_left_x[i]) can be set to 0, and the y-coordinate of the top-left position of the current sub-pic (e.g., subpic_ctu_top_left_y[i]) can be set to equal to the y-coordinate of the first sub-pic decoded before the current pic (e.g., subpic_ctu_top_left_y[i-1]) plus the height of the current sub-pic (e.g., subpic_height_minus1[i-1]). Additionally, the width of the current subpic (e.g., subpic_width_minus[i]) can be set to be equal to the width of the current pic (e.g., ((pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)-1).
[0329] The image decoding device can generate a prediction block for the current block based on inter-frame prediction information and winding information obtained from the bitstream (S2020).
[0330] In one implementation, a predicted block for the current block can be generated by performing wraparound motion compensation, based on a first flag (e.g., pps_ref_wraparound_enabled_flag) having a predetermined value (e.g., 1) that specifies enabling wraparound motion compensation for the current block.
[0331] In this implementation, winding motion compensation for the current block can be performed by modifying the x-coordinate of a reference block in a reference frame based on the winding offset. Alternatively, winding motion compensation can be performed by clipping the modified x-coordinate of the reference block within the boundary of the current sub-frame or the boundary of the reference frame. In this case, the boundary used for clipping can be determined based on whether the current sub-frame including the current block is independently encoded (e.g., `subpic_treated_as_pic_flag`). For example, when the current sub-frame is independently encoded (e.g., `subpic_treated_as_pic_flag == 1`), the x-coordinate of the reference block can be clipped within the boundary of the current sub-frame. Conversely, when the current sub-frame is not independently encoded (e.g., `subpic_treated_as_pic_flag == 0`), the x-coordinate of the reference block can be clipped within the boundary of the reference frame.
[0332] Furthermore, in this implementation, since the reference frame is scaled, the roll motion compensation for the current block can be skipped. Additionally, since the current sub-frame is independently encoded and has a width different from the width of the current frame, the roll motion compensation for the current block can also be skipped.
[0333] In this implementation, the winding motion compensation for the current block can be performed based on the left and right boundary positions, which are set based on whether the current sub-frame is independently encoded. For example, when the current sub-frame is independently encoded, the left and right boundary positions can be set to the left and right boundary positions of the current sub-frame. In contrast, when the current sub-frame is not independently encoded, the left and right boundary positions can be set to the left and right boundary positions of the reference frame. Therefore, since whether the current sub-frame is independently encoded is considered in the boundary positions used to clip the reference block, it is not necessary to determine whether the current sub-frame is independently encoded separately. Thus, the motion compensation process for generating the predicted block of the current block can be further simplified.
[0334] According to the image encoding / decoding method based on embodiments of this disclosure, even if the current sub-picture is encoded independently, scrolling motion compensation can be used based on predetermined conditions. Therefore, since sub-picture related encoding tools and scrolling motion compensation related encoding tools can be used together, encoding / decoding efficiency can be further improved.
[0335] The names of the syntax elements described in this disclosure may include information about the location of the corresponding syntax element being signaled. For example, a syntax element beginning with "sps_" may indicate that the corresponding syntax element is signaled in the Sequence Parameter Set (SPS). Additionally, syntax elements beginning with "pps_", "ph_", or "sh_" may indicate that the corresponding syntax element is signaled in the Picture Parameter Set (PPS), the Picture Header, and the Slice Header, respectively.
[0336] Although the exemplary methods of this disclosure described above are represented as a series of operations for clarity of description, they are not intended to limit the order in which the steps are performed, and these steps may be performed simultaneously or in different orders if necessary. To implement the method according to the invention, the described steps may further include other steps, including steps in addition to some steps, or may include additional steps in addition to some steps.
[0337] In this disclosure, the image encoding device or image decoding device that performs a predetermined operation (step) can perform an operation (step) that confirms the execution conditions or circumstances of the corresponding operation (step). For example, if it is described that a predetermined operation is performed when predetermined conditions are met, the image encoding device or image decoding device can perform the predetermined operation after determining whether the predetermined conditions are met.
[0338] The various embodiments of this disclosure are not a list of all possible combinations and are intended to describe representative aspects of this disclosure; the matters described in the various embodiments may be applied independently or in combination of two or more.
[0339] Various embodiments of this disclosure can be implemented in hardware, firmware, software, or a combination thereof. When this disclosure is implemented in hardware, it can be implemented using application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, etc.
[0340] Furthermore, the image decoding and image encoding devices applying the embodiments of this disclosure can be included in multimedia broadcasting transmission and receiving devices, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video chat devices, real-time communication devices such as video communication, mobile streaming devices, storage media, cameras, video-on-demand (VoD) service providers, OTT (over-the-top) video devices, internet streaming service providers, three-dimensional (3D) video devices, video telephony devices, medical video devices, etc., and can be used to process video signals or data signals. For example, OTT video devices can include game consoles, Blu-ray players, internet access televisions, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.
[0341] Figure 21 This is a view illustrating the content streaming system to which embodiments of this disclosure are applicable.
[0342] like Figure 21 As shown, the content streaming system applying the embodiments of this disclosure may mainly include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.
[0343] The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream and then sends the bitstream to the streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server can be omitted.
[0344] The bitstream can be generated by an image encoding method or image encoding device applying the embodiments of this disclosure, and the stream server can temporarily store the bitstream during the sending or receiving of the bitstream.
[0345] A streaming server sends multimedia data to a user's device based on a request from a web server, and the web server acts as a medium for informing the user of the service. When a user requests a service from the web server, the web server can deliver it to the streaming server, and the streaming server can send the multimedia data to the user. In this scenario, the content streaming system may include a separate control server. In this case, the control server is used to control the commands / responses between devices in the content streaming system.
[0346] A streaming server can receive content from media storage devices and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined period of time.
[0347] Examples of user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays), digital televisions, desktop computers, digital signage, etc.
[0348] In a content streaming system, each server can operate as a distributed server, in which case the data received from each server can be distributed.
[0349] Figure 22 This is a schematic illustration of an architecture for providing three-dimensional image / video services that can utilize embodiments of the present disclosure.
[0350] Figure 22 Examples of 360-degree or omnidirectional video / image processing systems can be given. Additionally, Figure 22 The system can be implemented, for example, in extended reality (XR) supporting devices. That is, the system can provide a method for providing virtual reality to users.
[0351] Extended reality collectively refers to virtual reality (VR), augmented reality (AR), and mixed reality (MR). VR technology only provides CG images of real-world objects or backgrounds, AR technology provides virtually created CG images on top of images of real objects, and MR technology is a computer graphics technology used to blend, combine, and provide virtual objects in the real world.
[0352] MR (Mixed Reality) and AR (Augmented Reality) technologies are similar in that real and virtual objects are displayed together. However, while virtual objects are used to complement real objects in AR, virtual and real objects are used with equal characteristics in MR.
[0353] XR technology is applicable to head-mounted displays (HMDs), head-up displays (HUDs), cellular phones, tablet PCs, laptops, desktops, TVs, digital signage, etc. Devices that utilize XR technology can be referred to as XR devices. XR devices may include a first digital device and / or a second digital device, which will be described below.
[0354] 360 content refers to the entire content used to implement and provide VR, and may include 360-degree video and / or 360-degree audio. 360-degree video can refer to video or image content required to provide VR by simultaneously capturing or playing in all directions (360 degrees or less). In the following, 360 video can refer to 360-degree video. 360-degree audio is also audio content used to provide VR, and can refer to spatial audio content that enables a sound source to be identified as being located in a specific three-dimensional space. 360-degree content can be generated, processed, and sent to a user who can use the 360-degree content to consume a VR experience. 360-degree video can be referred to as omnidirectional video, and 360-degree image can be referred to as omnidirectional image. In the following, the focus will be on 360-degree video, and embodiments of this disclosure are not limited to VR, but may include the processing of video / image content such as AR or MR. 360-degree video can refer to video or image displayed in a 3D space with various shapes based on a 3D model; for example, 360-degree video can be displayed on a sphere.
[0355] Specifically, this method proposes an efficient way to provide 360-degree video. To provide 360-degree video, firstly, 360-degree video can be captured by one or more cameras. The captured 360-degree video can be sent through a series of processes, and the data received by the receiver can be processed into raw 360-degree video and rendered. Therefore, 360-degree video can be provided to the user.
[0356] Specifically, the entire process for providing 360-degree video may include a capture process, a preparation process, a transmission process, a processing process, a rendering process, and / or a feedback process.
[0357] The capture process can refer to the process of capturing images or videos in multiple views using one or more cameras. Figure 22 The image / video data shown in 2210 can be generated through the capture process. Figure 22 The planes of 2210 can refer to images / videos of various views. Multiple captured images / videos can be referred to as raw data. Metadata related to the capture can be generated during the capture process.
[0358] For capture, a special camera designed for VR can be used. In some implementations, when providing 360-degree video of a computer-generated virtual space, capture via a real camera may not be necessary. In this case, the capture process can be simply replaced by a process for generating relevant data.
[0359] The preparation process can involve processing captured images / videos and metadata generated during the capture process. Captured images / videos may undergo stitching, projection, region-by-region packing, and / or encoding processes during preparation.
[0360] First, the individual images / videos can undergo a stitching process. This stitching process can be the process of connecting captured images / videos to generate a panoramic image / video or a spherical image / video.
[0361] Subsequently, the stitched images / videos can undergo a projection process. During projection, the stitched images / videos can be projected onto a 2D image. Depending on the context, this 2D image can be referred to as a 2D image frame. Projection onto a 2D image can be expressed as mapping to a 2D image. The projected image / video data can have... Figure 20 The form of the 2D image shown in 2220.
[0362] Video data projected onto a 2D image can undergo a region-by-region packing process to increase video encoding efficiency. Region-by-region packing refers to the process of dividing and processing video data projected onto a 2D image according to regions. Here, a region can refer to the area into which a 2D image projecting 360-degree video data is divided. Depending on the implementation, these regions can be obtained by equally or arbitrarily dividing the 2D image. Additionally, in some implementations, regions can be divided according to the projection scheme. The region-by-region packing process is optional and can be omitted during preparation.
[0363] In some implementations, the process may include rotating or rearranging regions on a 2D image to increase video coding efficiency. For example, rotating regions so that specific sides of the regions are closer together can increase coding efficiency.
[0364] In some implementations, the processing may include increasing or decreasing the resolution of specific regions to differentiate the resolution of different regions on the 360-degree video. For example, the resolution of a region corresponding to a relatively more important region on the 360-degree video may be higher than that of other regions. Video data projected onto a 2D image or region-by-region packaged video data may undergo an encoding process via a video codec.
[0365] In some implementations, the preparation process may also include an editing process. During the editing process, further editing of the image / video data before / after projection can be performed. Similarly, even during the preparation process, metadata about stitching / projection / encoding / editing can be generated. Additionally, metadata about the initial view or region of interest (ROI) of the video data projected onto the 2D image can be generated.
[0366] The transmission process can be the process of processing and transmitting image / video data and metadata that have undergone a preparation process. For transmission, processing according to any transmission process can be performed. Data processed for transmission can be transmitted via broadcast networks and / or broadband. The data can be transmitted to the receiver on demand. The receiver can receive the data via various paths.
[0367] The processing procedure can refer to the process of decoding received data and reprojecting the projected image / video data onto a 3D model. In this process, image / video data projected onto a 2D image can be reprojected into 3D space. Depending on the context, this process can be called mapping or projection. In this case, the 3D space can have a shape that varies depending on the 3D model. For example, the 3D model can include a sphere, cube, cylinder, or cone.
[0368] In some implementations, the processing may further include an editing process, an upscaling process, etc. During the editing process, further editing of the image / video data before / after reprojection can be performed. When the image / video data is scaled down, its size can be increased during the upscaling process by scaling the sample upwards. If necessary, a downscaling operation can be performed to reduce the size.
[0369] The rendering process can refer to the process of rendering and displaying image / video data reprojected into 3D space. In some expressions, reprojection and rendering can be collectively expressed as rendering on a 3D model. Images / videos reprojected onto (or rendered on) a 3D model can have... Figure 22 The shape shown in 2230. Figure 22 Example 2230 illustrates a reprojection onto a spherical 3D model. Users can view a portion of the rendered image / video using a VR display. In this case, the area viewed by the user can have... Figure 22 The shape shown in 2240.
[0370] A feedback process can refer to the process by which a transmitter transmits various feedback information that can be obtained during the display process. Through the feedback process, interactivity can be provided in 360-degree video consumption. In some implementations, head orientation information and viewport information indicating the area the user is currently viewing can be transmitted to the transmitter during the feedback process. In some implementations, the user can interact with those interactions implemented in the VR environment. In this case, information related to the interaction can be transmitted to the transmitter or service provider during the feedback process. In some implementations, a feedback process may not be performed.
[0371] Head orientation information can refer to information about the position, angle, and movement of the user's head. Based on this information, information about the area the user is currently viewing in a 360-degree video (i.e., viewport information) can be calculated.
[0372] Viewport information can be information about the area a user is currently viewing in a 360-degree video. From this, gaze analysis can be performed to determine how the user consumes the 360-degree video or the extent to which the user is looking at a specific area of the 360-degree video. Gaze analysis can be performed by a receiver and transmitted to a transmitter via a feedback channel. Devices such as VR displays can extract the viewport area based on the user's head position / orientation, the vertical or horizontal field of view (FOV) information supported by the device, etc.
[0373] Furthermore, 360-degree video / images can be processed based on sub-pictures. Projected or packaged images, including 2D images, can be divided into sub-pictures and processed on a per-sub-picture basis. For example, high resolution can be provided to specific sub-pictures based on the user's viewport, or only specific sub-pictures can be encoded and signaled to the receiving device (decoding device). In this case, the decoding device can receive the sub-picture bitstream, reconstruct / decode the specific sub-picture, and perform rendering based on the user's viewport.
[0374] In some implementations, the feedback information described above can be transmitted not only to the transmitter but also consumed at the receiver. That is, the receiver's decoding, reprojection, and rendering processes can be performed using the feedback information. For example, head orientation information and / or viewport information can be used to preferentially decode and render only the 360-degree video of the area currently being viewed by the user.
[0375] Here, the viewport or viewport area can refer to the area viewed by the user in a 360-degree video. The viewpoint can be the point in the 360-degree video that the user is viewing, and can also refer to the center point of the viewport area. That is, the viewport is the area centered on the viewpoint, and the size and shape of the area can be determined by the field of view (FOV).
[0376] In the overall architecture used to provide 360-degree video, image / video data that undergoes a series of processes such as capture / projection / encoding / transmission / decoding / reprojection / rendering can be referred to as 360-degree video data. The term 360-degree video data may include metadata or signaling information associated with this image / video data.
[0377] Standardized media file formats can be defined for storing and transmitting media data such as audio or video. In some implementations, media files may have a file format based on the ISO Basic Media File Format (BMFF).
[0378] The scope of this disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) for enabling the operation of methods according to various embodiments to be executed on a device or computer, and non-transitory computer-readable media having such software or commands stored thereon and executable on a device or computer.
[0379] Industrial applicability
[0380] The embodiments disclosed herein can be used to encode or decode images.
Claims
1. An image decoding method performed by an image decoding device, the image decoding method comprising the following steps: Obtain inter-frame prediction and winding information for the current block from the bitstream; as well as The prediction block for the current block is generated based on the inter-frame prediction information and the winding information. The winding information includes a first flag specifying whether winding motion compensation is enabled for the current frame including the current block. Wherein, based on the fact that at least one of the sub-frames included in the current frame is independently encoded and has a width different from the width of the current frame, the first flag has a first value specifying that the scrolling motion compensation is disabled for the current frame. Specifically, the prediction block is generated by performing the winding motion compensation, based on the second value of the first flag specifying that the winding motion compensation is enabled for the current frame. Specifically, the winding motion compensation is performed based on whether the current sub-frame, including the current block, is independently encoded, and also based on the boundary of the current sub-frame or the boundary of the reference frame of the current block. The winding motion compensation is performed based on the boundaries of the current sub-picture, which includes the current block, and is based on the fact that the current sub-picture is independently encoded. The winding motion compensation is performed based on the boundary of the reference frame of the current block, since the current sub-frame is not independently encoded.
2. The image decoding method according to claim 1, wherein, The winding motion compensation for the current block is performed based on the winding offset, and the winding offset is obtained from the bitstream based on the first flag having the second value.
3. The image decoding method according to claim 2, wherein, The winding motion compensation for the current block is performed by modifying the x-coordinate of a reference block in the reference frame based on the winding offset, the reference block being specified by the motion vector of the current block.
4. The image decoding method according to claim 3, wherein, The winding motion compensation for the current block is performed by cropping the modified x-coordinate of the reference block to the boundary range of the current sub-screen or the reference screen.
5. The image decoding method according to claim 1, wherein, Based on the scaling of the reference image, the winding motion compensation for the current block is skipped.
6. The image decoding method according to claim 1, wherein, Since the current sub-frame is independently encoded and has a width different from that of the current frame, the winding motion compensation for the current block is skipped.
7. The image decoding method according to claim 1, wherein, The winding motion compensation for the current block is performed based on the left and right boundary positions, which are set based on whether the current sub-screen is independently encoded.
8. The image decoding method according to claim 1, wherein, Based on enabling the wrap motion compensation for the current video sequence including the current block, each of the top-left position and width of the current sub-frame is set to a predetermined value.
9. An image encoding method performed by an image encoding device, the image encoding method comprising the following steps: Determine whether the winding motion compensation is applied to the current block; A predicted block for the current block is generated by performing inter-frame prediction based on determining whether the winding motion compensation is applied to the current block; as well as The inter-frame prediction information of the current block and the winding information for the winding motion compensation are encoded. The winding information includes a first flag specifying whether winding motion compensation is enabled for the current frame including the current block. Wherein, based on the fact that at least one of the sub-frames included in the current frame is independently encoded and has a width different from the width of the current frame, the first flag has a first value specifying that the scrolling motion compensation is disabled for the current frame. Specifically, the winding motion compensation is performed based on whether the current sub-frame, including the current block, is independently encoded, and also based on the boundary of the current sub-frame or the boundary of the reference frame of the current block. The winding motion compensation is performed based on the boundaries of the current sub-picture, which includes the current block, and is based on the fact that the current sub-picture is independently encoded. The winding motion compensation is performed based on the boundary of the reference frame of the current block, since the current sub-frame is not independently encoded.
10. The image encoding method according to claim 9, wherein, The winding motion compensation for the current block is performed by modifying the x-coordinate of a reference block in the reference frame based on a predetermined winding offset. The reference block is specified by the motion vector of the current block.
11. The image encoding method according to claim 10, wherein, The winding motion compensation for the current block is performed by cropping the modified x-coordinate of the reference block to the boundary range of the current sub-screen or the reference screen.
12. The image encoding method according to claim 9, wherein, Since the current sub-frame is independently encoded and has a width different from that of the current frame, the winding motion compensation for the current block is skipped.
13. The image encoding method according to claim 9, wherein, The winding motion compensation for the current block is performed based on the left and right boundary positions, which are set based on whether the current sub-screen is independently encoded.
14. A method for transmitting a bit stream, the method comprising the following steps: Determine whether the winding motion compensation is applied to the current block; A predicted block for the current block is generated by performing inter-frame prediction based on determining whether the winding motion compensation is applied to the current block; The inter-frame prediction information of the current block and the winding information for the winding motion compensation are encoded into the bit stream; as well as Send the bit stream, The winding information includes a first flag specifying whether winding motion compensation is enabled for the current frame including the current block. Wherein, based on the fact that at least one of the sub-frames included in the current frame is independently encoded and has a width different from the width of the current frame, the first flag has a first value specifying that the scrolling motion compensation is disabled for the current frame. Specifically, the winding motion compensation is performed based on whether the current sub-frame, including the current block, is independently encoded, and also based on the boundary of the current sub-frame or the boundary of the reference frame of the current block. The winding motion compensation is performed based on the boundaries of the current sub-picture, which includes the current block, and is based on the fact that the current sub-picture is independently encoded. The winding motion compensation is performed based on the boundary of the reference frame of the current block, since the current sub-frame is not independently encoded.
Citation Information
Patent Citations
Methods and apparatus for flexible grid regions
WO2020056247A1