Method and apparatus for sub-picture based image encoding / decoding and method of transmitting bitstream
By determining whether the current block is treated as a frame during image encoding/decoding and applying BDOF or PROF, the problem of low efficiency in high-resolution image transmission and storage is solved, achieving efficient encoding/decoding and bitstream processing.
Patent Information
- Application Number
- CN202411781793.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-25
- Filing Date
- 2020-09-23
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2040-09-23
AI Technical Summary
Existing technologies are inefficient in the transmission and storage of high-resolution and high-quality images, leading to increased costs. Improved encoding/decoding methods and equipment are needed to increase efficiency.
By determining whether the current block is treated as a frame, bidirectional optical flow (BDOF) or optical flow prediction refinement (PROF) is applied to extract prediction samples from the reference frame, and refinement prediction is performed based on the motion information of the current block, using signal notifications and positions within the limiting range to extract prediction samples.
It improves the efficiency of image encoding/decoding, supports sub-picture-based encoding/decoding, and realizes the determination of BDOF or PROF and the transmission and storage of bit streams, thereby improving transmission and storage efficiency.
Smart Images

Figure CN119363974B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to image encoding / decoding methods and apparatus, and methods for transmitting bit streams; more specifically, it relates to an image encoding / decoding method and apparatus for performing sub-picture encoding / decoding, and a method for transmitting bit streams generated by the image encoding method / apparatus of this disclosure. Background Technology
[0002] Recently, there has been an increasing demand for high-resolution and high-quality images, such as high-definition (HD) and ultra-high-definition (UHD) images, across various fields. With the improvement in image data resolution and quality, the amount of information or bits transmitted increases relative to existing image data. This increase in the amount of information or bits transmitted leads to increased transmission and storage costs.
[0003] Therefore, efficient image compression techniques are needed to effectively transmit, store, and reproduce information about high-resolution and high-quality images. Summary of the Invention
[0004] Technical issues
[0005] The purpose of this disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0006] Another object of this disclosure is to provide an image encoding / decoding method and apparatus for encoding / decoding images based on sub-pictures.
[0007] Another object of this disclosure is to provide an image encoding / decoding method and apparatus for performing BDOF or PROF based on the determination of whether the current sub-picture is treated as a picture.
[0008] Another object of this disclosure is to provide a method for transmitting a bitstream generated by an image encoding method or device according to this disclosure.
[0009] Another object of this disclosure is to provide a recording medium for storing bitstreams generated by an image encoding method or apparatus according to this disclosure.
[0010] Another object of this disclosure is to provide a recording medium that stores a bitstream received, decoded and used to reconstruct an image by an image decoding device according to this disclosure.
[0011] The technical problems solved by this disclosure are not limited to those described above. Other technical problems not described herein will become clear to those skilled in the art through the following description.
[0012] Technical solution
[0013] An image decoding method according to one aspect of this disclosure may include: determining whether bidirectional optical flow (BDOF) or optical flow prediction refinement (PROF) is applied to the current block; extracting prediction samples of the current block from a reference frame based on motion information of the current block based on the application of BDOF or PROF to the current block; and deriving refined prediction samples of the current block by applying BDOF or PROF to the current block based on the extracted prediction samples.
[0014] In the image decoding method disclosed herein, the step of extracting the prediction sample of the current block can be performed based on whether the current sub-picture, including the current block, is treated as a picture.
[0015] In the image decoding method disclosed herein, whether the current sub-picture is treated as a picture can be determined based on flag information signaled via a bitstream.
[0016] In the image decoding method disclosed herein, the flag information can be notified by signaling through a sequence parameter set (SPS).
[0017] In the image decoding method disclosed herein, the step of extracting the predicted sample of the current block can be performed based on the position of the predicted sample to be extracted, wherein the position of the predicted sample can be limited to a predetermined range.
[0018] In the image decoding method disclosed herein, the predetermined range can be specified by the boundary position of the current sub-picture, based on the current sub-picture being treated as the picture.
[0019] In the image decoding method disclosed herein, the position of the predicted sample to be extracted may include x-coordinates and y-coordinates. The x-coordinates may be limited to the range of the left and right boundaries of the current sub-frame, and the y-coordinates may be limited to the range of the upper and lower boundaries of the current sub-frame.
[0020] In the image decoding method disclosed herein, the left boundary position of the current sub-screen can be derived as the product of the position information of a predetermined unit at the left position of the current sub-screen and the width of the predetermined unit; the right boundary position of the current sub-screen can be derived by performing a "-1" operation on the product of the position information of a predetermined unit at the right position of the current sub-screen and the width of the predetermined unit; the upper boundary position of the current sub-screen can be derived as the product of the position information of a predetermined unit at the upper position of the current sub-screen and the height of the predetermined unit; and the lower boundary position of the current sub-screen can be derived by performing a "-1" operation on the product of the position information of a predetermined unit at the lower position of the current sub-screen and the height of the predetermined unit.
[0021] In the image decoding method disclosed herein, the predetermined unit can be a grid or a CTU.
[0022] In the image decoding method disclosed herein, since the current sub-picture is not treated as a picture, the predetermined range can be the range of the current picture that includes the current block.
[0023] An image decoding apparatus according to another aspect of this disclosure may include a memory and at least one processor. The at least one processor may determine whether bidirectional optical flow (BDOF) or optical flow prediction refinement (PROF) is applied to the current block; based on the application of BDOF or PROF to the current block, extract prediction samples of the current block from a reference frame of the current block based on motion information of the current block; and derive refined prediction samples of the current block by applying BDOF or PROF to the current block based on the extracted prediction samples.
[0024] In the image decoding device of this disclosure, the at least one processor can extract a prediction sample of the current block based on whether the current sub-picture, including the current block, is treated as a picture.
[0025] An image coding method according to another aspect of this disclosure may include: determining whether bidirectional optical flow (BDOF) or optical flow prediction refinement (PROF) is applied to the current block; extracting prediction samples of the current block from a reference frame based on motion information of the current block based on the application of BDOF or PROF to the current block; and deriving refined prediction samples of the current block by applying BDOF or PROF to the current block based on the extracted prediction samples.
[0026] In the image encoding method disclosed herein, the step of extracting the prediction sample of the current block can be performed based on whether the current sub-picture, including the current block, is treated as a picture.
[0027] According to another aspect of the transmission method of this disclosure, a bit stream generated by the image encoding method and / or the image encoding device of this disclosure can be sent to an image decoding device.
[0028] Furthermore, according to another aspect of this disclosure, a computer-readable recording medium can store a bitstream generated by the image encoding device or image encoding method of this disclosure.
[0029] The features described above in this brief overview are merely exemplary aspects of the following detailed description of this disclosure and do not limit the scope of this disclosure.
[0030] Beneficial effects
[0031] According to this disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.
[0032] Furthermore, according to this disclosure, an image encoding / decoding method and apparatus for encoding / decoding images based on sub-pictures can be provided.
[0033] Furthermore, according to this disclosure, an image encoding / decoding method and apparatus for performing BDOF or PROF based on determining whether the current sub-picture is treated as a picture can be provided.
[0034] Furthermore, according to this disclosure, a method for transmitting a bitstream generated by an image encoding method or device according to this disclosure can be provided.
[0035] Furthermore, according to this disclosure, it is possible to provide a recording medium for storing a bitstream generated by an image encoding method or apparatus according to this disclosure.
[0036] Furthermore, according to this disclosure, a recording medium can be provided that stores a bitstream received, decoded, and used to reconstruct an image by an image decoding device according to this disclosure.
[0037] Those skilled in the art will understand that the effects achievable through this disclosure are not limited to those specifically described above, and that other advantages of this disclosure will become clearer from the detailed description. Attached Figure Description
[0038] Figure 1 This is a view that schematically illustrates a video coding system to which embodiments of the present disclosure are applicable.
[0039] Figure 2 This is a view schematically illustrating an image encoding device to which embodiments of the present disclosure are applicable.
[0040] Figure 3 This is a view schematically illustrating an image decoding device to which embodiments of the present disclosure are applicable.
[0041] Figure 4 This is a flowchart illustrating a video / image coding method based on inter-frame prediction.
[0042] Figure 5 This is a view illustrating the configuration of the inter-frame prediction unit 180 according to this disclosure.
[0043] Figure 6 This is a flowchart illustrating a video / image decoding method based on inter-frame prediction.
[0044] Figure 7 This is a view illustrating the configuration of the inter-frame prediction unit 260 according to this disclosure.
[0045] Figure 8 It is a view that illustrates the motion that can be expressed in affine mode.
[0046] Figure 9 This is a view that represents the parametric model of the affine mode.
[0047] Figure 10This is a view that demonstrates the method for generating the affine merge candidate list.
[0048] Figure 11 This is a view illustrating the CPMV derived from neighboring blocks.
[0049] Figure 12 This is a view that illustrates neighboring blocks used to derive inheritance affine merge candidates.
[0050] Figure 13 This is an example view used to derive neighboring blocks for constructing affine merge candidates.
[0051] Figure 14 This is a view that demonstrates the methods for generating a list of affine MVP candidates.
[0052] Figure 15 This is a view of neighboring blocks in a TMVP pattern based on sub-blocks.
[0053] Figure 16 This is a view illustrating a method for deriving motion vector fields based on a sub-block-based TMVP pattern.
[0054] Figure 17 This is a view of a CU that is instantiated to perform BDOF.
[0055] Figure 18 This is a view illustrating the relationship between Δv(i,j), v(i,j), and the sub-block motion vectors.
[0056] Figure 19 This is a view illustrating an implementation of the syntax for signaling sub-screen syntax elements in SPS.
[0057] Figure 20 This is a view illustrating an implementation of an algorithm for deriving predetermined variables such as SubPicTop.
[0058] Figure 21 This is a view illustrating a method for encoding an image using a sub-screen by an encoding device according to an embodiment.
[0059] Figure 22 This is a view illustrating a method for decoding an image using a sub-screen by a decoding device according to an embodiment.
[0060] Figure 23 This is a view illustrating the process of deriving the predicted samples of the current block by applying BDOF.
[0061] Figure 24 This is a view illustrating the inputs and outputs of a BDOF process according to an embodiment of this disclosure.
[0062] Figure 25This is a view illustrating variables used for BDOF processing according to embodiments of this disclosure.
[0063] Figure 26 This is a view illustrating a method for generating prediction samples of each sub-block in the current CU based on whether or not BDOF is applied, according to an embodiment of the present disclosure.
[0064] Figure 27 This is a view illustrating a method for deriving the gradient, autocorrelation, and cross-correlation of the current sub-block according to an embodiment of this disclosure.
[0065] Figure 28 This is a view illustrating a method for deriving motion refinement (vx,vy) according to an embodiment of the present disclosure, deriving BDOF offset, and generating a prediction sample for the current sub-block.
[0066] Figure 29 This is a view illustrating the process of deriving the predicted samples for the current block by applying PROF.
[0067] Figure 30 This is a view illustrating an example of PROF processing according to this disclosure.
[0068] Figure 31 This is a view that illustrates the case where a reference sample needs to be extracted across the boundaries of a sub-screen.
[0069] Figure 32 yes Figure 31 A magnified view of the extracted region.
[0070] Figure 33 This is a view illustrating a reference sample extraction process according to an embodiment of the present disclosure.
[0071] Figure 34 This is a flowchart illustrating the extraction process of a reference sample according to this disclosure.
[0072] Figure 35 This is a view illustrating a portion of the fractional sample interpolation process according to this disclosure.
[0073] Figure 36 This is a view illustrating a portion of the sbTMVP derivation method according to this disclosure.
[0074] Figure 37 This is a view illustrating a method for deriving the position of a sub-screen boundary according to this disclosure.
[0075] Figure 38 This is a view illustrating the content streaming system to which embodiments of this disclosure are applicable. Detailed Implementation
[0076] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings to facilitate implementation by those skilled in the art. However, this disclosure can be implemented in various different forms and is not limited to the embodiments described herein.
[0077] In describing this disclosure, detailed descriptions of relevant known functions or constructions will be omitted if they unnecessarily obscure the scope of this disclosure. In the accompanying drawings, portions irrelevant to the description of this disclosure are omitted, and similar reference numerals are assigned to similar portions.
[0078] In this disclosure, when a component is "connected," "linked," or "coupled" to another component, it may include not only direct connections but also indirect connections where intermediate components exist. Furthermore, when a component "comprises" or "has" other components, unless otherwise stated, it means that other components may be included, not excluded.
[0079] In this disclosure, the terms first, second, etc., are used only for the purpose of distinguishing one component from other components and do not limit the order or importance of the components, unless otherwise stated. Accordingly, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.
[0080] In this disclosure, the components are distinguished from each other to clearly describe each feature, but this does not mean that the components must be separate. That is, multiple components may be integrated into a single hardware or software unit, or a single component may be distributed and implemented across multiple hardware or software units. Therefore, unless otherwise specified, implementations of these integrated or distributed components are included within the scope of this disclosure.
[0081] In this disclosure, the components described in the various embodiments are not necessarily essential components, and some components may be optional. Therefore, embodiments consisting of a subset of the components described in the embodiments are also included within the scope of this disclosure. Furthermore, embodiments that include other components besides those described in the various embodiments are also included within the scope of this disclosure.
[0082] This disclosure relates to the encoding and decoding of images. Unless redefined in this disclosure, the terms used herein may have the general meaning commonly used in the art to which this disclosure pertains.
[0083] In this disclosure, a "picture" generally refers to a unit representing an image within a specific time period, while a slice / tile is a coding unit that constitutes part of a picture. A picture can be composed of one or more slices / tiles. Furthermore, a slice / tile may include one or more coding tree units (CTUs).
[0084] In this disclosure, "pixel" or "pixel" can refer to the smallest unit that constitutes a frame (or image). Furthermore, "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or a pixel value, or it can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0085] In this disclosure, "unit" can refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information associated with that region. In some cases, the term "unit" may be used interchangeably with terms such as "sample array," "block," or "region." Generally, an M×N block may include a set (or array) of samples (or transform coefficients) with M columns and N rows.
[0086] In this disclosure, "current block" can mean one of "current coding block," "current coding unit," "coding target block," "decoding target block," or "processing target block." When performing prediction, "current block" can mean "current prediction block" or "prediction target block." When performing transform (inverse transform) / quantization (dequantization), "current block" can mean "current transform block" or "transform target block." When performing filtering, "current block" can mean "filter target block."
[0087] In this disclosure, the terms “ / ” or “,” can be interpreted as indicating “and / or”. For example, “A / B” and “A, B” can mean “A and / or B”. Furthermore, “A / B / C” and “A / B / C” can mean “at least one of A, B and / or C”.
[0088] In this disclosure, the term "or" should be interpreted to indicate "and / or". For example, the expression "A or B" may include 1) only "A", 2) only "B", or 3) both "A and B". In other words, in this disclosure, "or" should be interpreted to indicate "additionally or alternatively".
[0089] Overview of Video Encoding Systems
[0090] Figure 1 This is a schematic view of a video encoding system according to the present disclosure.
[0091] The video encoding system according to the embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may deliver encoded video and / or image information or data to the decoding device 20 in the form of a file or stream via a digital storage medium or network.
[0092] The encoding device 10 according to an embodiment may include a video source generator 11, an encoding unit 12, and a transmitter 13. The decoding device 20 according to an embodiment may include a receiver 21, a decoding unit 22, and a renderer 23. The encoding unit 12 may be referred to as a video / image encoding unit, and the decoding unit 22 may be referred to as a video / image decoding unit. The transmitter 13 may be included in the encoding unit 12. The receiver 21 may be included in the decoding unit 22. The renderer 23 may include a display, and the display may be configured as a separate device or an external component.
[0093] The video source generator 11 can acquire video / images through a process of capturing, compositing, or generating video / images. The video source generator 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device may include, for example, a computer, tablet computer, and smartphone, and can generate video / images (electronically). For example, virtual video / images can be generated by a computer, etc. In this case, the video / image capture process can be replaced by a process of generating related data.
[0094] The encoding unit 12 can encode the input video / image. For compression and encoding efficiency, the encoding unit 12 can perform a series of processes, such as prediction, transformation, and quantization. The encoding unit 12 can output encoded data (encoded video / image information) in the form of a bitstream.
[0095] Transmitter 13 can transmit encoded video / image information or data, output in bitstream form, to receiver 21 of decoding device 20 in the form of a file or stream via digital storage medium or network. Digital storage medium can include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. Transmitter 13 can include elements for generating media files according to a predetermined file format and may include elements for transmission via broadcast / communication networks. Receiver 21 can extract / receive bitstreams from storage medium or network and transmit the bitstreams to decoding unit 22.
[0096] The decoding unit 22 can decode video / images by performing a series of processes corresponding to the operations of the encoding unit 12, such as dequantization, inverse transform, and prediction.
[0097] Renderer 23 can render decoded video / images. The rendered video / images can be displayed on a monitor.
[0098] Overview of Image Encoding Devices
[0099] Figure 2 This is a schematic view illustrating an image encoding device to which embodiments of this disclosure may be applied.
[0100] like Figure 2 As shown, the image encoding device 100 may include an image segmenter 110, a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter-frame prediction unit 180, an intra-frame prediction unit 185, and an entropy encoder 190. The inter-frame prediction unit 180 and the intra-frame prediction unit 185 may be collectively referred to as "prediction units". The transformer 120, quantizer 130, dequantizer 140, and inverse transformer 150 may be included in a residual processor. The residual processor may also include a subtractor 115.
[0101] In some implementations, all or at least some of the components configuring the image encoding device 100 may be configured by a single hardware component (e.g., an encoder or a processor). Furthermore, the memory 170 may include a decoded screen buffer (DPB) and may be configured by a digital storage medium.
[0102] Image segmenter 110 can segment an input image (or picture or frame) input to image encoding device 100 into one or more processing units. For example, a processing unit may be called an encoding unit (CU). Encoding units can be obtained by recursively segmenting encoding tree units (CTUs) or maximum encoding units (LCUs) according to a quadtree / binary tree / tritree (QT / BT / TT) structure. For example, an encoding unit can be segmented into multiple encoding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the segmentation of encoding units, a quadtree structure can be applied first, followed by a binary tree structure and / or a ternary tree structure. The encoding process according to this disclosure can be performed based on the final encoding unit that is no longer segmented. The maximum encoding unit can be used as the final encoding unit, or a deeper encoding unit obtained by segmenting the maximum encoding unit can be used as the final encoding unit. Here, the encoding process may include prediction, transformation, and reconstruction processes, which will be described later. As another example, the processing unit of the encoding process may be a prediction unit (PU) or a transformation unit (TU). Prediction units and transform units can be partitioned or segmented from the final coding unit. Prediction units can be sample prediction units, and transform units can be units used to derive transform coefficients and / or units used to derive residual signals from transform coefficients.
[0103] The prediction unit (inter-frame prediction unit 180 or intra-frame prediction unit 185) can perform prediction on the block to be processed (the current block) and generate a prediction block that includes prediction samples of the current block. The prediction unit can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. The prediction unit can generate various information related to the prediction of the current block and transmit the generated information to the entropy encoder 190. The information about the prediction can be encoded in the entropy encoder 190 and output as a bitstream.
[0104] Intra-prediction unit 185 can predict the current block by referencing samples in the current frame. Depending on the intra-prediction mode and / or intra-prediction technique, the reference samples may be located among the neighbors of the current block or may be placed separately. Intra-prediction modes may include multiple non-directional modes and multiple directional modes. Non-directional modes may include, for example, DC mode and planar mode. Depending on the level of detail in the prediction direction, directional modes may include, for example, 33 or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. Intra-prediction unit 185 can determine the prediction mode to be applied to the current block by using prediction modes applied to neighboring blocks.
[0105] The inter-frame prediction unit 180 can deduce the prediction block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference frame. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, dual prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. The reference frame including the reference block and the reference frame including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a juxtaposed reference block, a juxtaposed CU (colCU), etc. The reference frame including the temporally neighboring block may be referred to as a juxtaposed frame (colPic). For example, the inter-frame prediction unit 180 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate to use to deduce the motion vector and / or reference frame index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame prediction unit 180 can use motion information from neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, residual signals may not be transmitted. In motion vector prediction (MVP) mode, motion vectors from neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be signaled by encoding motion vector differences and indicators of the motion vector predictors. The motion vector difference can refer to the difference between the motion vector of the current block and the motion vector predictor.
[0106] The prediction unit can generate prediction signals based on various prediction methods and techniques described below. For example, the prediction unit can apply not only intra-frame prediction or inter-frame prediction, but also both intra-frame prediction and inter-frame prediction simultaneously to predict the current block. The prediction method that simultaneously applies both intra-frame prediction and inter-frame prediction to predict the current block can be called Combined Intra-Frame and Inter-Frame Prediction (CIIP). Furthermore, the prediction unit can perform Intra-Frame Block Copy (IBC) to predict the current block. Intra-Frame Block Copy can be used for content image / video coding in games, such as Screen Content Coding (SCC). IBC is a method of predicting the current frame using a previously reconstructed reference block in the current frame at a predetermined distance from the current block. When IBC is applied, the position of the reference block in the current frame can be encoded as a vector (block vector) corresponding to the predetermined distance.
[0107] The prediction signal generated by the prediction unit can be used to generate a reconstructed signal or a residual signal. Subtractor 115 can generate a residual signal (residual block or residual sample array) by subtracting the prediction signal (prediction block or prediction sample array) output from the prediction unit from the input image signal (original block or original sample array). The generated residual signal can be transmitted to converter 120.
[0108] Transformer 120 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform techniques may include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loève Transform (KLT), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented graphically. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. Furthermore, the transform processing can be applied to square pixel blocks of the same size or to blocks of variable size instead of square.
[0109] Quantizer 130 quantizes the transform coefficients and transmits them to entropy encoder 190. Entropy encoder 190 encodes the quantized signal (information about the quantized transform coefficients) and outputs a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 130 can rearrange the block-form quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and generate information about the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients.
[0110] The entropy encoder 190 can perform various encoding methods, such as exponential Columbus coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 190 can encode, either together or separately, the information required for video / image reconstruction other than the quantization transform coefficients (e.g., values of syntax elements). The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream at the Network Abstraction Layer (NAL). The video / image information may also include information about various parameter sets, such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Furthermore, the video / image information may also include general constraint information. The signaled information, transmitted information, and / or syntax elements described in this disclosure can be encoded and included in the bitstream through the above encoding process.
[0111] The bitstream can be transmitted over a network or stored in a digital storage medium. The network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not shown) for transmitting the signal output from the entropy encoder 190 and / or a storage unit (not shown) for storing the signal may be included as internal / external components of the image encoding device 100. Alternatively, a transmitter may be provided as a component of the entropy encoder 190.
[0112] The quantization transform coefficients output from quantizer 130 can be used to generate residual signals. For example, the residual signals (residual blocks or residual samples) can be reconstructed by applying dequantization and inverse transform to the quantization transform coefficients through dequantizer 140 and inverse transformer 150.
[0113] Adder 155 adds the reconstructed residual signal to the prediction signal output from inter-frame prediction unit 180 or intra-frame prediction unit 185 to generate a reconstructed signal (reconstructed frame, reconstructed block, reconstructed sample array). If the block to be processed has no residual, such as in the case of applying skip mode, the prediction block can be used as a reconstructed block. Adder 155 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current frame, and can be used for inter-frame prediction of the next frame by filtering as described below.
[0114] Furthermore, as described below, Luminance Mapping and Chromatography Scaling (LMCS) is suitable for image encoding processing.
[0115] Filter 160 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 160 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 170, specifically in the DPB of memory 170. Various filtering methods can include, for example, deblocking filtering, sample adaptive offsetting, adaptive loop filtering, bilateral filtering, etc. Filter 160 can generate various filtering-related information and transmit the generated information to entropy encoder 190, as described later in the description of each filtering method. The filtering-related information can be encoded by entropy encoder 190 and output as a bitstream.
[0116] The modified reconstructed frame transmitted to memory 170 can be used as a reference frame in inter-frame prediction unit 180. When inter-frame prediction is applied by image encoding device 100, prediction mismatch between image encoding device 100 and image decoding device can be avoided and coding efficiency can be improved.
[0117] The DPB of memory 170 can store modified reconstructed frames for use as reference frames in inter-frame prediction unit 180. Memory 170 can store motion information of blocks from which motion information in the current frame is derived (or encoded) and / or motion information of already reconstructed blocks in the frame. The stored motion information can be transmitted to inter-frame prediction unit 180 and used as motion information for spatially or temporally neighboring blocks. Memory 170 can store reconstructed samples of reconstructed blocks in the current frame and can transmit the reconstructed samples to intra-frame prediction unit 185.
[0118] Overview of image decoding devices
[0119] Figure 3 This is a schematic view illustrating an image decoding device to which embodiments of the present disclosure may be applied.
[0120] like Figure 3 As shown, the image decoding device 200 may include an entropy decoder 210, a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame prediction unit 260, and an intra-frame prediction unit 265. The inter-frame prediction unit 260 and the intra-frame prediction unit 265 may be collectively referred to as "prediction units". The dequantizer 220 and the inverse transformer 230 may be included in a residual processor.
[0121] According to an implementation, all or at least some of the components of the image decoding device 200 can be configured by hardware components (e.g., a decoder or a processor). Furthermore, the memory 250 may include a decoded screen buffer (DPB) or may be configured by a digital storage medium.
[0122] The image decoding device 200, having received a bitstream including video / image information, can perform operations related to... Figure 2 The image is reconstructed by processing corresponding to the processing performed by the image encoding device 100. For example, the image decoding device 200 can perform decoding using a processing unit applied in the image encoding device. Therefore, the decoding processing unit can be, for example, an encoding unit. The encoding unit can be obtained by segmenting a coding tree unit or a maximum coding unit. The reconstructed image signal decoded and output by the image decoding device 200 can be reproduced by a reproduction device (not shown).
[0123] Image decoding device 200 can receive data in bitstream form from... Figure 2The signal output by the image encoding device. The received signal can be decoded by the entropy decoder 210. For example, the entropy decoder 210 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets, such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). In addition, the video / image information may also include general constraint information. The image decoding device can also decode the picture based on the information about the parameter sets and / or general constraint information. The information and / or syntax elements notified / received by signals described in this disclosure can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 210 decodes the information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC, and outputs the values of the syntax elements required for image reconstruction and the quantized values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine the context model using information about the target syntax element, decoding information of neighboring blocks and the target block, or information about symbols / bins decoded in the previous stage, perform arithmetic decoding on the bins based on the determined context model by predicting the occurrence probability of the bins, and generate symbols corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method can update the context model after determining the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin. The prediction-related information in the information decoded by the entropy decoder 210 can be provided to the prediction units (inter-frame prediction unit 260 and intra-frame prediction unit 265), and the residual value of entropy decoding performed in the entropy decoder 210, i.e., the quantization transform coefficients and related parameter information, can be input to the dequantizer 220. In addition, the filtering information in the information decoded by the entropy decoder 210 can be provided to the filter 240. Furthermore, the receiver (not shown) for receiving signals output from the image encoding device may be further configured as an internal / external element of the image decoding device 200, or the receiver may be a component of the entropy decoder 210.
[0124] Furthermore, the image decoding apparatus according to this disclosure can be referred to as a video / image / screen decoding apparatus. The image decoding apparatus can be divided into an information decoder (video / image / screen information decoder) and a sample decoder (video / image / screen sample decoder). The information decoder may include an entropy decoder 210. The sample decoder may include at least one of a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame prediction unit 260, or an intra-frame prediction unit 265.
[0125] Dequantizer 220 can dequantize the quantized transform coefficients and output transform coefficients. Dequantizer 220 can rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement can be performed based on the coefficient scan order performed in the image encoding device. Dequantizer 220 can obtain transform coefficients by performing dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information).
[0126] The inverse transformer 230 can perform inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).
[0127] The prediction unit can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. The prediction unit can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on the prediction information output from the entropy decoder 210, and can determine a specific intra-frame / inter-frame prediction mode (prediction technique).
[0128] Similar to that described in the prediction unit of the image coding device 100, the prediction unit can generate a prediction signal based on various prediction methods (techniques) described later.
[0129] Intra-prediction unit 265 can predict the current block by referring to samples in the current frame. The description of intra-prediction unit 185 also applies to intra-prediction unit 265.
[0130] The inter-frame prediction unit 260 can deduce the prediction block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference frame. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, dual prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. For example, the inter-frame prediction unit 260 can configure a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference frame index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the inter-frame prediction mode of the current block.
[0131] Adder 235 can generate a reconstruction signal (reconstructed frame, reconstruction block, reconstruction sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including inter-frame prediction unit 260 and / or intra-frame prediction unit 265). The description of adder 155 also applies to adder 235.
[0132] Furthermore, as described below, Luminance Mapping and Chromaticity Scaling (LMCS) is suitable for image decoding processing.
[0133] Filter 240 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 240 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 250, specifically in the DPB of memory 250. Various filtering methods may include, for example, deblocking filtering, adaptive sample shifting, adaptive loop filtering, bilateral filtering, etc.
[0134] The (modified) reconstructed frame stored in the DPB of memory 250 can be used as a reference frame in inter-frame prediction unit 260. Memory 250 can store motion information of blocks from which motion information in the current frame is derived (or decoded) and / or motion information of already reconstructed blocks in the frame. The stored motion information can be transmitted to inter-frame prediction unit 260 to be used as motion information for spatially or temporally neighboring blocks. Memory 250 can store reconstructed samples of reconstructed blocks in the current frame and transmit the reconstructed samples to intra-frame prediction unit 265.
[0135] In this disclosure, the embodiments described in the filter 160, inter-frame prediction unit 180 and intra-frame prediction unit 185 of the image encoding device 100 can be equally or correspondingly applied to the filter 240, inter-frame prediction unit 260 and intra-frame prediction unit 265 of the image decoding device 200.
[0136] Overview of inter-frame prediction
[0137] Image encoding / decoding devices can perform inter-frame prediction on a block-by-block basis to derive prediction samples. Inter-frame prediction can refer to predictions derived in a way that depends on data elements from frames other than the current frame. When inter-frame prediction is applied to the current block, the prediction block for the current block can be derived based on a reference block in a reference frame specified by motion vectors.
[0138] In this context, to reduce the amount of motion information transmitted in inter-frame prediction mode, the motion information of the current block can be derived based on the correlation between the motion information of neighboring blocks and the current block. This motion information can be derived on a block, sub-block, or sample basis. Motion information may include motion vectors and reference frame indices. It may also include inter-frame prediction type information. Here, inter-frame prediction type information can refer to the direction information of inter-frame prediction. Inter-frame prediction type information can indicate whether to use L0 prediction, L1 prediction, or dual prediction to predict the current block.
[0139] When inter-frame prediction is applied to the current block, the neighboring blocks of the current block can include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in a reference frame. The reference frame that includes the reference block of the current block and the reference frame that includes the temporally neighboring block can be the same or different. The temporally neighboring block can be referred to as the juxtaposed reference block or juxtaposed CU (colCU), and the reference frame that includes the temporally neighboring block can be referred to as the juxtaposed frame (colPic).
[0140] Furthermore, a candidate list of motion information can be constructed based on the neighboring blocks of the current block, and in this case, a signal can be used to indicate which candidate's flag or index information to use in order to derive the motion vector and / or reference screen index of the current block.
[0141] Depending on the inter-frame prediction type, motion information can include L0 motion information and / or L1 motion information. The motion vector in the L0 direction can be defined as the L0 motion vector or MVL0, and the motion vector in the L1 direction can be defined as the L1 motion vector or MVL1. Prediction based on the L0 motion vector can be defined as L0 prediction, prediction based on the L1 motion vector can be defined as L1 prediction, and prediction based on both L0 and L1 motion vectors can be defined as dual prediction. Here, the L0 motion vector can refer to the motion vector associated with the reference frame list L0, and the L1 motion vector can refer to the motion vector associated with the reference frame list L1.
[0142] Reference screen list L0 can include screens that precede the current screen in output order as reference screens, and reference screen list L1 can include screens that follow the current screen in output order. Previous screens can be defined as forward (reference) screens, and subsequent screens can be defined as backward (reference) screens. Furthermore, reference screen list L0 can also include screens that follow the current screen in output order as reference screens. In this case, within reference screen list L0, previous screens can be indexed first, and then subsequent screens can be indexed. Reference screen list L1 can also include screens that precede the current screen in output order as reference screens. In this case, within reference screen list L1, subsequent screens can be indexed first, and then previous screens can be indexed. Here, the output order can correspond to the screen order count (POC) order.
[0143] Figure 4 This is a flowchart illustrating a video / image coding method based on inter-frame prediction.
[0144] Figure 5 This is a view illustrating the configuration of the inter-frame predictor 180 according to this disclosure.
[0145] Figure 4 The encoding method can be determined by Figure 2 The image encoding device performs the following steps: Specifically, step S410 can be performed by the inter-frame predictor 180, and step S420 can be performed by the residual processor. Specifically, step S420 can be performed by the subtractor 115. Step S430 can be performed by the entropy encoder 190. The prediction information of step S430 can be derived by the inter-frame predictor 180, and the residual information of step S430 can be derived by the residual processor. The residual information is information about the residual samples. The residual information may include information about the quantization transform coefficients used for the residual samples. As described above, the residual samples can be derived into transform coefficients by the transformer 120 of the image encoding device, and the transform coefficients can be derived into quantization transform coefficients by the quantizer 130. The information about the quantization transform coefficients can be encoded by the entropy encoder 190 through the residual encoding process.
[0146] The image coding device can perform inter-frame prediction for the current block (S410). The image coding device can deduce the inter-frame prediction mode and motion information for the current block and generate prediction samples for the current block. Here, the inter-frame prediction mode determination, motion information deduction, and prediction sample generation processes can be performed simultaneously, or any one of them can be performed before other processes. For example, as... Figure 5 As shown, the inter-frame prediction unit 180 of the image coding apparatus may include a prediction mode determination unit 181, a motion information derivation unit 182, and a prediction sample derivation unit 183. The prediction mode determination unit 181 determines the prediction mode for the current block, the motion information derivation unit 182 derives the motion information for the current block, and the prediction sample derivation unit 183 derives the prediction samples for the current block. For example, the inter-frame prediction unit 180 of the image coding apparatus can search for blocks similar to the current block within a predetermined region (search region) of a reference frame using motion estimation, and derive a reference block whose difference from the current block is equal to or less than a predetermined criterion or minimum value. Based on this, a reference frame index indicating the reference frame in which the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The image coding apparatus can determine the mode applicable to the current block among various inter-frame prediction modes. The image coding apparatus can compare rate distortion (RD) costs for various prediction modes and determine the optimal inter-frame prediction mode for the current block. However, the methods by which an image coding device determines the inter-frame prediction mode of the current block are not limited to the examples above, and various methods can be used.
[0147] For example, the inter-frame prediction mode of the current block can be determined as at least one of the following: merge mode, merge skip mode, motion vector prediction (MVP) mode, symmetric motion vector difference (SMVD) mode, affine mode, sub-block based merge mode, adaptive motion vector resolution (AMVR) mode, history-based motion vector predictor (HMVP) mode, pairwise average merge mode, merge mode with motion vector difference (MMVD) mode, decoder-side motion vector refinement (DMVR) mode, combined inter-frame and intra-frame prediction (CIIP) mode, or geometric segmentation mode (GPM).
[0148] For example, when a skip mode or merge mode is applied to the current block, the image encoding device can deduce merge candidates from neighboring blocks of the current block and use the deduced merge candidates to construct a merge candidate list. Alternatively, the image encoding device can deduce reference blocks from among the reference blocks indicated by the merge candidates included in the merge candidate list, whose difference from the current block is equal to or less than a predetermined criterion or minimum value. In this case, a merge candidate associated with the deduced reference block can be selected, and merge index information indicating the selected merge candidate can be generated and signaled to the image decoding device. Motion information of the current block can be deduced using the motion information of the selected merge candidate.
[0149] As another example, when the MVP mode is applied to the current block, the image encoding device can derive motion vector predictor (MVP) candidates from the neighboring blocks of the current block and use the derived MVP candidates to construct an MVP candidate list. Alternatively, the image encoding device can use the motion vector of an MVP candidate selected from the MVP candidates included in the MVP candidate list as the MVP of the current block. In this case, for example, the motion vector of a reference block derived through the above motion estimation can be used as the motion vector of the current block, and the MVP candidate with the motion vector having the smallest difference from the motion vector of the current block can be the selected MVP candidate. The motion vector difference (MVD) can be derived as the difference obtained by subtracting the MVP from the motion vector of the current block. In this case, the index information indicating the selected MVP candidate and information about the MVD can be signaled to the image decoding device. Additionally, when the MVP mode is applied, the value of the reference frame index can be constructed as reference frame index information and signaled separately to the image decoding device.
[0150] The image coding device can derive residual samples based on the predicted samples (S420). The image coding device can derive residual samples by comparing the original samples of the current block with the predicted samples. For example, residual samples can be derived by subtracting the corresponding predicted samples from the original samples.
[0151] The image encoding device can encode image information including prediction information and residual information (S430). The image encoding device can output the encoded image information in the form of a bitstream. The prediction information can include prediction mode information (e.g., skip flag, merge flag, or mode index, etc.) and information about motion information as information related to the prediction process. In the prediction mode information, the skip flag indicates whether the skip mode is applied to the current block, and the merge flag indicates whether the merge mode is applied to the current block. Alternatively, the prediction mode information can indicate one of a variety of prediction modes, such as the mode index. When the skip flag and the merge flag are 0, it can be determined that the MVP mode is applied to the current block. The information about motion information can include candidate selection information (e.g., merge index, MVP flag, or MVP index) as information for deriving motion vectors. In the candidate selection information, the merge index can be signaled when the merge mode is applied to the current block, and can be information for selecting one of the merge candidates included in the merge candidate list. In the candidate selection information, the MVP flag or MVP index can be signaled when the MVP mode is applied to the current block, and can be information for selecting one of the MVP candidates in the MVP candidate list. Specifically, the MVP flag can be signaled using the syntax elements mvp_10_flag or mvp_11_flag. Additionally, motion information can include information about the aforementioned MVD and / or reference frame index information. Furthermore, motion information can include indications of whether L0 prediction, L1 prediction, or dual prediction is applied. Residual information is about the residual samples. Residual information can include information about the quantization transform coefficients used for the residual samples.
[0152] The output bitstream can be stored in (digital) storage media and sent to an image decoding device or it can be sent to an image decoding device via a network.
[0153] As described above, the image coding device can generate a reconstructed frame (including a frame of reconstructed samples and reconstructed blocks) based on reference samples and residual samples. This is to allow the image coding device to derive the same prediction result as that performed by the image decoding device, thereby improving coding efficiency. Therefore, the image coding device can store the reconstructed frame (or reconstructed samples and reconstructed blocks) in memory and use it as a reference frame for inter-frame prediction. As mentioned above, the in-loop filtering process is also applied to the reconstructed frame.
[0154] Figure 6 This is a flowchart illustrating a video / image decoding method based on inter-frame prediction.
[0155] Figure 7 This is a view illustrating the configuration of the inter-frame prediction unit 260 according to this disclosure.
[0156] An image decoding device can perform operations corresponding to those performed by an image encoding device. The image decoding device can perform predictions for the current block and derive prediction samples based on received prediction information.
[0157] Figure 6 The decoding method can be derived from Figure 3 The image decoding device performs the steps S610 to S630. Steps S610 to S630 can be performed by the inter-frame prediction unit 260, and the prediction information of step S610 and the residual information of step S640 can be obtained from the bitstream by the entropy decoder 210. The residual processor of the image decoding device can derive the residual samples of the current block based on the residual information (S640). Specifically, the dequantizer 220 of the residual processor can perform dequantization based on the quantization transform coefficients derived from the residual information to derive the transform coefficients, and the inverse transformer 230 of the residual processor can perform an inverse transform on the transform coefficients to derive the residual samples of the current block. Step S650 can be performed by the adder 235 or the reconstructor.
[0158] Specifically, the image decoding device can determine the prediction mode of the current block based on the received prediction information (S610). The image decoding device can determine which inter-frame prediction mode is applied to the current block based on the prediction mode information in the prediction information.
[0159] For example, a skip flag can be used to determine whether a skip mode applies to the current block. Alternatively, a merge flag can be used to determine whether a merge mode or an MVP mode applies to the current block. Alternatively, a mode index can be used to select one of various inter-frame prediction mode candidates. Inter-frame prediction mode candidates may include skip mode, merge mode, and / or MVP mode, or may include various inter-frame prediction modes as described below.
[0160] The image decoding device can deduce motion information for the current block based on the determined inter-frame prediction mode (S620). For example, when a skip mode or merge mode is applied to the current block, the image decoding device can construct a merge candidate list (described below) and select one of the merge candidates included in the merge candidate list. The selection can be performed based on the aforementioned candidate selection information (merge index). The motion information of the current block can be deduced using the motion information of the selected merge candidate. For example, the motion information of the selected merge candidate can be used as the motion information of the current block.
[0161] As another example, when the MVP mode is applied to the current block, the image decoding device can construct an MVP candidate list and use the motion vector of the MVP candidate selected from the MVP candidates included in the MVP candidate list as the MVP of the current block. Selection can be performed based on the aforementioned candidate selection information (MVP flag or MVP index). In this case, the MVD of the current block can be derived based on information about the MVD, and the motion vector of the current block can be derived based on the MVP and MVD of the current block. Additionally, the reference frame index of the current block can be derived based on reference frame index information. The frame indicated by the reference frame index in the reference frame list of the current block can be derived as the reference frame to be referenced for inter-frame prediction of the current block.
[0162] The image decoding device can generate a prediction sample for the current block based on the motion information of the current block (S630). In this case, a reference frame can be derived based on the reference frame index of the current block, and the prediction sample for the current block can be derived using samples of the reference block indicated by the motion vector of the current block on the reference frame. In some cases, a prediction sample filtering process can also be performed on all or some of the prediction samples of the current block.
[0163] For example, such as Figure 7 As shown, the inter-frame prediction unit 260 of the image decoding device may include a prediction mode determination unit 261, a motion information derivation unit 262, and a prediction sample derivation unit 263. In the inter-frame prediction unit 260 of the image decoding device, the prediction mode determination unit 261 can determine the prediction mode of the current block based on the received prediction mode information, the motion information derivation unit 262 can derive the motion information (motion vector and / or reference frame index, etc.) of the current block based on the received motion information, and the prediction sample derivation unit 263 can derive the prediction samples of the current block.
[0164] The image decoding device can generate residual samples for the current block based on the received residual information (S640). The image decoding device can generate reconstructed samples for the current block based on the predicted samples and residual samples, and generate a reconstructed image based on this (S650). Thereafter, the in-loop filtering process is applied to the reconstructed image as described above.
[0165] As described above, the inter-frame prediction process may include the steps of determining an inter-frame prediction mode, deriving motion information based on the determined prediction mode, and performing prediction (generating prediction samples) based on the derived motion information. As described above, the inter-frame prediction process may be performed by an image encoding device and an image decoding device.
[0166] The steps for deriving motion information based on the prediction pattern will be described in more detail below.
[0167] As described above, inter-frame prediction can be performed using the motion information of the current block. The image coding device can derive the optimal motion information for the current block through a motion estimation process. For example, the image coding device can use original blocks in the original frame of the current block on a fractional pixel basis to search for highly correlated similar reference blocks within a predetermined search range in the reference frame, and use these to derive motion information. The similarity of the blocks can be calculated based on the sum of absolute differences (SAD) between the current block and the reference blocks. In this case, motion information can be derived based on the reference block with the smallest SAD in the search area. The derived motion information can be signaled to the image decoding device according to various methods based on inter-frame prediction modes.
[0168] When a merge mode is applied to the current block, the motion information of the current block is not sent directly; instead, the motion information of neighboring blocks is used to deduce the motion information of the current block. Therefore, the motion information of the current prediction block can be indicated by sending flag information indicating that a merge mode is used and candidate selection information (e.g., a merge index) indicating which neighboring block is used as a merge candidate. In this disclosure, since the current block is a prediction execution unit, the current block can be used to mean the same thing as the current prediction block, and neighboring blocks can be used to mean the same thing as neighboring prediction blocks.
[0169] Image encoding devices can search for candidate blocks to merge to derive motion information for the current block in order to execute a merge mode. For example, up to five candidate blocks can be used, but this is not a limitation. A maximum number of candidate blocks can be sent in the slice header or tile group header, but this is not a limitation. After finding candidate blocks, the image encoding device can generate a list of candidate blocks and select the candidate block with the lowest RD cost as the final candidate block.
[0170] The merge candidate list can use, for example, five merge candidate blocks. For instance, four spatial merge candidates and one temporal merge candidate could be used.
[0171] Overview of Affine Modes
[0172] The following section describes affine patterns as an example of inter-frame prediction modes in detail. In traditional video coding / decoding systems, only one motion vector is used to represent the motion information of the current block (translation motion model). However, in traditional methods, optimal motion information is expressed only at the block level, but optimal motion information cannot be expressed at the pixel level. To address this issue, affine motion modes that define block motion information at the pixel level have been proposed. According to the affine mode, the motion vectors of individual pixels and / or sub-block units of the block can be determined using two to four motion vectors associated with the current block.
[0173] Compared to existing motion information expressed using translation (or displacement) of pixel values, in affine mode, the motion information of individual pixels can be expressed using at least one of translation, scaling, rotation, or shearing.
[0174] Figure 8 It is a view that illustrates the motion that can be expressed in affine mode.
[0175] exist Figure 8 In the motion shown, the affine pattern used to express the motion information of each pixel through translation, scaling, or rotation can be a similar or simplified affine pattern. The affine patterns described below may refer to similar or simplified affine patterns.
[0176] Motion information in affine mode can be represented using two or more control point motion vectors (CPMVs). CPMVs can be used to derive the motion vector for a specific pixel position in the current block. In this case, the set of motion vectors for each pixel and / or sub-block of the current block can be defined as an affine motion vector field (affine MVF).
[0177] Figure 9 This is a view that represents the parametric model of the affine mode.
[0178] When an affine pattern is applied to the current block, an affine MVF can be derived using either a 4-parameter model or a 6-parameter model. In this case, a 4-parameter model can refer to a model type using two CPMVs, and a 6-parameter model can refer to a model type using three CPMVs. Figure 9 (a) and Figure 9 (b) shows the CPMV used in the 4-parameter model and the 6-parameter model, respectively.
[0179] When the current block position is (x, y), the motion vector based on the pixel position can be derived according to Equation 1 or Equation 2. For example, the motion vector based on the 4-parameter model can be derived according to Equation 1, and the motion vector based on the 6-parameter model can be derived according to Equation 2.
[0180] [Formula 1]
[0181]
[0182] [Equation 2]
[0183]
[0184] In Equations 1 and 2, mv0 = {mv_0x, mv_0y} can be the CPMV at the top-left corner of the current block, mv1 = {mv_1x, mv_1y} can be the CPMV at the top-right corner of the current block, and mv2 = {mv_2x, mv_2y} can be the CPMV at the bottom-left corner of the current block. In this case, W and H correspond to the width and height of the current block, respectively, and mv = {mv_x, mv_y} can refer to the motion vector of pixel position {x, y}.
[0185] In encoding / decoding processing, the affine MVF can be determined on a pixel and / or predefined sub-block basis. When determining the affine MVF on a pixel basis, motion vectors can be derived based on individual pixel values. Furthermore, when determining the affine MVF on a sub-block basis, the motion vector of the corresponding block can be derived based on the center pixel value of the sub-block. The center pixel value can refer to a virtual pixel located at the center of the sub-block or the bottom-right pixel among the four pixels at the center. Additionally, the center pixel value can be a specific pixel within the sub-block or a pixel representing the sub-block. In this disclosure, the case of determining the affine MVF on a 4×4 sub-block basis will be described. However, this is only for ease of description, and the size of the sub-blocks can be varied.
[0186] That is, when affine prediction is available, the motion model applicable to the current block can include three models: a translational motion model, a 4-parameter affine motion model, and a 6-parameter affine motion model. Here, the translational motion model can represent the model used by the existing block unit motion vectors, the 4-parameter affine motion model can represent the model used by two CPMVs, and the 6-parameter affine motion model can represent the model used by three CPMVs. Affine modes can be divided into detailed modes based on the motion information encoding / decoding method. For example, affine modes can be further divided into affine MVP mode and affine merging mode.
[0187] When an affine merge mode is applied to the current block, the CPMV can be derived from neighboring blocks of the current block that are encoded / decoded in affine mode. An affine merge mode can be applied to the current block when at least one of its neighboring blocks is encoded / decoded in affine mode. That is, when an affine merge mode is applied to the current block, the CPMV of the current block can be derived using the CPMVs of neighboring blocks. For example, the CPMV of a neighboring block can be determined as the CPMV of the current block, or the CPMV of the current block can be derived based on the CPMVs of neighboring blocks. When deriving the CPMV of the current block based on the CPMVs of neighboring blocks, at least one encoding parameter of the current block or a neighboring block can be used. For example, the CPMV of a neighboring block can be modified based on the size of the neighboring blocks and the size of the current block and used as the CPMV of the current block.
[0188] Furthermore, affine merging that derives MV on a sub-block basis can be referred to as a sub-block merging pattern, which can be specified by a merge_subblock_flag with a first value (e.g., 1). In this case, the affine merging candidate list described below can be referred to as the sub-block merging candidate list. In this case, candidates deduced as sbTMVP can be further included in the sub-block merging candidate list. In this case, candidates deduced as sbTMVP can be used as candidates at index #0 of the sub-block merging candidate list. In other words, candidates deduced as sbTMVP can be placed before the inherited affine candidates and constructed affine candidates described below in the sub-block merging candidate list.
[0189] For example, an affine mode flag can be defined to specify whether an affine mode applies to the current block, which can be signaled at at least one higher level of the current block (e.g., sequence, picture, slice, tile, tile group, tile, etc.). For example, the affine mode flag can be named sps_affine_enabled_flag.
[0190] When applying affine merge mode, the affine merge candidate list can be configured to derive the CPMV of the current block. In this case, the affine merge candidate list can include at least one of inherited affine merge candidates, constructed affine merge candidates, or zero merge candidates. When the neighboring blocks of the current block are encoded / decoded in affine mode, inherited affine merge candidates can refer to candidates derived using the CPMV of neighboring blocks. Constructed affine merge candidates can refer to candidates that derive the individual CPMVs based on the motion vectors of neighboring blocks at each control point (CP). Furthermore, zero merge candidates can refer to candidates consisting of CPMVs of size 0. In the following description, CP can refer to a specific location of the block used to derive the CPMV. For example, CP can be the individual vertex locations of the block.
[0191] Figure 10 This is a view that demonstrates the method for generating the affine merge candidate list.
[0192] Reference Figure 10 The flowchart shows that affine merge candidates can be added to the affine merge candidate list in the following order: inheriting affine merge candidates (S1210), constructing affine merge candidates (S1220), and zero merge candidates (S1230). If, even after all inherited and constructed affine merge candidates have been added to the affine merge candidate list, the number of candidates included in the list still does not meet the maximum number of candidates, zero merge candidates can be added. In this case, zero merge candidates can be added until the number of candidates in the affine merge candidate list meets the maximum number of candidates.
[0193] Figure 11This is a view illustrating the control point motion vector (CPMV) derived from the neighboring blocks.
[0194] For example, up to two inheritance affine merge candidates can be derived, each of which can be derived based on at least one of the left neighbor block and the upper neighbor block.
[0195] Figure 12 This is a view that is used to derive neighboring blocks for affine inheritance candidates.
[0196] Based on the inheritance affine merge candidate derived from the left neighbor block. Figure 12 The affine merge candidate derived from at least one of the neighboring blocks A0 or A1 can be based on the above neighboring blocks. Figure 12 The candidate is derived from at least one of the neighboring blocks B0, B1, or B2. In this case, the scan order of the neighboring blocks can be A0 to A1 and B0, B1, and B2, but is not limited to this. For each of the left and top neighboring blocks, the inherited affine merge candidate can be derived based on the first neighboring block available in the scan order. In this case, redundancy checks may not be performed between the candidates derived from the left neighboring block and the top neighboring block.
[0197] For example, such as Figure 11 As shown, when the left neighboring block A is encoded / decoded in an affine mode, at least one of the motion vectors v2, v3, and v4 corresponding to the CP of the neighboring block A can be derived. When the neighboring block A is encoded / decoded using a 4-parameter affine model, the inherited affine merge candidate can be derived using v2 and v3. In contrast, when the neighboring block A is encoded / decoded using a 6-parameter affine model, the inherited affine merge candidate can be derived using v2, v3, and v4.
[0198] Figure 13 This is an example view used to derive neighboring blocks for constructing affine merge candidates.
[0199] Constructing an affine candidate can mean having a candidate CPMV that is derived using the combined motion information of neighboring blocks. The motion information of each CP can be derived using the spatial or temporal neighboring blocks of the current block. In the following description, CPMVk can mean representing the motion vector of the k-th CP. For example, referring to Figure 13 CPMV1 can be determined as the first available motion vector among the motion vectors of B2, B3, and A2, and in this case, the scan order can be B2, B3, and A2. CPMV2 can be determined as the first available motion vector among the motion vectors of B1 and B0, and in this case, the scan order can be B1 and B0. CPMV3 can be determined as one of the motion vectors of A1 and A0, and in this case, the scan order can be A1 and A0. When TMVP applies to the current block, CPMV4 can be determined as the motion vector of the time neighboring block T.
[0200] After deriving the four motion vectors for each CP, affine merge candidates can be derived based on these. The affine merge candidates can be configured by including at least two motion vectors selected from the four motion vectors of each derived CP. For example, an affine merge candidate can consist of at least one of the following in this order: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, or {CPMV1, CPMV3}. A affine merge candidate consisting of three motion vectors can be a candidate for a 6-parameter affine model. In contrast, a affine merge candidate consisting of two motion vectors can be a candidate for a 4-parameter affine model. To avoid the scaling process of motion vectors, combinations of relevant CPMVs can be ignored and not used to derive the affine merge candidate when the reference frame indices of the CPs are different from each other.
[0201] When the affine MVP mode is applied to the current block, the encoder / decoder can derive two or more CPMV predictors and CPMVs for the current block and derive the CPMV difference based on them. In this case, the CPMV difference can be signaled from the encoder to the decoder. The image decoder can derive the CPMV predictors for the current block, reconstruct the signaled CPMV difference, and then derive the CPMV for the current block based on the CPMV predictors and CPMV difference.
[0202] Furthermore, the affine MVP mode can be applied to the current block when the affine merge mode or sub-block-based TMVP is not applied (e.g., the affine merge flag or the value of merge_subblock_flag is 0). Alternatively, the affine MVP mode can be applied to the current block when the value of inter_affine_flag is 1. Additionally, the affine MVP mode can be represented as the affine CP MVP mode. The affine MVP candidate list described below can be referred to as the control point motion vector prediction sub-candidate list.
[0203] When the affine MVP pattern is applied to the current block, the affine MVP candidate list can be configured to deduce the CPMV of the current block.
[0204] In this scenario, the affine MVP candidate list can include at least one of the following: inherited affine MVP candidates, constructed affine MVP candidates, translational affine MVP candidates, or zero MVP candidates. For example, the affine MVP candidate list can include at most n (e.g., n = 2) candidates.
[0205] In this context, an inherited affine MVP candidate can refer to a candidate derived from the CPMV of neighboring blocks when the current block is encoded / decoded in affine mode. A constructed affine MVP candidate can refer to a candidate derived by generating a combination of CPMVs based on the motion vectors of CP neighboring blocks. A zero MVP candidate can refer to a candidate consisting of CPMVs with a value of 0. The derivation methods and characteristics of inherited and constructed affine MVP candidates are the same as those described above, and therefore their descriptions will be omitted.
[0206] When the maximum number of candidates in the affine MVP candidate list is 2, if the current number of candidates is less than 2, you can add affine MVP candidates for construction, translational motion affine MVP candidates, and zero MVP candidates. Specifically, translational motion affine MVP candidates can be derived in the following order.
[0207] For example, CPMV0 can be used as an affine MVP candidate when the number of candidates included in the affine MVP candidate list is less than 2 and CPMV0 used to construct the affine MVP candidate is valid. That is, affine MVP candidates whose motion vectors of CP0, CP1, and CP2 are all CPMV0 can be added to the affine MVP candidate list.
[0208] Next, when the number of candidates in the affine MVP candidate list is less than 2 and CPMV1 used to construct the affine MVP candidate is valid, CPMV1 can be used as an affine MVP candidate. That is, affine MVP candidates whose motion vectors of CP0, CP1, and CP2 are all CPMV1 can be added to the affine MVP candidate list.
[0209] Next, when the number of candidates in the affine MVP candidate list is less than 2 and the CPMV2 used to construct the affine MVP candidate is valid, CPMV2 can be used as an affine MVP candidate. That is, affine MVP candidates whose motion vectors of CP0, CP1, and CP2 are all CPMV2 can be added to the affine MVP candidate list.
[0210] Regardless of the above conditions, when the number of candidates in the affine MVP candidate list is less than 2, the Temporal Motion Vector Predictor (TMVP) for the current block can be added to the affine MVP candidate list. Regardless of the above, when the number of candidates in the affine MVP candidate list is less than 2, a zero MVP candidate can be added to the affine MVP candidate list.
[0211] Figure 14 This is a view that demonstrates the methods for generating a list of affine MVP candidates.
[0212] Reference Figure 14The flowchart shows how candidates can be added to the affine MVP candidate list in the following order: inheriting affine MVP candidates (S1610), constructing affine MVP candidates (S1620), translating affine MVP candidates (S1630), and zero MVP candidates (S1640). As described above, steps S1620 to S1640 can be performed based on whether the number of candidates included in the affine MVP candidate list in each step is less than 2.
[0213] The scan order of inheriting affine MVP candidates can be equal to the scan order of inheriting affine merge candidates. However, in the case of inheriting affine MVP candidates, only neighboring blocks that reference the same reference frame as the current block can be considered. Redundancy checks can be omitted when an inheriting affine MVP candidate is added to the affine MVP candidate list.
[0214] To derive the construction of affine MVP candidates, we can only consider... Figure 13 The spatial neighboring blocks are shown. Furthermore, the scan order for constructing affine MVP candidates can be equal to the scan order for constructing affine merge candidates. Additionally, to derive the construction of affine MVP candidates, the reference frame index of neighboring blocks can be checked, and in the scan order, the first neighboring block that is inter-coded and references the same reference frame as the current block can be used.
[0215] Overview of Sub-Block-Based Temporal Motion Vector Prediction (SbTMVP) Pattern
[0216] The following section describes in detail the sub-block-based TMVP mode as an example of inter-frame prediction modes. Based on the sub-block-based TMVP mode, the motion vector field (MVF) of the current block can be derived, and motion vectors can be derived on a sub-block basis.
[0217] Unlike the traditional TMVP mode, which executes on a unit-by-unit basis, motion vectors can be encoded / decoded on a unit-by-unit basis for applying the sub-block-based TMVP mode. Furthermore, while temporal motion vectors in the traditional TMVP mode can be derived from juxtaposed blocks in a juxtaposed frame, in the sub-block-based TMVP mode, the motion vector field can be derived from a reference block in the juxtaposed frame, specified by motion vectors derived from neighboring blocks of the current block. Hereinafter, the motion vectors derived from neighboring blocks can be referred to as the motion shift or representative motion vector of the current block.
[0218] Figure 15 This is a view of neighboring blocks in a TMVP pattern based on sub-blocks.
[0219] When a sub-block-based TMVP pattern is applied to the current block, neighboring blocks can be identified to determine motion shifts. For example, it can be done according to... Figure 15The order of blocks A1, B1, B0, and A0 is used to scan neighboring blocks for determining motion shift. As another example, the neighboring blocks for determining motion shift can be limited to specific neighboring blocks of the current block. For example, the neighboring block for determining motion shift can always be determined as block A1. When a neighboring block has a motion vector from a reference col frame, the corresponding motion vector can be determined as the motion shift. The motion vector determined as the motion shift can be referred to as the temporal motion vector. Furthermore, when the aforementioned motion vector cannot be derived from neighboring blocks, the motion shift can be set to (0,0).
[0220] Figure 16 This is a view illustrating a method for deriving motion vector fields based on a sub-block-based TMVP pattern.
[0221] Next, the reference block on the juxtaposed screen specified by the motion shift can be determined. For example, motion information based on the sub-block (motion vector or reference screen index) can be obtained from the col screen by adding the motion shift to the coordinates of the current block. Figure 16 In the example shown, assume the motion shift is the motion vector of block A1. By applying the motion shift to the current block, a sub-block (col sub-block) corresponding to each sub-block configured in the col frame can be specified. Then, using the motion information of the corresponding sub-block (col sub-block) in the col frame, the motion information of each sub-block of the current block can be derived. For example, the motion information of the corresponding sub-block can be obtained from its center position. In this case, the center position can be the position of the bottom right sample among the four samples located at the center of the corresponding sub-block. When the motion information of a specific sub-block of the col block corresponding to the current block is unavailable, the motion information of the center sub-block of the col block can be determined as the motion information of the corresponding sub-block. When deriving the motion vector of the corresponding sub-block, similar to the TMVP process described above, the reference frame index and the motion vector of the current sub-block can be switched. That is, when deriving the motion vector based on the sub-block, the POC of the reference frame of the reference block can be considered to perform motion vector scaling.
[0222] As mentioned above, the motion vector field or motion information of the current block derived from the sub-block can be used to derive the sub-block-based TMVP candidate for the current block.
[0223] The merge candidate list configured on a sub-block basis is defined below as the sub-block unit merge candidate list. The affine merge candidates and the sub-block-based TMVP candidates mentioned above can be merged to configure the sub-block unit merge candidate list.
[0224] Furthermore, a sub-block-based TMVP mode flag can be defined to indicate whether a sub-block-based TMVP mode applies to the current block. This flag can be signaled at at least one level within the current block's higher-level hierarchy (e.g., sequence, picture, slice, tile, tile group, tile, etc.). For example, the sub-block-based TMVP mode flag could be named `sps_sbtmvp_enabled_flag`. When a sub-block-based TMVP mode applies to the current block, sub-block-based TMVP candidates can first be added to the sub-block unit merge candidate list, and then affine merge candidates can be added to the sub-block unit merge candidate list. Additionally, the maximum number of candidates that can be included in the sub-block unit merge candidate list can be signaled. For example, the maximum number of candidates that can be included in the sub-block unit merge candidate list could be 5.
[0225] The size of the sub-blocks used to derive the candidate list for sub-block cell merging can be signaled or preset to M×N. For example, M×N can be 8×8. Therefore, the affine mode or the sub-block-based TMVP mode applies to the current block only if the current block size is 8×8 or larger.
[0226] The following describes implementations of the prediction execution method of this disclosure. It is possible to... Figure 4 Step S410 or Figure 6 In step S630, the following prediction execution method is performed.
[0227] A prediction block for the current block can be generated based on motion information derived from a prediction pattern. The prediction block (the predicted block) can include prediction samples (an array of prediction samples) for the current block. When the motion vector of the current block specifies partial sample units, an interpolation process can be performed, and thus, prediction samples for the current block can be derived based on reference samples, using partial samples within the reference frame as units. When affine inter-frame prediction is applied to the current block, prediction samples can be generated based on sample / sub-block units MV. When dual prediction is applied, prediction samples derived by a weighted sum or weighted average (based on phase) of prediction samples derived from L0 prediction (i.e., using MVL0 and predictions of reference frames within the reference frame list L0) and prediction samples derived from L1 prediction (i.e., using MLV1 and predictions of reference frames within the reference frame list L1) can be used as prediction samples for the current block. When dual prediction is applied and the reference frames used for L0 prediction and L1 prediction are located in different time directions relative to the current frame (i.e., if it corresponds to dual prediction and bidirectional prediction), this can be called true dual prediction.
[0228] In an image decoding device, reconstructed samples and reconstructed images can be generated based on derived prediction samples, and then an in-loop filtering process can be performed. Conversely, in an image encoding device, residual samples can be derived based on the derived prediction samples, and image information including both prediction and residual information can be encoded.
[0229] Double prediction (BCW) with CU-level weights
[0230] When dual prediction is applied to the current block as described above, the prediction samples can be derived based on a weighted average. Traditionally, the dual prediction signal (i.e., the dual prediction sample) can be derived by a simple average of the L0 prediction signal (L0 prediction sample) and the L1 prediction signal (L1 prediction sample). That is, the dual prediction sample is derived by averaging the L0 prediction sample based on the L0 reference frame and MVL0 with the L1 prediction sample based on the L1 reference frame and MVL1. However, according to this disclosure, when dual prediction is applied, the dual prediction signal (dual prediction sample) can be derived as follows by a weighted average of the L0 prediction signal and the L1 prediction signal.
[0231] [Formula 3]
[0232] P bi-pred =((8-w)*P0+w*P1+4)>>3
[0233] In equation 3 above, P bi-pred This represents the dual prediction signal (dual prediction block) derived through weighted averaging, where P0 and P1 represent the L0 prediction sample (L0 prediction block) and the L1 prediction sample (L1 prediction block), respectively. Additionally, (8-w) and w represent the weights applied to P0 and P1, respectively.
[0234] When generating a dual prediction signal via weighted averaging, five weights are allowed. For example, the weight w can be selected from {-2, 3, 4, 5, 10}. For each dual prediction CU, the weight w can be determined using one of two methods. As the first method, when the current CU is not in merging mode (non-merging CU), the weight index can be signaled along with the motion vector difference. For example, the bitstream can include information about the weight index after information about the motion vector difference. As the second method, when the current CU is in merging mode (merging CU), the weight index can be derived from neighboring blocks based on the merge candidate index (merging index).
[0235] The generation of dual prediction signals by weighted averaging can be restricted to applications only to CUs with a size comprising 256 or more samples (luminance component samples). That is, dual prediction by weighted averaging can be performed only for CUs where the product of the width and height of the current block is 256 or greater. Furthermore, the weight w can be one of the five weights described above, and different numbers of weights can be used. For example, depending on the characteristics of the current image, five weights can be used for low-latency frames, and three weights can be used for non-low-latency frames. In this case, the three weights could be {3, 4, 5}.
[0236] By applying a fast search algorithm, image encoding devices can determine weight indices without significantly increasing complexity. In this case, the fast search algorithm can be summarized as follows. Hereinafter, unequal weights mean that the weights applied to P0 and P1 are not equal. Conversely, equal weights mean that the weights applied to P0 and P1 can be equal.
[0237] - When AMVR mode is applied together with the resolution of motion vectors that adaptively change, unequal weights can be conditionally checked for each of the 1-pixel motion vector resolution and the 4-pixel motion vector resolution when the current frame is a low-latency frame.
[0238] - When affine modes are applied together and the affine mode is selected as the optimal mode for the current block, the image coding device can perform affine motion estimation (ME) with unequal weights for each mode.
[0239] - When the two reference frames used for dual prediction are equal, unequal weights can be checked conditionally only.
[0240] -Unequal weights can be skipped when predetermined conditions are met. Predetermined frames can be based on the POC distance between the current frame and the reference frame, quantization parameters (QP), time levels, etc.
[0241] The weight index of BCW can be encoded using a context-coded bin and one or more subsequent bypass-coded bins. The first context-coded bin specifies whether equal weights are used. When unequal weights are used, additional bins can be bypass-coded and signaled. Additional bins can be signaled to specify which weights to use.
[0242] Weighted prediction (WP) is a tool for efficiently encoding images, including faded ones. Based on weighted prediction, weighting parameters (weights and offsets) can be signaled to each reference frame included in each of the reference frame lists L0 and L1. Then, when motion compensation is performed, the weights and offsets can be applied to the corresponding reference frames. Weighted prediction and BCW can be used for different types of images. To avoid interaction between weighted prediction and BCW, for a CU using weighted prediction, the BCW weight index can be omitted from signaling. In this case, the weight can be inferred as 4. That is, equal weights can be applied.
[0243] In the case of a CU applying the merge mode, the weight index can be inferred from neighboring blocks based on the merge candidate index. This can be applied to both the general merge mode and the inherited affine merge mode.
[0244] When constructing an affine merging pattern, affine motion information can be configured based on the motion information of up to three blocks. The BCW weight index of the CU used to construct the affine merging pattern can be set to the BCW weight index of the first CP in the merging. That is, BCW may not be applied to CUs encoded in CIIP mode. For example, the BCW weight index of a CU encoded in CIIP mode can be set to a value with specified equal weights.
[0245] Bidirectional optical flow (BDOF)
[0246] According to this disclosure, BDOF can be used to refine dual prediction signals. When dual prediction is applied to the current block (e.g., CU), BDOF generates prediction samples by calculating refined motion information. Therefore, the process of calculating refined motion information by applying BDOF can be included in the motion information derivation steps described above.
[0247] For example, BDOF can be applied at the 4×4 sub-block level. That is, BDOF can be performed within the current block in units of 4×4 sub-blocks.
[0248] For example, BODF can be applied to CUs that satisfy at least one or all of the following conditions.
[0249] -CU uses true dual-predictive mode encoding, meaning that one of the two reference frames is displayed before the current frame in the order of display, and the other is displayed after the current frame in the order of display.
[0250] -CU is not in affine mode or ATMVP merge mode
[0251] -CU has over 64 luminance samples
[0252] -The height and width of the CU are 8 or more brightness samples.
[0253] -BCW weight index specifies equal weights, that is, equal weights are applied to L0 prediction samples and L1 prediction samples.
[0254] - Weighted prediction (WP) should not be applied to the current CU.
[0255] -CIIP mode is used for the current CU
[0256] Furthermore, BDOF can be applied only to the luminance component. However, this disclosure is not limited thereto; BDOF can be applied to the chrominance component or both the luminance and chrominance components.
[0257] BDOF mode is based on the concept of optical flow. That is, it assumes that the motion of the object is smooth. When applying BDOF, motion refinement (v) can be calculated for each 4×4 sub-block. x ,v y Motion refinement can be computed by minimizing the difference between the L0 and L1 predicted samples. Motion refinement can be used to adjust the double-predicted sample values within a 4×4 sub-block.
[0258] The process of executing BDOF will be described in more detail below.
[0259] First, the horizontal gradients of the two predicted signals can be calculated. and vertical gradient In this case, k can be 0 or 1. The gradient can be calculated by directly calculating the difference between two adjacent samples. For example, the gradient can be calculated as follows.
[0260] [Formula 4]
[0261]
[0262] In equation 4 above, I (k) (i,j) represents the sample value of the coordinate (i,j) of the predicted signal in list k (k=0,1). For example, I (0) (i,j) can represent the sample value at position (i,j) in the L0 prediction block, I (1) (i,j) can represent the sample value at position (i,j) in the L1 prediction block. In Equation 4 above, the first shift shift1 can be determined based on the bit depth of the luminance component. For example, when the bit depth of the luminance component is bitDepth, shift1 can be determined as max(6, bitDepth-6).
[0263] As described above, after calculating the gradients, the autocorrelation and cross-correlation between gradients S1, S2, S3, S5 and S6 can be calculated as follows.
[0264] [Formula 5]
[0265] S1=∑(i,j)∈Ω Abs(ψ x (i,j)), S3=∑ (i,j)∈Ω θ(i,j)·Sign(ψ x (i,j))
[0266]
[0267] in
[0268]
[0269] θ(i,j)=(I (1) (i,j)>>n b )-(I (0) (i,j)>>n b )
[0270] Where Ω is the 6×6 window surrounding the 4×4 sub-block.
[0271] In equation 5 above, n a and n b It can be set to min(1, bitDepth-11) and min(4, bitDepth-8) respectively.
[0272] Motion refinement (v) x ,v y The above autocorrelation and cross-correlation between gradients can be used to derive the following.
[0273] [Formula 6]
[0274]
[0275] in S_(2,s)=S_2&(2^(n_(S_2))-1),th′ BIO =2 13-BD .and It is the floor function.
[0276] In equation 6 above, n S2 It can be 12. Based on the derived motion refinement and gradient, the following adjustments can be performed on each sample in the 4×4 sub-block.
[0277] [Formula 7]
[0278]
[0279] Finally, the predicted samples pred of the CU with BDOF can be calculated by adjusting the double predicted samples of the CU as follows. BDOF .
[0280] [Formula 8]
[0281] pred BDOF (x,y)=(I (0) (x,y)+I (1) (x,y)+b(x,y)+o offset )>>shift
[0282] In the above formula, n a n b and n S2 The values can be 3, 6, and 12. These values can be selected to ensure that the multiplier does not exceed 15 bits in BDOF processing and that the bit width of the intermediate parameters remains within 32 bits.
[0283] To derive the gradient values, we can generate a list k (k = 0, 1) of predicted samples I that exist outside the current CU. (k) (i,j). Figure 17 This is a view of a CU that is instantiated to perform BDOF.
[0284] like Figure 17 As shown, to perform BDOF, rows / columns can be extended around the boundary of the CU. To control the computational complexity of generating predicted samples outside the boundary, the extended region ( Figure 17 The predicted samples in the white area (in the image) can be generated using a bilinear filter, CU( Figure 17 Predicted samples in the gray area (in the diagram) can be generated using a normal 8-tap motion-compensated interpolation filter. Sample values at extended locations can be used solely for gradient calculation. When sample values and / or gradient values located outside the CU boundary are needed to perform the remaining steps of BDOF processing, they can be padded (repeated) and used with the nearest neighboring sample values and / or gradient values.
[0285] When the width and / or height of a CU is greater than 16 luminance samples, the corresponding CU can be divided into sub-blocks with a width and / or height of 16 luminance samples. The boundaries of the sub-blocks can be handled in the same way as the CU boundaries described above in BDOF processing. The maximum cell size for performing BDOF processing can be limited to 16×16.
[0286] For each sub-block, it can be determined whether to perform BDOF. That is, BDOF processing for each sub-block can be skipped. For example, when the SAD value between the initial L0 predicted sample and the initial L1 predicted sample is less than a predetermined threshold, BDOF processing can be omitted from the corresponding sub-block. In this case, when the width and height of the corresponding sub-block are W and H, the predetermined threshold can be set to (8*W*(H>>1). Considering the complexity of additional SAD calculation, the SAD between the initial L0 predicted sample and the initial L1 predicted sample calculated in the DMVR processing can be reused.
[0287] BDOF may not be applied when it is available for the current block BCW, for example, when the BCW weight index specifies unequal weights. Similarly, BDOF may not be applied when it is available for the current block WP, for example, when at least one of the two reference frames has a luma_weight_lx_flag of 1. In this case, luma_weight_lx_flag may be information indicating whether the weighting factor of the WP used for lx prediction (x is 0 or 1) exists in the bitstream or information indicating whether the WP is applied to the luma component of lx prediction. BDOF may not be applied when the CU is encoded in Symmetric MVD (SMVD) mode or CIIP mode.
[0288] Optical flow prediction refinement (PROF)
[0289] The following describes a method for refining sub-block-based affine motion compensation prediction blocks by applying optical flow. Prediction samples generated by performing sub-block-based affine motion compensation can be refined based on differences derived from the optical flow equation. In this disclosure, this refinement of prediction samples can be referred to as optical flow prediction refinement (PROF). PROF enables pixel-level granular inter-frame prediction without increasing memory access bandwidth.
[0290] The parameters of the affine motion model can be used to derive the motion vectors of individual pixels in the control unit (CU). However, due to the high complexity and increased memory access bandwidth resulting from pixel-based affine motion compensation prediction, sub-block-based affine motion compensation prediction can be performed. When performing sub-block-based affine motion compensation prediction, the CU can be divided into 4×4 sub-blocks, and motion vectors can be determined for each sub-block. In this case, the motion vectors of each sub-block can be derived from the CPMV of the CU. Sub-block-based affine motion compensation involves a trade-off between coding efficiency and complexity and memory access bandwidth. Since motion vectors are derived on a sub-block basis, complexity and memory access bandwidth are reduced, but prediction accuracy is reduced.
[0291] Therefore, finer-grained motion compensation can be achieved by applying optical flow to sub-block-based affine motion compensation predictions.
[0292] As mentioned above, the brightness prediction samples can be refined by adding the difference derived from the optical flow equation after performing sub-block-based affine motion compensation. More specifically, PROF can be performed in the following four steps.
[0293] Step 1) Generate the predicted sub-block I(i,j) by performing sub-block-based affine motion compensation.
[0294] Step 2) Calculate the spatial gradient g of the predicted sub-block at each sample location. x (i,j) and g y (i,j). In this case, a 3-tap filter can be used, and the filter coefficients can be [-1,0,1]. For example, the spatial gradient can be calculated as follows.
[0295] [Formula 9]
[0296] g x (i,j)=I(i+1,j)-I(i-1,j)
[0297] g y (i,j)=I(i,j+1)-I(i,j-1)
[0298] To compute the gradient, the predicted sub-block can be extended by one pixel on each side. In this case, to reduce memory bandwidth and complexity, the pixels extending the boundary can be copied from the nearest integer number of pixels in the reference image. Therefore, additional interpolation of the filled region can be skipped.
[0299] Step 3) The brightness prediction refinement (ΔI(i,j)) can be calculated using the optical flow equation. For example, the following formula can be used.
[0300] [Formula 10]
[0301] ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j)
[0302] In the above formula, Δv(i,j) represents the difference between the pixel motion vector (pixel MV, v(i,j)) calculated at the sample position (i,j) and the sub-block MV of the sub-block to which the sample (i,j) belongs.
[0303] Figure 18 This is a view illustrating the relationship between Δv(i,j), v(i,j), and the sub-block motion vectors.
[0304] exist Figure 18 In the example shown, for instance, the motion vector v(i,j) at the top-left sample position of the current sub-block is compared with the motion vector v of the current sub-block. SB The difference between them can be represented by a thick dashed arrow, and the vector represented by the thick dashed arrow can correspond to Δv(i,j).
[0305] The affine model parameters and pixel position from the center of the sub-block remain unchanged. Therefore, Δv(i,j) can be calculated only for the first sub-block and can be reused for other sub-blocks in the same CU. Assuming the horizontal and vertical offsets from the pixel position to the center of the sub-block are x and y respectively, Δv(x,y) can be derived as follows.
[0306] [Equation 11]
[0307]
[0308] For a 4-parameter affine model
[0309]
[0310] For a 6-parameter affine model
[0311]
[0312] In the above text, (v 0x ,v 0y ), (v 1x ,v 1y ) and (v 2x ,v 2y ) correspond to the top left CPMV, top right CPMV, and bottom left CPMV respectively, and w and h represent the width and height of CU respectively.
[0313] Step 4) Finally, the final prediction block I'(i,j) can be generated based on the calculated brightness prediction refinement ΔI(i,j) and prediction sub-block I(i,j). For example, the final prediction block I' can be generated as follows.
[0314] [Equation 12]
[0315] I′(i,j)=I(i,j)+ΔI(i,j)
[0316] Overview of sub-screen
[0317] As mentioned above, quantization and dequantization of the luminance and chrominance components can be performed based on quantization parameters. Furthermore, a frame to be encoded can be divided into multiple CTUs, slices, tiles, or blocks; additionally, a frame can be divided into multiple sub-frames.
[0318] Within a frame, sub-frames can be encoded or decoded regardless of whether the preceding sub-frames have been encoded or decoded. For example, different quantizations or different resolutions can be applied to multiple sub-frames.
[0319] Furthermore, sub-pictures can be processed like individual pictures. For example, the picture to be encoded can be a 360-degree image / video or a projected picture or a packaged picture in an omnidirectional image / video.
[0320] In this implementation, a portion of the screen can be rendered or displayed based on the viewport of the user terminal (e.g., a head-mounted display). Therefore, to achieve low latency, at least one sub-screen covering the viewport in a sub-screen that constitutes a screen can be encoded or decoded preferentially or independently of the remaining sub-screens.
[0321] The encoded result of a sub-picture can be referred to as a sub-bitstream, sub-stream, or simply a bitstream. A decoding device can decode a sub-picture from a sub-bitstream, sub-stream, or bitstream. In this case, high-level syntaxes (HLS) such as PPS, SPS, VPS, and / or DPS (decoding parameter sets) can be used for encoding / decoding the sub-picture.
[0322] In this disclosure, the High-Level Syntax (HLS) may include at least one of the following: APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, or slice header syntax. For example, APS (APS syntax) or PPS (PPS syntax) may include information / parameters commonly applicable to one or more slices or frames. SPS (SPS syntax) may include information / parameters commonly applicable to one or more sequences. VPS (VPS syntax) may include information / parameters commonly applicable to multiple layers. DPS (DPS syntax) may include information / parameters commonly applicable to the entire video. For example, DPS may include information / parameters related to the concatenation of encoded video sequences (CVS).
[0323] Definition of sub-screen
[0324] Subframes can construct rectangular regions of an encoded frame. The size of a subframe can be set differently within a frame. For all frames belonging to a sequence, the size and position of a specific individual subframe can be set equally. Individual subframe sequences can be decoded independently. Tiles and slices (and CTBs) can be restricted not to cross subframe boundaries. To this end, the encoding device can perform encoding so that subframes are decoded independently. This may require semantic constraints in the bitstream. Furthermore, the arrangement of tiles, slices, and blocks within subframes can be constructed differently for each frame in a sequence.
[0325] Sub-screen design purpose
[0326] Sub-picture design aims to abstract or encapsulate a range smaller than the picture level or larger than the slice or tile group level. Therefore, VCL NAL units of a subset of Motion Constant Tiles Sets (MCTS) can be extracted from one VVC bitstream and redirected to another VVC bitstream without difficulties such as VCL-level modifications. Here, MCTS is an encoding technique that allows spatial and temporal independence between tiles. When applying MCTS, it is impossible to refer to information about tiles not included in the MCTS to which the current tile belongs. When an image is divided into MCTS and encoded, independent transmission and decoding of MCTS are possible.
[0327] The advantage of this sub-screen design is that it changes the viewing direction in a hybrid resolution viewport-dependent 360° flow scheme.
[0328] Sub-screen usage
[0329] Subpicks are required in viewport-dependent 360° schemes that provide extended real-space resolution on the viewport. For example, schemes involving tiles covering the viewport derived from 6K (6144×3072) ERP (equirectangular projection) or cubemap projection (CMP) resolutions with equivalent 4K decoding capabilities (HEVC level 5.1) are included in OMAF sections D.6.3 and D.6.4 and adopted in the VR Industry Forum Guidelines. This resolution is known to be suitable for head-mounted displays using quad HD (2560×1440) display panels.
[0330] Encoding: The content can be encoded using two spatial resolutions, one with a cube face size of 1656×1536 and the other with a cube face size of 768×768. A 6×4 tile grid can be used in both bitstreams, and MCTS can encode at each tile position.
[0331] Streaming MCTS Selection: 12 MCTS can be selected from a high-resolution bitstream, and 12 additional MCTS can be obtained from a low-resolution bitstream. Therefore, a hemisphere (180° × 180°) of streaming content can be generated from a high-resolution bitstream.
[0332] Merging and decoding using MCTS and bitstreams: MCTS of a single time instance can be merged into an encoded picture with a resolution of 1920×4608 conforming to HEVC Level 5.1. In another option for merging the picture, four tile columns have a width value of 768, two tile columns have a width value of 384, and three tile rows have a height value of 768, thus constructing a picture composed of 3840×2304 luma samples. Here, the units for width and height can be units representing the number of luma samples.
[0333] Sub-screen signaling
[0334] It is possible Figure 19 The signaling for the sub-screen is executed at the SPS level. Figure 19 This shows the syntax used to notify sub-screen syntax elements with signals in SPS. The following will describe... Figure 19 Syntax elements.
[0335] The syntax element `pic_width_max_in_luma_samples` can be referenced in SPS to specify the maximum width of each decoded frame in units of luminance samples. The value of `pic_width_max_in_luma_samples` is greater than 0 and can be an integer multiple of `andMinCbSizeY`. Here, `MinCbSizeY` is a variable specifying the minimum size of the luminance component coding block.
[0336] The syntax element `pic_height_max_in_luma_samples` can be referenced in SPS to specify the maximum height of each decoded frame in units of luminance samples. `pic_height_max_in_luma_samples` is greater than 0 and can have values that are integer multiples of `MinCbSizeY`.
[0337] The syntax element `subpic_grid_col_width_minus1` can be used to specify the width of each element of the subpic identifier grid. For example, `subpic_grid_col_width_minus1` can specify the width of each element of the subpic identifier grid in units of 4 samples, and the value obtained by adding 1 to `subpic_grid_col_width_minus1` can also specify the width of each element of the subpic identifier grid in units of 4 samples. The length of the syntax element can be Ceil(Log2(pic_width_max_in_luma_samples / 4)) bits.
[0338] Therefore, the variable NumSubPicGridCols, which specifies the number of columns in the sub-screen grid, can be derived as follows.
[0339] NumSubPicGridCols=(pic_width_max_in_luma_samples+subpic_grid_col_width_minus1*4+3) / (subpic_grid_col_width_minus1*4+4)
[0340] The syntax element `subpic_grid_row_height_minus1` can be used to specify the height of each element of the subpic identifier grid. For example, `subpic_grid_row_height_minus1` can specify the height of each element of the subpic identifier grid in units of 4 samples. The value obtained by adding 1 to `subpic_grid_row_height_minus1` can also specify the height of each element of the subpic identifier grid in units of 4 samples. The length of the syntax element can be Ceil(Log2(pic_height_max_in_luma_samples / 4)) bits.
[0341] Therefore, the variable NumSubPicGridRows, which specifies the number of rows in the sub-picture grid, can be derived as follows.
[0342] NumSubPicGridRows=(pic_height_max_in_luma_samples+subpic_grid_row_height_minus1*4+3) / (subpic_grid_row_height_minus1*4+4)
[0343] The syntax element `subpic_grid_idx[i][j]` can specify the subpic index at grid position (i,j). The length of the syntax element can be Ceil(Log2(max_subpics_minus1+1)) bits.
[0344] The variables SubPicTop[subpic_grid_idx[i][j]], SubPicLeft[subpic_grid_idx[i][j]], SubPicWidth[subpic_grid_idx[i][j]], SubPicHeight[subpic_grid_idx[i][j]] and NumSubPics can be as follows Figure 20 The algorithm is derived in the same way.
[0345] The syntax element `subpic_treated_as_pic_flag[i]` specifies whether a subpic is treated the same as a normal picture during decoding processes outside of the in-loop filtering process. For example, a first value (e.g., 0) of `subpic_treated_as_pic_flag[i]` specifies that the i-th subpic of each encoded picture in the CVS is not treated as a picture during decoding processes outside of the in-loop filtering process. A second value (e.g., 1) of `subpic_treated_as_pic_flag[i]` specifies that the i-th subpic of each encoded picture in the CVS is treated as a picture during decoding processes outside of the in-loop filtering process. When the value of `subpic_treated_as_pic_flag[i]` is not obtained from the bitstream, the value of `subpic_treated_as_pic_flag[i]` can be deduced to be the first value (e.g., 0).
[0346] The syntax element `loop_filter_across_subpic_enabled_flag[i]` specifies whether in-loop filtering is performed across the boundaries of the i-th subpic of each encoded frame belonging to CVS. For example, a first value (e.g., 0) of `loop_filter_across_subpic_enabled_flag[i]` specifies that in-loop filtering is not performed across the boundaries of the i-th subpic of each encoded frame belonging to CVS. A second value (e.g., 1) of `loop_filter_across_subpic_enabled_flag[i]` specifies that in-loop filtering can be performed across the boundaries of the i-th subpic of each encoded frame belonging to CVS. When the value of `loop_filter_across_subpic_enabled_flag[i]` is not obtained from the bitstream, the value of `loop_filter_across_subpic_enabled_flag[i]` can be deduced as the second value.
[0347] Furthermore, for bitstream consistency, the following constraints can be applied. For any two subpics, subpicA and subpicB, when the index of subpicA is less than the index of subpicB, all coded NAL units of subpicA may have a lower decoding order than all coded NAL units of subpicB. Alternatively, after decoding, the shape of the subpic needs to have a perfect left boundary and a perfect top boundary that constitute the picture boundary, or the boundary of the previously decoded subpic.
[0348] Overview of sub-picture-based encoding and decoding
[0349] The following disclosure relates to the encoding / decoding of the aforementioned screen and / or sub-screens. The encoding device can encode the current screen based on the sub-screen structure. Alternatively, the encoding device can encode at least one sub-screen that constitutes the current screen and output a (sub)bitstream including (encoded) information about at least one (encoded) sub-screen.
[0350] The decoding device can decode at least the sub-frame belonging to the current frame based on a (sub)bit stream that includes (encoded) information about at least one sub-frame.
[0351] Figure 21 This is a view illustrating a method for encoding an image using sub-pictures according to an embodiment of the encoding apparatus. The encoding apparatus can divide an input picture into multiple sub-pictures (S2110). The encoding apparatus can encode at least one sub-picture using information about the sub-pictures (S2120). For example, each sub-picture can be independently separated and encoded using information about the sub-pictures. Next, the encoding apparatus can output a bitstream by encoding image information including information about the sub-pictures (S2130). Here, the bitstream of the sub-pictures can be referred to as a sub-stream or sub-bitstream.
[0352] Information about sub-pictures will be described differently in this disclosure. For example, there may be information about whether in-loop filtering can be performed across the boundaries of sub-pictures, information about the sub-picture region, information about the grid spacing of the sub-picture, etc.
[0353] Figure 22 This is a view illustrating a method by which a decoding device, according to an embodiment, decodes an image using sub-pictures. The decoding device can obtain information about the sub-pictures from the bitstream (S2210). Next, the decoding device can deduce at least one sub-picture (S2220) and decode the at least one sub-picture (S2230).
[0354] In this way, the decoding device can decode at least one sub-picture, and thus output at least one decoded sub-picture or the current picture including at least one sub-picture. The bitstream may include sub-streams of sub-pictures or sub-bitstreams.
[0355] As mentioned above, information about sub-pictures can be constructed in the HLS of the bitstream. The decoding device can deduce at least one sub-picture based on the information about the sub-pictures. The decoding device can decode the sub-pictures based on methods such as CABAC, prediction methods, residual processing methods (transform, quantization), and in-loop filtering methods.
[0356] When outputting decoded sub-pictures, the decoded sub-pictures can be output together in the form of OPS (Output Subpicture Set). For example, when the current picture is related to a 360° or omnidirectional image / video and is partially rendered, only some sub-pictures can be decoded, and some or all of the decoded sub-pictures can be rendered according to the user's viewport.
[0357] When information indicating the availability of in-loop filtering across sub-picture boundaries is provided, the decoding device can perform in-loop filtering (e.g., deblocking filter) on the sub-picture boundaries located between two sub-pictures. Furthermore, when the sub-picture boundary is equal to the picture boundary, in-loop filtering of the sub-picture boundary may not be applied.
[0358] This disclosure relates to sub-screen-based encoding / decoding. The embodiments of this disclosure to which BDOF and PROF are applicable will be described in more detail below.
[0359] As mentioned above, applying BDOF in inter-frame prediction processing to refine reference samples in motion compensation processing can increase image compression performance. BDOF can be performed in normal mode. That is, BDOF is not performed in affine mode, GPM mode, or CIIP mode.
[0360] Figure 23 This is a view illustrating the process of deriving the predicted samples of the current block by applying BDOF.
[0361] Figure 23 The BDOF-based inter-frame prediction process can be performed by image encoding and image decoding devices.
[0362] First, in step S2310, motion information for the current block can be derived. The motion information for the current block can be derived using various methods described in this disclosure. For example, the motion information for the current block can be derived using a regular merge mode, an MMVD mode, or an AMVP mode. The motion information may include dual-predictive motion information (L0 motion information and L1 motion information). For example, L0 motion information may include MVL0 (L0 motion vector) and refIdxL0 (L0 reference frame index), and L1 motion information may include MVL1 (L1 motion vector) and refIdxL1 (L1 reference frame index).
[0363] Subsequently, the predicted sample of the current block can be derived based on the deduced motion information of the current block (S2320). Specifically, the L0 predicted sample of the current block can be derived based on the L0 motion information. In addition, the L1 predicted sample of the current block can be derived based on the L1 motion information.
[0364] Subsequently, the BDOF offset can be derived based on the derived prediction samples (S2330). The BDOF in step S2330 can be performed according to the method described in this disclosure. For example, the BDOF offset can be derived based on the gradient (based on phase) of the L0 prediction sample and the gradient (based on phase) of the L1 prediction sample.
[0365] Subsequently, based on the LX (X = 0 or 1) prediction sample and the BDOF offset, the refined prediction sample for the current block can be derived (S2340). The refined prediction sample can be used to generate the final prediction block for the current block.
[0366] Image encoding devices can be based on Figure 23 The method generates a predicted sample for the current block, which is then compared with the original sample to derive the residual sample. As mentioned above, information about the residual sample (residual information) can be included in the image / video information and encoded, and output as a bitstream. Additionally, as mentioned above, the image decoding device can be based on... Figure 23 The method generates a predicted sample of the current block and a residual sample obtained based on the residual information in the bit stream to generate the reconstructed current block.
[0367] Figure 24 This is a view illustrating the inputs and outputs of a BDOF process according to an embodiment of this disclosure.
[0368] like Figure 24 As shown, the input to BDOF processing can include the width nCbW and height CbH of the current block, the predicted sub-blocks predSamplesL0 and predSamplesL1 with a predetermined length (e.g., 2) of the boundary region extension, the prediction direction indices predFlagL0 and predFlagL1, and the reference screen indices refIdxL0 and refIdxL1. Additionally, the input to BDOF processing can also include the BDOF utilization flag bdofUtilizationFlag. In this case, the BDOF utilization flag can be an input specifying whether BDOF is applied to the corresponding sub-block on a per-sub-block basis within the current block.
[0369] In addition, BDOF processing can generate refined prediction blocks pbSamples by applying BDOF based on the input information.
[0370] Figure 25 This is a view illustrating variables used for BDOF processing according to embodiments of this disclosure. Figure 25 It can be Figure 24 Subsequent processing.
[0371] like Figure 25As shown, in order to perform BDOF processing, the input bit depth bitDepth of the current block can be set to BitDepth. Y In this case, BitDepth Y It can be derived based on information about the bit depth signaled via the bit stream. Furthermore, various right shifts can be set based on the bit depth. For example, the first shift (shift1), second shift (shift2), third shift (shift3), and fourth shift (shift4) can be set as follows: Figure 24 The derivation shown is based on bit depth. Additionally, the offset 4 can be set based on shift4. Furthermore, the variable mvRefineThres, used to specify the limiting range for motion refinement, can be set based on bit depth. This will be described below. Figure 24 The uses of the various variables described in the text.
[0372] Figure 26 This is a view illustrating a method for generating prediction samples of each sub-block in the current CU based on whether or not BDOF is applied, according to an embodiment of the present disclosure. Figure 26 It can be Figure 25 Subsequent processing.
[0373] It can execute on each sub-block in the current CU. Figure 26 The processing is shown, and in this case, the size of the sub-block can be 4×4. When the BDOF utilization flag bdofUtilizationFlag of the current sub-block is the first value (false, "0"), BDOF may not be applied to the current sub-block. In this case, the predicted samples of the current sub-block are derived from the weighted sum of the L0 and L1 predicted samples, and in this case, the weights applied to the L0 predicted samples and the weights applied to the L1 predicted samples can be the same. Figure 26 The shift4 and offset4 used in equation (1) can be Figure 17 The value set in [the document]. When the BDOF utilization flag bdofUtilizationFlag for the current sub-block is the second value (true, "1"), BDOF can be applied to the current sub-block. In this case, a prediction sample for the current sub-block can be generated by the BDOF processing according to this disclosure.
[0374] Figure 27 This is a view illustrating a method for deriving the gradient, autocorrelation, and cross-correlation of the current sub-block according to an embodiment of this disclosure. Figure 27 It can be Figure 26 Subsequent processing.
[0375] Execute on each sub-block in the current CU Figure 27 The processing shown is as follows, and in this case, the size of the sub-block can be 4×4.
[0376] according to Figure 27 Based on equations (1) and (2), the positions (h) of each sample position (x, y) in the current sub-block can be derived. x ,h y ). Then, the horizontal and vertical gradients at each sample location can be derived according to equations (3) to (6). Then, the variables used to derive autocorrelation and cross-correlation (the first intermediate parameter diff and the second intermediate parameters tempH and tempV) can be derived according to equations (7) to (9). For example, the first intermediate parameter diff can be derived using the values obtained by applying a right shift of shift2 to the predicted samples predSamplesL0 and predSamplesL1 of the current block. For example, the second intermediate parameters tempH and tempV can be derived by applying a right shift of shift3 to the sum of the gradients in the L0 direction and the L1 direction as shown in equations (8) and (9). Then, autocorrelation and cross-correlation can be derived based on the derived first and second intermediate parameters according to equations (10) to (16).
[0377] Figure 28 This illustrates the derivation of motion refinement based on embodiments of the present disclosure (v x ,v y A view of the method for deriving BDOF offsets and generating predicted samples for the current sub-block. Figure 28 It can be Figure 27 Subsequent processing.
[0378] Execute on each sub-block in the current CU Figure 28 The processing shown is as follows, and in this case, the size of the sub-block can be 4×4.
[0379] according to Figure 28 The motion refinement (v) can be derived from equations (1) and (2). x ,v y Motion refinement can be limited to the range specified by mvRefineThres. Furthermore, based on the motion refinement and gradient, the BDOF offset bdofOffset can be derived according to equation (3). The derived BDOF offset can be used to generate prediction samples pbSamples for the current sub-block according to equation (4).
[0380] By continuously executing the reference Figures 24 to 28 The described method can implement BDOF processing according to the first embodiment of this disclosure. In accordance with... Figures 24 to 28In an embodiment, the first shift shift1 is set to Max(6, bitDepth - 6), and mvRefineThres is set to 1 << Max(, bitDepth - 7). Therefore, the bit width of predSample and each parameter of BDOF according to BitDepth can be derived as shown in the following table.
[0381] [Table 1]
[0382]
[0383] As described above, by applying BDOF in the inter-frame prediction process to refine the reference sample in the motion compensation process, the compression performance of the image can be increased. BDOF can be executed in the normal mode. That is, BDOF is not executed in the case of the affine mode, GPM mode, or CIIP mode.
[0384] As a method similar to BDOF, PROF can be performed on a block encoded in the affine mode. As described above, by refining the reference sample in each 4×4 sub-block via PROF, the compression performance of the image can be increased.
[0385] According to the present disclosure, the above-mentioned affine motion (sub-block motion) information of the current block can be derived, and the affine motion information can be refined through the above PROF process or the prediction sample derived based on the affine motion information can be refined.
[0386] Figure 29 is a view illustrating the process of deriving the prediction sample of the current block by applying PROF.
[0387] Figure 29 The inter-frame prediction process based on PROF can be executed by an image encoding device and an image decoding device.
[0388] First, in step S2910, the motion information of the current block can be derived. The motion information of the current block can be derived by various methods described in the present disclosure. For example, the motion information of the current block can be derived by the methods described in the above affine mode or sub-block-based TMVP mode. The motion information may include the sub-block motion information of the current block. The sub-block motion information may include dual-prediction sub-block motion information (L0 sub-block motion information and L1 sub-block motion information). For example, the L0 sub-block motion information may include sbMVL0 (L0 sub-block motion vector) and refIdxL0 (L0 reference picture index), and the L1 sub-block motion information may include sbMVL1 (L1 sub-block motion vector) and refIdxL1 (L1 reference picture index).
[0389] Subsequently, the predicted samples of the current block can be derived based on the motion information derived from the current block (S2920). Specifically, the L0 predicted samples of each sub-block of the current block can be derived based on the L0 sub-block motion information. In addition, the L1 predicted samples of each sub-block of the current block can be derived based on the L1 sub-block motion information.
[0390] Subsequently, the PROF offset can be derived based on the derived prediction samples (S2930). The PROF in step S2930 can be performed according to the methods described in this disclosure. For example, the differential motion vector diffMv and gradient of the LX (X = 0 or 1) prediction samples can be calculated, and based on these, the PROF offset dI or ΔI can be derived according to the methods described in this disclosure. Various examples in this disclosure relate to differential motion vector derivation, gradient derivation, and / or PROF offset derivation.
[0391] Subsequently, based on the LX (X = 0 or 1) prediction samples and the PROF offset, the refined prediction samples for the current block can be derived (S2940). The refined prediction samples can be used to generate the final prediction block for the current block. For example, the final prediction block for the current block can be generated by weighted summing of the refined L0 prediction samples and the refined L1 prediction samples.
[0392] Image encoding devices can be based on Figure 29 The method generates a predicted sample for the current block, which is then compared with the original sample to derive the residual sample. Information about the residual sample (residual information) can be included in the image / video information and encoded, and output as a bitstream as described above. Additionally, the image decoding device can, as described above, base its analysis on... Figure 29 The method generates a predicted sample of the current block and a residual sample obtained based on the residual information in the bit stream to generate the reconstructed current block.
[0393] Figure 30 This is a view illustrating an example of PROF processing according to this disclosure.
[0394] according to Figure 30 For example, PROF processing can be performed using the width sbWidth, height sbHeight, predicted sub-blocks predSamples with a predetermined borderExtention length for the current sub-block, and the differential motion vector diffMv as input. In this case, for example, the predicted sub-blocks could be generated by performing affine motion compensation. As a result of performing PROF processing, refined predicted sub-blocks pbSamples can be generated.
[0395] To perform PROF processing, a predetermined first shift, shift1, can be calculated. This first shift can be based on the bit depth (BitDepth) of the luminance component.Y is derived. For example, the first shift can be derived as the maximum of 6 and (BitDepth Y - 6).
[0396] Thereafter, for each sample position (x, y) of the input prediction sub - block, the horizontal gradient gradientH,g x and the vertical gradient gradientV,g y can be calculated. The horizontal gradient and the vertical gradient can be calculated according to Figure 30 Equations (1) and (2) respectively.
[0397] Thereafter, based on the horizontal gradient, the vertical gradient, and the differential motion vector diffM v , the PROF offset dI or ΔI for each sample position can be calculated. For example, the PROF offset can be calculated according to Figure 30 Equation (3). In Equation (3), the differential motion vector diffM v used to calculate the PROF offset can mean Δv described in Figure 18 . In this case, diffMv can be clipped as follows by dmvLimit, and dmvLimit can be calculated based on BitDepth Y as follows.
[0398] [Equation 13]
[0399] dmvLimit = 1 << Max(5, BitDepthY - 7),
[0400] diffMv[x][y][i]= Clip3(-dmvLimit, dmvLimit - 1, diffMv[x][y][i])
[0401] Thereafter, based on the calculated PROF offset and the prediction sub - block predSamples, the refined prediction sub - block pbsamples can be derived. For example, the refined prediction sub - block can be derived according to Figure 30 Equation (4).
[0402] According to Figure 30 's example, the first shift shift1) can be set to Max(6, bitDepth - 6), dmvLimit can be set to 1 << Max(5, bitDepth - 7). Additionally, diffMv can be clipped within the range of [-dmvLimit, dmvLimit - 1]. Therefore, the bit width of predSample and each parameter of the PROF according to BitDepth Y can be derived as shown in the following table.
[0403] [Table 2]
[0404]
[0405] This disclosure provides various implementations for situations where a subpic is treated as a picture during subpic-based encoding / decoding (e.g., subpic_treated_as_pic_flag == 1). For example, in BDOF or PROF processing, the reference sample extraction process can be constrained to not refer to reference samples included in subpics different from the subpic to which the current block belongs. Furthermore, constraints based on the subpic being treated as a picture can be added in temporal motion vector prediction sub-derivation processing, bilinear interpolation processing of luminance samples, 8-tap interpolation filtering processing of luminance samples, and / or chrominance sample interpolation processing.
[0406] As mentioned above, in BDOF and / or PROF processing, a reference sample extending a predetermined length around the boundary of the current block can be used to calculate the gradient. However, when the current sub-frame to which the current block belongs is treated as a frame, the range of the reference sample used to calculate the gradient needs to be limited to the same sub-frame as the current sub-frame.
[0407] As mentioned above, the predicted sample can be modified using bdofOffset in the case of BDOF, and using dI in the case of PROF. bdofOffset and dI can be obtained based on the gradient. The gradient can be derived based on the difference between reference samples in the reference image. To derive the gradient, a reference sample extraction process (extracting reference samples from the reference image) and an 8-tap interpolation filtering process can be performed. The output of the reference sample extraction process can be a brightness sample at an integer pixel location.
[0408] Figure 31 This is a view that illustrates the case where a reference sample needs to be extracted across the boundaries of a sub-screen.
[0409] exist Figure 31 In this context, the region to be extracted in the reference frame (the extraction region) can be specified based on the motion vector of the current block included in the current sub-frame [1] of the current frame. In this case, the extraction region can exist across the boundary of the sub-frame [1] in the reference frame. That is, the reference sample to be extracted can be included in a sub-frame different from the current sub-frame [1].
[0410] Figure 32 yes Figure 31 A magnified view of the extracted region.
[0411] like Figure 32As shown, the reference samples to be extracted can be included in sub-screens (sub-screens [0], [2], and [3]) that are different from the current sub-screen [1]. Consider Figure 32 As shown, when a sub-screen is treated as a screen, the range of reference samples to be extracted needs to be limited to the same sub-screen as the current sub-screen.
[0412] Figure 33 This is a view illustrating a reference sample extraction process according to an embodiment of the present disclosure.
[0413] exist Figure 33 In the reference sample extraction process, information about the motion vector of the current block and the reference image refPicLX can be input. L Information about the motion vector of the current block specifies the extraction region and can be the integer sample position (xInt) derived from the motion vector of the current block. L ,yInt L ).
[0414] Figure 33 The output of the reference sample extraction process can be the prediction block of the current block. In this case, the prediction block can be the prediction block to be refined by BDOF or PROF.
[0415] like Figure 33 As shown, it can be determined whether the current subpic is treated as a picture. For example, when subpic_treated_as_pic_flag is 1, it can be determined that the current subpic is treated as a picture and the extraction area can be restricted to the same subpic as the current subpic.
[0416] When the current sub-screen is treated as the screen, such as Figure 33 As shown in Equation (1), the x-coordinate xInt of the location of the sample to be extracted can be bounded within the range of [SubPicLeftBoundaryPos, SubPicRightBoundaryPos]. SubPicLeftBoundaryPos can specify the position of the left boundary of the current sub-screen. In addition, SubPicRightBoundaryPos can specify the position of the right boundary of the current sub-screen. According to Equation (1) above, since the x-coordinate of the reference sample to be extracted exists within the range of SubPicLeftBoundaryPos to SubPicRightBoundaryPos, reference samples located outside the left or right boundary of the sub-screen are not extracted. The method for deriving SubPicLeftBoundaryPos and SubPicRightBoundaryPos will be described later.
[0417] Similarly, when the current sub-screen is treated as a screen, such as Figure 33 As shown in Equation (2), the y-coordinate yInt of the location of the sample to be extracted can be limited to the range [SubPicTopBoundaryPos, SubPicBotBoundaryPos]. SubPicTopBoundaryPos can specify the position of the upper boundary of the current sub-screen. In addition, SubPicBotBoundaryPos can specify the position of the lower boundary of the current sub-screen. According to Equation (2) above, the y-coordinate of the reference sample to be extracted exists within the range of SubPicTopBoundaryPos to SubPicBotBoundaryPos, and reference samples existing outside the upper or lower boundary of the sub-screen are not extracted. The method of deriving SubPicTopBoundaryPos and SubPicBotBoundaryPos will be described later.
[0418] When the current sub-picture is not treated as a picture, for example, when subpic_treated_as_pic_flag is 0, the coordinates of the reference sample to be extracted can be derived according to equations (3) and (4). According to equations (3) and (4), the coordinates of the reference sample to be extracted are not limited by the boundary position of the sub-picture. According to equations (3) and (4), the coordinates of the reference sample to be extracted can be limited to the range of the current picture.
[0419] Subsequently, according to equation (5), the extraction of reference samples from the reference image can be performed based on the coordinates (xInt, yInt) of the reference sample to be extracted.
[0420] Figure 34 This is a flowchart illustrating the extraction process of a reference sample according to this disclosure.
[0421] Figure 33 and Figure 34 The reference sample extraction process can be performed by image encoding and image decoding devices used to perform BDOF and / or PROF.
[0422] Reference Figure 34 First, it can be determined whether the current subpic is treated as a pic (S3410). The determination of step S3410 can be performed based on subpic_treated_as_pic_flag.
[0423] When the current sub-picture is treated as a picture (e.g., subpic_treated_as_pic_flag == 1), the extraction position can be derived (S3420), and the derived extraction position can be limited (S3430). The derivation and limiting of the extraction position can be based on... Figure 33The limiting in step S3430 can be performed by using equations (1) and (2). When the derived extraction position is outside the boundary of the current sub-screen, the corresponding extraction position can be changed to a position in the current sub-screen (e.g., the boundary position of the current sub-screen).
[0424] When the current sub-picture is not treated as a picture (e.g., subpic_treated_as_pic_flag == 0), the extraction position can be deduced (S3440). For example, the extraction position can be deduced based on... Figure 33 Equations (3) and (4) are used to execute the procedure.
[0425] Subsequently, reference sample extraction can be performed based on the extraction location derived in step S3430 or step S3440 (S3450). For example, reference sample extraction can be based on... Figure 33 The formula (5) is used to execute the operation.
[0426] according to Figure 34 In the implementation shown, when a sub-picture is treated as a picture, reference samples outside the boundary of the current sub-picture may not be extracted during the reference sample extraction process. That is, reference samples in the reference picture that belong to a sub-picture different from the current sub-picture may not be referenced.
[0427] The fractional sample interpolation process according to another embodiment of this disclosure will now be described.
[0428] As described above, when the motion vector of the current block specifies a fractional sample unit, an interpolation process can be performed, and the predicted sample of the current block can be derived accordingly based on the reference sample of the fractional sample unit in the reference frame.
[0429] Figure 35 This is a view illustrating a portion of the fractional sample interpolation process according to this disclosure.
[0430] To perform fractional sample interpolation, the variables fRefWidth and fRefHeight can be derived. For example... Figure 35 As shown, fRefWidth and fRefHeight can be derived differently depending on whether the current subpic is treated as a pic (subpic_treated_as_pic_flag).
[0431] When the current subpic is treated as a picture, for example, when subpic_treated_as_pic_flag is 1, fRefWidth and fRefHeight can be derived as follows.
[0432] fRefWidth=(SubPicWidth[SubPicIdx]*(subpic_grid_col_width_minus1+1)*4)
[0433] fRefHeight=(SubPicHeight[SubPicIdx]*(subpic_grid_row_height_minus1+1)*4)
[0434] In the above text, `SubPicWidth[SubPicIdx]` and `SubPicHeight[SubPicIdx]` can refer to the width and height of the current subpicture, respectively. In this case, the width and height of the subpicture can be represented by the number of grid cells configured in the subpicture. For example, a subpicture width of 4 can mean that the corresponding subpicture includes four grid cells in the horizontal direction. Additionally, `subpic_grid_col_width_minus1` and `subpic_grid_row_height_minus1` can refer to the width and height of the grid cells configured in the subpicture, respectively. In this case, the width and height of the grid cells can be represented in units of 4 pixels. For example, a grid width of 4 can mean that the grid width is 16 (4×4) pixels.
[0435] When the current subpic is not treated as a picture, for example, when subpic_treated_as_pic_flag is 0, fRefWidth and fRefHeight can be derived as the width PicOutputWidthL and height PicOutputHeightL of the output image of the reference picture, respectively.
[0436] The fRefWidth and fRefHeight derived as described above can be used based on Figure 35 Equations (1) and (2) are used to derive the scaling factors hori_scale_fp and vert_scale_fp. Subsequently, based on equations (3) to (6), the position (refx) of the reference sample specified by the motion vector can be derived. L ,refy L Based on the position derived from the reference sample, the integer position (xInt) can be derived. L ,yInt L ) and fractional position (xFrac) L ,yFrac L The aforementioned reference sample extraction or 8-tap interpolation filtering can be performed based on the derived integer and / or fractional positions.
[0437] The following describes another embodiment of the sbTMVP derivation method according to this disclosure.
[0438] For reference Figure 16 The described motion shift can be derived and applied to the current block, thereby specifying the sub-blocks (col sub-blocks) in the col frame that correspond to the various sub-blocks configured in the current block. Subsequently, using the motion information of the corresponding sub-blocks (col sub-blocks) in the col frame, the motion information of each sub-block in the current block can be derived. Based on the derived motion information of the sub-blocks, sbTMVP can be derived.
[0439] Figure 36 This is a view illustrating a portion of the sbTMVP derivation method according to this disclosure.
[0440] Figure 36 This illustrates a portion of the processing following the derivation of the motion shift of the current block during the sbTMVP derivation process. Specifically, Figure 36 The process shown includes deriving the position of the corresponding col sub-block for each sub-block configured in the current block.
[0441] according to Figure 36 First, based on equations (1) and (2), the position (xSb, ySb) of the lower right center sample of the current sub-block can be derived. Then, the position (xColSb, yColSb) of the col sub-block can be derived based on whether the current sub-picture is treated as a picture (e.g., subpic_treated_as_pic_flag).
[0442] Specifically, when subpic_treated_as_pic_flag is 1, the y-coordinate yColSb of the col sub-block can be derived according to equation (3). In this case, yColSb derived from ySb and motion shift tempMv can be bounded to a position specified by SubPicTopBoundaryPos, SubPicBotBoundaryPos, and the y-coordinate yCtb of the current CTB. Through equation (3), the y-coordinate of the col sub-block exists in the current CTB and the current sub-picture. Additionally, when subpic_treated_as_pic_flag is 0, the y-coordinate yColSb of the col sub-block can be derived according to equation (4). In this case, yColSb derived from ySb and motion shift tempMv can be bounded to a position specified by the y-coordinate yCtb of the current CTB and the height of the current picture. Through equation (4), the y-coordinate of the col sub-block exists in the current CTB and the current picture.
[0443] Similarly, when subpic_treated_as_pic_flag is 1, the x-coordinate xColSb of the col sub-block can be derived according to equation (5). In this case, xColSb, derived from xSb and the motion shift tempMv, can be bounded to a position specified by SubPicLeftBoundaryPos, SubPicRightBoundaryPos, and the x-coordinate xCtb of the current CTB. Through equation (5), the x-coordinate of the col sub-block exists in the current CTB and the current sub-picture. Additionally, when subpic_treated_as_pic_flag is 0, the x-coordinate yColSb of the col sub-block can be derived according to equation (6). In this case, xColSb, derived from xSb and the motion shift tempMv, can be bounded to a position specified by the x-coordinate xCtb of the current CTB and the width of the current picture. Through equation (6), the x-coordinate of the col sub-block exists in the current CTB and the current picture.
[0444] According to this disclosure, when the current sub-screen is treated as a screen, the col sub-block used to derive sbTMVP for the current sub-block exists in the same sub-screen as the current sub-screen.
[0445] The following describes a method for deriving the sub-screen boundary position according to another embodiment of the present disclosure.
[0446] Figure 37 This is a view illustrating a method for deriving the position of a sub-screen boundary according to this disclosure.
[0447] According to this disclosure, the positions of the left boundary of a sub-screen, SubPicLeftBoundaryPos, the right boundary of a sub-screen, SubPicRightBoundaryPos, the top boundary of a sub-screen, SubPicTopBoundaryPos, and the bottom boundary of a sub-screen, SubPicBotBoundaryPos, can be derived. In this disclosure, SubPicIdx can be an index used to identify each sub-screen in the current screen.
[0448] When the current sub-picture specified by SubPicIdx is treated as a picture, SubPicLeftBoundaryPos can be derived based on the position information SubPicLeft of the predetermined cell specifying the left position of the current sub-picture and the width information subpic_grid_col_width_minus1 of the corresponding cell. The predetermined cell can be a grid. However, this disclosure is not limited to this; for example, the predetermined cell can be a CTU. When the predetermined cell is a CTU, the left position of the current sub-picture can be derived as the product of the position information of the CTU cell specifying the left position of the current sub-picture and the size of the CTU.
[0449] When the current sub-screen is treated as a screen, SubPicRightBoundaryPos can be derived based on the position information of a predetermined cell specifying the left position of the current sub-screen, the width information of a predetermined cell specifying the width of the current sub-screen, and the width information of the corresponding cell. For example, the position information of a predetermined cell specifying the right position of the current sub-screen can be derived by adding the left position and the width of the current sub-screen. SubPicRightBoundaryPos can be derived by performing a "-1" operation on the final calculated value. The predetermined cell can be a grid. However, this disclosure is not limited to this; for example, the predetermined cell can be a CTU. When the predetermined cell is a CTU, the right position of the current sub-screen can be derived by performing a "-1" operation on the product of the position information of the CTU cell specifying the right position of the current sub-screen and the size of the CTU. The position information of the CTU cell specifying the right position of the current sub-screen can be derived as the sum of the position information of the CTU cell specifying the left position of the current sub-screen and the width information of the CTU cell specifying the width of the current sub-screen.
[0450] Similarly, SubPicTopBoundaryPos can be derived based on the position information SubPicTop of a predetermined cell specifying the upper position of the current sub-screen and the height information subpic_grid_low_height_minus1 of the corresponding cell. The predetermined cell can be a grid. However, this disclosure is not limited thereto; for example, the predetermined cell can be a CTU. When the predetermined cell is a CTU, the upper position of the current sub-screen can be derived as the product of the position information of the CTU cell specifying the upper position of the current sub-screen and the size of the CTU.
[0451] When the current sub-screen is treated as a screen, SubPicBotBoundaryPos can be derived based on the position information of a predetermined cell at the upper position of the current sub-screen, the height information of a predetermined position at the height of the current sub-screen, and the height of the corresponding cell. For example, the position information of a predetermined cell at the lower position of the current sub-screen can be derived by adding the upper position of the current sub-screen to the height of the current sub-screen. SubPicBotBoundaryPos can be derived by performing a "-1" operation on the final calculated value. The predetermined cell can be a grid. However, this disclosure is not limited to this; for example, the predetermined cell can be a CTU. When the predetermined cell is a CTU, the lower position of the current sub-screen can be derived by performing a "-1" operation on the product of the position information of the CTU cell at the lower position of the current sub-screen and the size of the CTU. The position information of the CTU cell at the lower position of the current sub-screen can be derived as the sum of the position information of the CTU cell at the upper position of the current sub-screen and the height information of the CTU cell at the height of the current sub-screen.
[0452] The various embodiments described in this disclosure can be implemented individually or in combination with other embodiments. Alternatively, some embodiments can be added to another embodiment, or some embodiments can be replaced by other embodiments.
[0453] Although the exemplary methods of this disclosure described above are represented as a series of operations for clarity of description, they are not intended to limit the order in which the steps are performed, and these steps may be performed simultaneously or in different orders if necessary. To implement the method according to the invention, the described steps may further include other steps, including steps in addition to some steps, or may include additional steps in addition to some steps.
[0454] In this disclosure, the image encoding device or image decoding device that performs a predetermined operation (step) can perform an operation (step) that confirms the execution conditions or circumstances of the corresponding operation (step). For example, if it is described that a predetermined operation is performed when predetermined conditions are met, the image encoding device or image decoding device can perform the predetermined operation after determining whether the predetermined conditions are met.
[0455] The various embodiments of this disclosure are not a list of all possible combinations and are intended to describe representative aspects of this disclosure; the matters described in the various embodiments may be applied independently or in combination of two or more.
[0456] Various embodiments of this disclosure can be implemented in hardware, firmware, software, or a combination thereof. When this disclosure is implemented in hardware, it can be implemented using application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, etc.
[0457] Furthermore, the image decoding and image encoding devices applying the embodiments of this disclosure can be included in multimedia broadcasting transmission and receiving devices, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video chat devices, real-time communication devices such as video communication, mobile streaming devices, storage media, cameras, video-on-demand (VoD) service providers, OTT (over-the-top) video devices, internet streaming service providers, three-dimensional (3D) video devices, video telephony devices, medical video devices, etc., and can be used to process video signals or data signals. For example, OTT video devices can include game consoles, Blu-ray players, internet access televisions, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.
[0458] Figure 38 This is a view illustrating a content streaming system to which embodiments of the present disclosure can be applied.
[0459] like Figure 38 As shown, the content streaming system applying the embodiments of this disclosure may mainly include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.
[0460] The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream and then sends the bitstream to the streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server can be omitted.
[0461] The bitstream can be generated by an image encoding method or image encoding device applying the embodiments of this disclosure, and the stream server can temporarily store the bitstream during the sending or receiving of the bitstream.
[0462] A streaming server sends multimedia data to a user's device based on a request from a web server, and the web server acts as a medium for informing the user of the service. When a user requests a service from the web server, the web server can deliver it to the streaming server, and the streaming server can send the multimedia data to the user. In this scenario, the content streaming system may include a separate control server. In this case, the control server is used to control the commands / responses between devices in the content streaming system.
[0463] A streaming server can receive content from media storage devices and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined period of time.
[0464] Examples of user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays), digital televisions, desktop computers, digital signage, etc.
[0465] In a content streaming system, each server can operate as a distributed server, in which case the data received from each server can be distributed.
[0466] The scope of this disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) for enabling the operation of methods according to various embodiments to be executed on a device or computer, and non-transitory computer-readable media having such software or commands stored thereon and executable on a device or computer.
[0467] Industrial applicability
[0468] The embodiments disclosed herein can be used to encode or decode images.
Claims
1. An image decoding method performed by an image decoding device, the image decoding method comprising the following steps: Determine whether bidirectional optical flow (BDOF) or optical flow prediction refinement (PROF) is applied to the current block; Based on BDOF or PROF applied to the current block, a predicted sample of the current block is generated from a reference image of the current block based on the motion information of the current block; as well as By applying BDOF or PROF to the current block based on the generated prediction samples, refined prediction samples for the current block are derived. The step of generating the predicted sample for the current block is based on whether the current sub-image including the current block is treated as an image. The step of generating the predicted sample for the current block includes obtaining a reference sample at a specified location within the reference image and shifting the value of the reference sample to the left. The location within the reference image is limited to a predetermined range. Wherein, based on the current sub-image being treated as an image, the predetermined range is specified by the boundary position of the current sub-image. The position within the reference image includes x-coordinates and y-coordinates. The x-coordinate is limited to the range of the left and right boundary positions of the current sub-image, and The y-coordinate is limited to the range between the upper and lower boundaries of the current sub-image. The left boundary position of the current sub-image is derived as the product of the left position of the current sub-image in units of a predetermined size and the predetermined size. The right boundary position of the current sub-image is derived by performing a "-1" operation on the product of the right position of the current sub-image (in units of the predetermined size) and the predetermined size. Wherein, the upper boundary position of the current sub-image is derived as the product of the upper position of the current sub-image in units of the predetermined size and the predetermined size, and The lower boundary position of the current sub-image is derived by performing a "-1" operation on the product of the lower position of the current sub-image in units of the predetermined size and the predetermined size.
2. The image decoding method according to claim 1, wherein, Whether the current sub-image is treated as an image is determined based on flag information signaled via a bitstream.
3. The image decoding method according to claim 2, wherein, The flag information is communicated via a signal using the sequence parameter set SPS.
4. The image decoding method according to claim 1, wherein, The predetermined size is a grid or CTU.
5. The image decoding method according to claim 1, wherein, Since the current sub-image is not treated as an image, the predetermined range is the range of the current image that includes the current block.
6. An image encoding method performed by an image encoding device, the image encoding method comprising the following steps: Determine whether bidirectional optical flow (BDOF) or optical flow prediction refinement (PROF) is applied to the current block; Based on BDOF or PROF applied to the current block, a predicted sample of the current block is generated from a reference image of the current block based on the motion information of the current block; as well as By applying BDOF or PROF to the current block based on the generated prediction samples, refined prediction samples for the current block are derived. The step of generating the predicted sample for the current block is based on whether the current sub-image including the current block is treated as an image. The step of generating the predicted sample for the current block includes obtaining a reference sample at a specified location within the reference image and shifting the value of the reference sample to the left. The location within the reference image is limited to a predetermined range. Wherein, based on the current sub-image being treated as an image, the predetermined range is specified by the boundary position of the current sub-image. The position within the reference image includes x-coordinates and y-coordinates. The x-coordinate is limited to the range of the left and right boundary positions of the current sub-image, and The y-coordinate is limited to the range between the upper and lower boundaries of the current sub-image. The left boundary position of the current sub-image is derived as the product of the left position of the current sub-image in units of a predetermined size and the predetermined size. The right boundary position of the current sub-image is derived by performing a "-1" operation on the product of the right position of the current sub-image (in units of the predetermined size) and the predetermined size. Wherein, the upper boundary position of the current sub-image is derived as the product of the upper position of the current sub-image in units of the predetermined size and the predetermined size, and The lower boundary position of the current sub-image is derived by performing a "-1" operation on the product of the lower position of the current sub-image in units of the predetermined size and the predetermined size.
7. A method for transmitting a bit stream, said bit stream being generated by the image encoding method according to claim 6.
Citation Information
Patent Citations
Video coding / decoding method and device
CN101610413A
Inter prediction refinement based on bi-directional optical flow (BIO)
US20180262773A1